Technical Portfolio + Evidence Hub

Tran Dang Khoa — Senior Software Engineer

AI Systems, Model Inference & Platforms

Large-model deployment, multimodal AI, semantic retrieval, speech synthesis, image generation, accelerator-aware inference, model portability and reproducible releases.

Selected engineering evidence

  • 8 TPU devices used for distributed large-model inference.
  • 68/68 intended Top-1 checks for bilingual and cross-modal retrieval.
  • 141 automated tests passed for MOSS-TTS qualification.
  • 1247/1247 weights loaded in TranslateGemma verification.
  • Approximately 100K-record semantic-search corpus / vector-index scale.
  • Two public JAX / TPU converted model distributions.

Featured AI systems

  • Gemma 4 31B — TPU multimodal service.
  • Mage-Flow-Turbo / Edit-Turbo — JAX / Keras 3 / Orbax conversion and TPU qualification.
  • MOSS-TTS-v1.5 — original checkpoint qualified on two Tesla T4 GPUs.
  • VoxCPM2 — CPU, one-T4 and two-T4 runtime.
  • WeMM-Embedding-9B — bilingual and cross-modal retrieval.
  • Qwen3 Embedding + Reranker — two-stage semantic retrieval.
  • Muse-Glimmer-30B + DFlash2 — self-hosted speculative-decoding stack.