JIAXUAN LUO
Research Engineer · Streaming Multimodal Systems
I'm a research engineer focused on streaming multimodal systems—real-time speech/omni models, post-training, and efficient inference optimization.
I work across model and systems research: developing learning and agent capabilities for live multimodal streams, then making them run efficiently under tight latency, memory, and reliability constraints.
Currently, I work on efficient TPU/GPU inference for autonomous-driving systems at Waymo (Alphabet) and contribute to SGLang-Omni as a core contributor. Previously, I spent three years at Alibaba building large-scale recommendation systems and conducted research with Prof. Lei Li at CMU LTI.
News
| Jul 15, 2026 | Released AutoTerm-SST, an adaptive terminology-memory system for simultaneous speech translation. |
|---|---|
| Jul 01, 2026 | Leading the TTS runtime refactor in SGLang-Omni. |
| Apr 20, 2026 | Our work on security vulnerabilities in agent-generated code was accepted to ICML 2026. |
| Feb 01, 2026 | Joined Waymo (Alphabet) as a Machine Learning Engineer, working on efficient TPU/GPU inference for autonomous-driving systems. |
| Jan 29, 2026 | Released RASST, retrieval-augmented simultaneous speech translation over partial audio. |
| Jan 15, 2025 | Joined Prof. Lei Li’s group at CMU LTI as a research assistant. |
Selected Work
SGLang-Omni
Core contributor leading the TTS runtime refactor and working on mixed-chunk scheduling, tensor-parallel correctness, topology-aware collectives, and shared multimodal serving infrastructure.
AutoTerm-SST
Adaptive, budgeted terminology memory for streaming speech translation without per-session glossary setup or domain selection.
RASST
Multi-scale speech–text retrieval over partial audio, paired with training that teaches a Speech LLM when to use terminology hints during simultaneous translation.
sst-rl-framework
A reusable GRPO post-training framework for speech RL with modular rollouts, rewards, typed adapters, and task-specific plug-ins.
CausalCache
Fixed-budget visual-memory reallocation for long-horizon GUI agents, combining history-gated KV adaptation with conditional restoration of task-relevant past screens.
SUSVIBES
A real-world benchmark for the security of code generated by software-engineering agents, published at ICML 2026.
Selected Publications
-
CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI AgentsUnder review at AAAI 2027, 2026
-
Experience
Waymo (Alphabet)
Efficient TPU/GPU inference for autonomous-driving systems, with a focus on determinism, memory efficiency, and runtime performance.
Carnegie Mellon University, LTI
Streaming speech agents, retrieval, reinforcement learning, and efficient inference; advised by Prof. Lei Li.
TikTok
Post-training and low-latency serving for reasoning agents.
Alibaba Group
Built multi-objective recommendation models (MMoE/MTL), causal uplift estimators, and constrained RL policies for large-scale marketplace ranking and coupon allocation. Improved long-term metric stability and budget-constrained ROI, while delivering ~9% relative lift in targeted conversion experiments.
Education
Johns Hopkins University
M.S. in Computer Science
Central South University
B.S. in Computer Science · Turing Honors Program