JIAXUAN LUO

Research Engineer · Streaming Multimodal Systems

I'm a research engineer focused on streaming multimodal systems—real-time speech/omni models, post-training, and efficient inference optimization.

I work across model and systems research: developing learning and agent capabilities for live multimodal streams, then making them run efficiently under tight latency, memory, and reliability constraints.

Currently, I work on efficient TPU/GPU inference for autonomous-driving systems at Waymo (Alphabet) and contribute to SGLang-Omni as a core contributor. Previously, I spent three years at Alibaba building large-scale recommendation systems and conducted research with Prof. Lei Li at CMU LTI.

News

Jul 15, 2026 Released AutoTerm-SST, an adaptive terminology-memory system for simultaneous speech translation.
Jul 01, 2026 Leading the TTS runtime refactor in SGLang-Omni.
Apr 20, 2026 Our work on security vulnerabilities in agent-generated code was accepted to ICML 2026.
Feb 01, 2026 Joined Waymo (Alphabet) as a Machine Learning Engineer, working on efficient TPU/GPU inference for autonomous-driving systems.
Jan 29, 2026 Released RASST, retrieval-augmented simultaneous speech translation over partial audio.
Jan 15, 2025 Joined Prof. Lei Li’s group at CMU LTI as a research assistant.

Selected Work

Open source · ML systems

SGLang-Omni

Core contributor leading the TTS runtime refactor and working on mixed-chunk scheduling, tensor-parallel correctness, topology-aware collectives, and shared multimodal serving infrastructure.

Research system · Streaming speech

AutoTerm-SST

Adaptive, budgeted terminology memory for streaming speech translation without per-session glossary setup or domain selection.

Research · Retrieval for Speech LLMs

RASST

Multi-scale speech–text retrieval over partial audio, paired with training that teaches a Speech LLM when to use terminology hints during simultaneous translation.

Research infrastructure · Speech RL

sst-rl-framework

A reusable GRPO post-training framework for speech RL with modular rollouts, rewards, typed adapters, and task-specific plug-ins.

Research · GUI agent memory

CausalCache

Fixed-budget visual-memory reallocation for long-horizon GUI agents, combining history-gated KV adaptation with conditional restoration of task-relevant past screens.

Research · Agent security

SUSVIBES

A real-world benchmark for the security of code generated by software-engineering agents, published at ICML 2026.

Selected Publications

  1. CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
    Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang, and Haojian Huang
    Under review at AAAI 2027, 2026
  2. Retrieval-Augmented Simultaneous Speech Translation
    Jiaxuan Luo, Siqi Ouyang, Jiaxing Xu, and Lei Li
    arXiv preprint arXiv:2601.22777, 2026
  3. Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
    Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, and Lei Li
    In International Conference on Machine Learning, 2026
  4. AutoTerm-SST: Adaptive Terminology Memory for Simultaneous Speech Translation
    Jiaxuan Luo, Siqi Ouyang, and Lei Li
    Under review for EMNLP 2026 System Demonstrations, 2026

Experience

2026–present

Waymo (Alphabet)

Machine Learning Engineer

Efficient TPU/GPU inference for autonomous-driving systems, with a focus on determinism, memory efficiency, and runtime performance.

2026–present

SGLang-Omni

Core Contributor

Real-time multimodal inference and shared TTS runtime systems.

2025–2026

Carnegie Mellon University, LTI

Research Assistant

Streaming speech agents, retrieval, reinforcement learning, and efficient inference; advised by Prof. Lei Li.

2025

TikTok

Research Engineer Intern

Post-training and low-latency serving for reasoning agents.

2021–2024

Alibaba Group

Algorithm Engineer (MLE)

Built multi-objective recommendation models (MMoE/MTL), causal uplift estimators, and constrained RL policies for large-scale marketplace ranking and coupon allocation. Improved long-term metric stability and budget-constrained ROI, while delivering ~9% relative lift in targeted conversion experiments.

Education

Johns Hopkins University

M.S. in Computer Science

2025

Central South University

B.S. in Computer Science · Turing Honors Program

2021