Why the Best Inference Configuration Can Fail in Production
An interactive guide to the workload, memory, routing, and capacity assumptions behind an inference benchmark winner.
Inference engines, KV cache, batching, and the GPUs underneath. I contribute to SGLang and vLLM, and write here about what I learn along the way.
An interactive guide to the workload, memory, routing, and capacity assumptions behind an inference benchmark winner.
A short pointer to the LMSYS article on work I contributed to: serving MOSS-TTS-Local Transformer v1.5 on SGLang-Omni with native streaming at 48 kHz stereo.
How we serve Ming-Omni in SGLang: the unified multimodal architecture, and the optimizations — an encoder kernel fix, tensor parallelism, CFM CUDA-graph capture, and streaming TTS — behind fast omni serving.