
SGLang Omni: Serving Omni & Multimodal Models | Nebius Science Paper Club
Join a session with Jiaxin Deng of the SGLang Omni team.
SGLang-Omni is a high-performance, multi-stage serving runtime for omni, speech, and TTS models, extending SGLang beyond single-loop autoregressive decoding to pipelines that mix text, image, audio, and video in and out.
The talk will cover SGLang Omni’s design along with broader SGLang LLM inference work, followed by Q&A and open discussion with Nebius researchers and the community.
About Nebius Science Paper Club
Nebius Science Paper Club is a webinar series led by Nebius researchers. Each session brings together paper authors and practitioners to discuss new ideas, research, and discoveries in AI.
The session is open to researchers, engineers, students, and anyone curious about the paper.
Key takeaways
SGLang Omni coordinates multiple inference stages to serve voice and omni models efficiently.
It reuses SGLang’s autoregressive inference capabilities and adds scheduling across components such as encoders, language models and audio decoders.
Each stage can use an execution strategy suited to its workload.
Try Nebius AI Cloud console today
