MLPerf® Storage v3.0 results: Leading RetinaNet training on Nebius Object Storage
The MLPerf® Storage v3.0 results are in. Nebius Object Storage fed 768 simulated NVIDIA B200 accelerators on the RetinaNet training workload, the highest accelerator count of any submission this round, and ran the highest Unet3D accelerator count among object storage submissions. This post covers the full results, methodology, and what they mean for training directly on object storage.
Incident post-mortem analysis: us-central1 service disruption on August 19, 2026
A detailed analysis of the incident on August 19, 2026 that led to service outages in the us-central1 region. The incident was caused by a storm-related event at the data center facility that disabled the cooling infrastructure.
Nebius Token Factory Becomes First AI Cloud to Adopt NVIDIA Groq 3 LPX
Nebius is he first AI cloud to adopt NVIDIA Groq 3 LPX, to Nebius Token Factory, adding generation-optimized performance purpose-built for agentic AI. Independently benchmarked at 3,400 output tokens per second for a single user on Google Gemma 4 31B, NVIDIA Groq 3 LPX pairs with NVIDIA Vera Rubin NVL72 so developers get it through the same Token Factory catalog and API they already use.
Serverless AI Builders Challenge: winners announced
The results are in for the Nebius Serverless AI Builders Challenge. In total, 35 projects qualified across AI & ML, Scientific AI & Healthcare, and robotics, each with a public repo and a technical blog post. Discover the top three winners, special award recipients, and cast your vote for the Community Choice award.
Nebius Academy: closing the gap between AI tools and the people who use them
Nebius Academy is the education and research hub of Nebius. This guide lays out everything it offers in one place: AI adoption programs and readiness assessments for organizations, role-based AI cloud certifications with free credits, open courses for professionals, and research grants that fund academic work with GPU compute.
The whole AI stack moves faster in the open
Nebius was built on the conviction that AI must stay competitive and diverse. Open source is essential to keeping it that way.
Training speculative decoders: removing the logits and attention bottlenecks
Training speculative decoding draft heads at scale introduces two bottlenecks: the memory cost of the language-modeling loss and the attention pattern used by EAGLE-style draft heads. In this article, we show how Streaming Cross Entropy removes the vocabulary-logits bottleneck and block-sparse FlashAttention addresses the draft-head attention pattern, improving memory and performance efficiency for long-context training.
Inside the Nebius + PyTorch DeepSeek V3 recipe: NVSHMEM and DeepEP for wide expert parallelism
Nebius and PyTorch built a high-performance DeepSeek V3 training recipe on TorchTitan, running on 256 Nebius B200 GPUs with Soperator. This post dives deeper into why GPU-initiated RDMA with NVSHMEM and DeepEP make such an impact in improving performance for large MoE models, and walks through the full IBGDA and DeepEP setup you can run on Nebius today.
SlimSpec: faster speculative decoding without cutting the vocabulary
SlimSpec is a low-rank draft LM-head architecture for speculative decoding that reduces projection cost while preserving full-vocabulary support. In this article, we examine why the draft LM head becomes a bottleneck, how SlimSpec differs from vocabulary-truncation methods and what our EAGLE-3 experiments show about end-to-end decoding performance.
LangChain tunes Deep Agents for NVIDIA Nemotron 3 Ultra: top open-model accuracy at 10x lower cost with Nebius Agents Blueprint
LangChain has released a Deep Agents profile tuned for NVIDIA Nemotron 3 Ultra, achieving top agent accuracy among open models at roughly 10x lower cost than closed alternatives — with the model itself untouched. The Nebius Agents Blueprint already pairs Deep Agents with Nemotron 3 Ultra on Nebius Token Factory, and support for the tuned profile is coming soon.
Train the draft model for your workload
Custom Speculator Training is now in Token Factory. Move from production data to a workload-specific draft model, then deploy it alongside the base model, in one workflow.
AI infrastructure that speaks your language
Introducing Nebius Echo, an AI agent built right into the Nebius console. Ask questions, inspect resources, and create infrastructure in plain language. No setup required; find it right inside the console.