SlimSpec: faster speculative decoding without cutting the vocabulary
SlimSpec is a low-rank draft LM-head architecture for speculative decoding that reduces projection cost while preserving full-vocabulary support. In this article, we examine why the draft LM head becomes a bottleneck, how SlimSpec differs from vocabulary-truncation methods and what our EAGLE-3 experiments show about end-to-end decoding performance.
LangChain tunes Deep Agents for NVIDIA Nemotron 3 Ultra: top open-model accuracy at 10x lower cost with Nebius Agents Blueprint
LangChain has released a Deep Agents profile tuned for NVIDIA Nemotron 3 Ultra, achieving top agent accuracy among open models at roughly 10x lower cost than closed alternatives — with the model itself untouched. The Nebius Agents Blueprint already pairs Deep Agents with Nemotron 3 Ultra on Nebius Token Factory, and support for the tuned profile is coming soon.
Train the draft model for your workload
Custom Speculator Training is now in Token Factory. Move from production data to a workload-specific draft model, then deploy it alongside the base model, in one workflow.
Nebius AI Cloud “Aether 3.6”: Operating production AI with more control, efficiency, and confidence
Introducing Aether 3.6. Our latest quarterly platform release responds to how AI teams now operate at scale and with maturity, delivering a smoother developer experience, a stronger security and compliance foundation, advances across our in-house storage portfolio, and natural-language control over your AI cloud.
AI infrastructure that speaks your language
Introducing Nebius Echo, an AI agent built right into the Nebius console. Ask questions, inspect resources, and create infrastructure in plain language. No setup required; find it right inside the console.
MLPerf® Training 6.0: Leading NVIDIA HGX B300 and competitive NVIDIA GB300 NVL72 results on NVIDIA Blackwell Ultra systems
MLPerf® Training 6.0 results are in. Across six configurations on NVIDIA Blackwell Ultra systems, Nebius posted the #1 single-node HGX B300 times for Llama-3.1-8B and GPT-OSS 20B pre-training and came within 3.1% of the fastest GB300 NVL72 result on all three benchmarks at 72 GPUs. This post covers the full results and methodology.
Nebius Cloud Logs are now available in Datadog: Trace AI incidents across every layer
Nebius AI Cloud Logs now stream into Datadog Log Management. If you already run on Datadog, you can investigate your Nebius workloads alongside the rest of your stack and trace an AI incident across every layer without switching tools.
Introducing the Nebius Agents Blueprint: open architecture for production-ready AI agents
Today we’re introducing the Nebius Agents Blueprint, an open reference architecture for building, operating, and continuously improving AI agents in production. This post covers the six-component composable stack — inference, orchestration, retrieval, grounding, observability, and simulation — and the case study behind it: a compliance audit agent that achieved 72% lower cost and 20% higher precision over a GPT-based prototype by improving the system, not the model.
Building a compliance audit agent using Nebius Agents Blueprint
Using Nebius Agents Blueprint, we built a regulatory compliance audit agent and ran the same audit task through four configurations — from a $470 prototype to a production system at 72% lower cost with 1.00 recall. This post covers what changed at each step, what each configuration revealed about the one before it, and the production failures that only became visible once the operational architecture existed to find them.
NVIDIA retail AI blueprints, now running on Nebius
Nebius has collaborated with NVIDIA to bring two retail AI blueprints to production on Nebius infrastructure: the NVIDIA Agentic Commerce Blueprint and the NVIDIA Retail Catalog Enrichment Blueprint for automated product content. Both are open-source reference architectures built on NVIDIA NIMs that retailers and developers can customize and deploy today with a 1-click deployment on Nebius AI Cloud.
Building transaction foundation models on Nebius AI Cloud
Today, we’re exploring how transaction foundation models move from developer examples to production systems. Using NVIDIA’s TFM blueprint and Revolut’s PRAGMA model, we show how Nebius AI Cloud supports the full lifecycle — from GPU-accelerated data preparation and multi-node training to managed inference on Token Factory.
Run physical AI workflows, not glue code
The Nebius Physical AI Workbench turns NVIDIA Cosmos 3, NVIDIA Isaac Sim, NVIDIA Isaac GR00T and other Physical AI tools into composable building blocks that agents can wire together. We are building it in the open, and it is available now on GitHub.