All the latest features and innovation our engineering teams have been delivering across our AI Cloud. This quarter, we gave teams more control over cloud operations and spending, faster paths from development to deployment, and more ways to run models and agents in production.
What changed: Cloud Interconnect provides a dedicated private connection between Nebius AI Cloud and your network through neutral interconnection facilities, with MACsec encryption on the last mile. Aggregate available 100G and 400G ports for higher bandwidth.
Why it matters: Connect on-premises data centers, colocation sites, or another public cloud without routing traffic over the public internet, and without a bilateral agreement between providers.
What changed:Preemptible VMs now use dynamic pricing that reflects demand for each GPU type and region. View the current price and up to 30 days of price history before launching a workload.
Why it matters: Use pricing trends to schedule workloads that can tolerate interruption and take advantage of lower-cost capacity when demand is quieter.
What changed:Data Transfer Service now supports Nebius Shared Filesystem, through filesystem buckets, as a source or destination, alongside S3-compatible object storage.
Why it matters: Move datasets into and out of Nebius Shared Filesystem without writing custom transfer scripts or operating your own transfer infrastructure. Simplify migration and data movement across regions.
What changed: Serverless AI removes more of the setup between an idea and a running workload
Serverless AI Devlab: Launch a persistent CPU or GPU workspace using a curated template or your own image, and access it through a managed HTTPS URL.
Direct Python execution: Run Python files as Serverless Jobs without first building a Docker container.
File injection: Supply local files or inline content to Jobs, Endpoints and Devlab environments at launch.
Pre-built templates: Start with validated configurations, aligned framework versions and a recommended GPU type.
Why it matters: Spend less time packaging code, configuring environments and passing files into workloads, and more time developing, testing and iterating.
What changed: Native cloud security posture management support, starting with Wiz, gives security teams visibility into misconfigurations in their Nebius environment.
Why it matters: Teams using Wiz can assess their Nebius infrastructure alongside the rest of their cloud estate.
What changed: Every public Nebius service is now accessible over HTTP and JSON using your existing IAM credentials. A published OpenAPI specification helps tools and agents discover and use the APIs.
Why it matters: Bring Nebius infrastructure management into scripts, CI pipelines and agent workflows using a familiar API format.
What changed: Custom Speculator Training lets you train a workload-specific draft model on your own data and deploy it alongside a supported base model.
Why it matters: Tailoring speculative decoding to your traffic can improve inference throughput and latency for recurring workloads.
What changed: The expanded model lineup includes GLM-5.2, GLM-5.3, GLM-5.3-Flash, Kimi K2.7 Code, Kimi K3, NVIDIA Nemotron 3.5 Lightning, DeepSeek V4 Flash-0731, DeepSeek V4 Pro 0813, DeepSeek V4.1 Flash, Qwen 3.8-27B and K2 Horizon. An updated Model Catalog and more comprehensive model pages make discovery easier and clarify the distinction between models and deployed endpoints.
Why it matters: Explore a broader selection of models for your workloads and understand your deployment options in one place.
What changed:Expanded observability includes active GPU counts, GPU utilization, KV-cache hit rate and clearer streaming-latency reporting.
Why it matters: Understand how serving resources are being used, investigate bottlenecks and make better-informed decisions about capacity and performance.
What changed: Batch Inference 2.0 (in private beta for selected customers) runs asynchronous inference on dedicated batch capacity that scales to zero between jobs.
Why it matters: Process workloads that do not need an immediate response without keeping batch compute running between jobs.
What changed: Tavily improved evidence ranking, deduplication, contradiction handling, index coverage and freshness. In its published evaluation, Tavily ranked #1 in accuracy across the providers tested on SealQA-Hard, SealQA-0 and SimpleQA-Verified.
Why it matters: Agents get more relevant evidence, with fewer redundant or conflicting results to work through.
What changed: Tavily expanded coverage of people, companies and news to support workflows across sales, go-to-market, compliance, onboarding and market intelligence.
Why it matters: Help agents research prospects, understand markets, verify entities and gather evidence for risk assessments.
What changed: New integrations connect Tavily’s real-time search and extraction directly to Convex, OpenCode and NanoClaw.
Why it matters: Give applications, coding agents and assistants in Slack, Telegram and the CLI access to current web evidence with less custom integration work.