Nebius Token Factory Becomes First AI Cloud to Adopt NVIDIA Groq 3 LPX

Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX, boosting inference token generation for NVIDIA Vera Rubin NVL72 in Nebius Token Factory. NVIDIA Groq 3 LPX is purpose-built for extremely fast token generation, which is the rate at which tokens are produced and value is realized for an individual user or agent. An extension of the NVIDIA Vera Rubin platform and in conjunction with Token Factory production inference platform, Groq 3 LPX gives developers a major boost in inference performance for agentic use cases where speed and responsiveness matter.

Production agentic workloads are changing inference requirements

Inference has two distinct jobs: context and generation. Context involves ingesting long prompts, large codebases, and persistent multi-turn agent state. Generation produces the tokens a user or agent actually sees and acts on.

An agent can make dozens of sequential model calls to complete a single task. Each step depends on the one before it, so latency compounds across the entire workflow. As a result, the speed of generation increasingly determines how quickly an agent can get useful work done.

NVIDIA Groq 3 LPX is purpose-built for long-context, low-latency performance. Each NVIDIA LPX rack couples 256 LPU accelerators with 128 GB of on-chip SRAM and 640 TB/s of scale-up bandwidth, fully liquid-cooled using the NVIDIA MGX rack architecture.

Fastest Gemma 4 31B inference performance ever recorded

Artificial Analysis benchmarked NVIDIA Groq 3 LPX running Gemma 4 31B at 3,400 output tokens per second for a single user, the fastest performance ever recorded for that model. With NVIDIA Vera Rubin NVL72, NVIDIA projects up to 35x higher inference throughput per megawatt for 2T-parameter models at long-context and low-latency than GB200 NVL72, a measure of what extreme codesign unlocks at agentic-scale token volume.

Faster generation, no re-architecture

Token Factory provides production inference across leading open models, with serverless and dedicated endpoints, autoscaling, observability, function calling, structured outputs, and the infrastructure required to operate models at scale.

A fast token means very little if getting it requires re-architecting a stack for new silicon. Nebius’s approach is the opposite: NVIDIA Vera Rubin NVL72 and Groq 3 LPX will run on the same Token Factory platform developers are already running at scale, initially supporting a subset of models alongside familiar features like autoscaling, and observability.

Developers get speed, scale, and lowest token cost

Token Factory already ships the pieces multi-agent systems depend on: native function calling, structured JSON outputs, and built-in safety guardrails for tool-calling agents, on top of dedicated endpoints with sub-second latency targets. With NVIDIA Groq 3 LPX on that same platform, developers building multi-agent pipelines, real-time coding assistants, and latency-sensitive reasoning applications get faster generation the same way they call every other Token Factory endpoint today: same platform, same operating model, faster tokens where the workload demands it.

Get started on Token Factory

For teams already running production traffic on Token Factory, adding NVIDIA Groq 3 LPX to a latency-sensitive workload is a model-selection change, not a migration: no new SDK, no new vendor relationship, no new billing setup.

Explore Nebius Token Factory

Explore Nebius AI Cloud

Sign in to save this post