Run open models at production scale

Access leading open models, optimize inference, choose how you deploy, and adapt models with your own data — all through one platform.

Discover the latest open models

Compare leading text, code and reasoning models, then choose the right fit for your workload. Explore all models

MiniMax M3

A 128B open model with a 1M-token context window for reasoning, coding and long-context workloads.

Kimi K3

A next-generation open model built for advanced software engineering and autonomous coding-agent workflows.

DeepSeek V4 Pro

An 862B open model with a 1M-token context window for complex coding and long-context workloads.

From model access to production in one platform

Everything you need to serve, scale and improve open models — without stitching together separate tools

Optimized inference

Run open models on serving stacks tuned for latency, throughput and cost.

Scalable infrastructure

Move from early testing to sustained production traffic on infrastructure built for demanding AI workloads.

Post-training

Adapt open models with your own data, then deploy the resulting model on Token Factory.

Choose how you deploy

Start quickly with serverless inference, or choose dedicated capacity for greater production control.

Serverless inference

Explore models and start building through a familiar API, without planning infrastructure upfront.

Best for: evaluation, prototyping and variable workloads.

Dedicated endpoints

Run models on isolated capacity when you need greater control over performance, scaling, model configuration or regional placement.

Best for: sustained traffic, custom models and production-critical workloads

Built for real production workloads

Put open models to work across high-value applications.

Coding agents

Build code generation, review, debugging and agentic software-development experiences.

AI search

Build agents that can retrieve, verify and reason over current information.

Fraud and risk

Use AI agents to investigate complex signals, support risk workflows and keep human review in the loop.

Proven in production

See how teams use Token Factory to run demanding AI workloads at scale.

Revolut is powering FinCrime agents and a support-chat system at global scale.

Up to 1.2 million support tickets per month

Sword Health scales Dawn from a 30B model to 200B+ while keeping safety-sensitive conversations responsive.

Tail latency: over 20 seconds → under 12 seconds

Prosus is running high-volume open-model workloads through Dedicated Endpoints.

Up to 200B tokens/day · up to 26× lower cost

Build on open models. Keep control of what matters

Choose the models that fit your workload, deploy them the way your product needs, and improve them with your own data. Token Factory gives you one path from first experiment to production.