DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

The NVIDIA Nemotron 3 Super is a state-of-the-art 120-billion parameter hybrid Mixture-of-Experts (MoE) model designed to bridge the gap between high-compute efficiency and extreme accuracy. Engineered specifically for the next generation of AI development, Nemotron 3 Super excels in multi-agent applications, specialized agentic systems, and complex reasoning tasks. By utilizing a sophisticated architecture that activates only 12 billion parameters at any given time, the model provides the performance of a massive LLM with the agility required for real-time, collaborative AI workflows.
NVIDIA has optimized the Nemotron 3 Super (specifically the 120B-A12B variant) to run numerous collaborating agents simultaneously on a single GPU. This is achieved through a Latent Mixture-of-Experts (LatentMoE) framework, which projects tokens into a smaller latent dimension for routing, significantly reducing compute overhead.
The Nemotron 3 Super demonstrates specialized capabilities in agentic workflows, scientific reasoning, and autonomous software engineering. It consistently outperforms peer models in its parameter class across critical benchmarks.
| Benchmark Category | Benchmark Name | Score | Metric |
|---|---|---|---|
| General Knowledge | MMLU-Pro | 83.73 | Accuracy (%) |
| Reasoning | AIME25 (No Tools) | 90.21 | Accuracy (%) |
| Reasoning | HMMT Feb25 (No Tools) | 93.67 | Accuracy (%) |
| Coding | LiveCodeBench (v5) | 81.19 | Pass@1 (%) |
| Human Preference | Arena-Hard-V2 | 73.88 | Score |
| Benchmark | Nemotron 3 Super | Qwen3.5-122B-A10B | GPT-OSS-120B |
|---|---|---|---|
| MMLU-Pro | 83.73 | 86.70 | 81.00 |
| HMMT Feb25 | 93.67 | 91.40 | 90.00 |
| RULER @ 1M Context | 91.75 | 91.33 | 22.30 |
| SWE-Bench | 60.47 | 66.40 | 41.90 |
Nemotron 3 Super is available for public deployment via the DeepInfra inference cloud. DeepInfra provides an OpenAI-compatible endpoint, making it easy for developers to integrate the model into existing applications.
Access requires an API key obtained from your DeepInfra dashboard. Include this key in the Authorization header of your requests:
Authorization: Bearer YOUR_DEEPINFRA_API_KEY
import requests
import os
DEEPINFRA_API_KEY = os.getenv("DEEPINFRA_API_KEY", "YOUR_DEEPINFRA_API_KEY")
url = "https://api.deepinfra.com/v1/openai/chat/completions"
headers = {
"Authorization": f"Bearer {DEEPINFRA_API_KEY}",
"Content-Type": "application/json"
}
payload = {
"model": "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B",
"messages": [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "Explain the concept of quantum entanglement in simple terms."}
],
"max_new_tokens": 150,
"temperature": 0.7
}
response = requests.post(url, headers=headers, json=payload)
print(response.json())cURL Example
curl -X POST \
https://api.deepinfra.com/v1/openai/chat/completions \
-H "Authorization: Bearer YOUR_DEEPINFRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B",
"messages": [
{"role": "user", "content": "Explain quantum entanglement."}
],
"max_new_tokens": 150
}'DeepInfra offers a highly competitive, usage-based pricing model for Nemotron 3 Super, allowing developers to scale from prototyping to enterprise production without massive upfront costs.
For users looking to deploy the model on private infrastructure, the following hardware configurations are recommended:
* Minimum: 8× NVIDIA H100-80GB GPUs.
* Optimized: Fully compatible with NVIDIA Grace Blackwell (GB200) systems. On B200/B300 hardware, the BF16 checkpoint can fit on as few as 2 GPUs due to increased HBM capacity.
The NVIDIA Nemotron 3 Super represents a significant milestone in Mixture-of-Experts technology. By combining a massive 120B parameter knowledge base with a highly efficient 12B active parameter execution, it offers a unique value proposition: enterprise-grade reasoning and multi-agent collaboration at a fraction of the traditional compute cost. Whether you are building autonomous software agents or processing million-token documents, Nemotron 3 Super provides the accuracy and efficiency required for modern AI systems.
For the latest updates and community milestones, visit the official NVIDIA news section or the DeepInfra blog.
Step 3.7 Flash is Live on DeepInfra: An Agentic, Multimodal Model Built for ProductionStepFun's Step 3.7 Flash is now live on DeepInfra. It's a 198B-parameter sparse MoE vision-language model with just ~11B active parameters per token, a 256K context window, and three selectable reasoning levels—purpose-built for high-throughput agentic workflows that combine perception, search, and reasoning.
GLM-5.1 Pricing Guide: API Cost Comparison & Analysis<p>Provider choice for GLM-5.1 is a real economic decision. Across 10 benchmarked API providers, blended pricing runs from $0.74 to $1.70 per 1M tokens, output speed from 33.8 to 175.2 t/s, and the fastest provider is 5.2x quicker than the slowest. For teams deploying at scale, that spread determines whether this model fits a production […]</p>
Gemma 4 Model Overview: Features, Architecture & Use Cases<p>Gemma 4 is Google DeepMind’s latest family of open-weight models, released on April 3, 2026 under the Apache 2.0 license. The family spans four model sizes — from edge-optimized variants for mobile devices to a 31B dense model for server-side deployments — with every model supporting multimodal input, built-in reasoning, and a context window of […]</p>
© 2026 DeepInfra. All rights reserved.