We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

How DeepInfra Built on NVIDIA's Inference Stack and Why It Paid Off
Published on 2026.06.30 by Aray Sultanbekova
How DeepInfra Built on NVIDIA's Inference Stack and Why It Paid Off

How DeepInfra Built on NVIDIA's Inference Stack and Why It Paid Off

When we built DeepInfra, we made a deliberate bet on the NVIDIA inference software stack. Not as a hedge — as a conviction. Today, that bet is paying off in ways that are easy to measure.

The stack

DeepInfra runs on Blackwell-generation GPUs including B300s and our inference stack is built on TensorRT-LLM and NVIDIA Dynamo for distributed serving. We use ModelOpt to quantize models to NVFP4 weights, which reduces memory and compute requirements without meaningful accuracy loss.

These aren't just checkboxes. The components of the NVIDIA inference software stack work together — NVFP4 reduces memory pressure. Dynamo handles KV-aware routing and disaggregated prefill/decode. TensorRT-LLM makes sure the kernels are actually optimized for the hardware underneath. When they work together, the result is production economics that go beyond benchmark performance. Read more about it here.

What it looks like in practice: DeepSeek V4

The clearest proof point is what happened when DeepSeek V4 dropped.

We served it in production on day 0. We launched on Hopper first, then measured performance on B300. The result: 4x better performance. A workload that previously required 4×H200 now runs on a single B300 at higher tokens per second.

"That's not a theoretical improvement. That's a direct reduction in infrastructure cost for the same production traffic."

It happened because the software stack — TensorRT-LLM, Dynamo, NVFP4 — was already in place and ready to take advantage of the hardware the moment it was available. A direct line to the NVIDIA team also helped. Getting a new frontier model deployed and serving production traffic at scale from day one requires more than good software, it requires coordination.

What this means for developers building on DeepInfra

As NVIDIA continues to optimize the inference software stack, those improvements flow through to every model on DeepInfra. Developers don't have to do anything. The same API call gets faster and cheaper as the stack compounds underneath.

That's the reason we invested early in this ecosystem and why we keep investing. The flywheel is real.

Get started

Explore the models available on DeepInfra, including DeepSeek V4 Pro on NVIDIA Blackwell, at deepinfra.com.

Related articles
DeepSeek V4 Pro: Model Overview, Features & Performance GuideDeepSeek V4 Pro: Model Overview, Features & Performance Guide<p>DeepSeek V4 Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model from DeepSeek, released on April 24, 2026 under the MIT license. It is designed for advanced reasoning, complex software engineering, and long-running agentic tasks, and arrives alongside DeepSeek-V4-Flash, a lighter 284B-parameter variant built for faster, lower-cost inference. The V4 series is DeepSeek&#8217;s first two-tier lineup [&hellip;]</p>
Introducing NVIDIA Nemotron 3 Nano Omni on DeepInfraIntroducing NVIDIA Nemotron 3 Nano Omni on DeepInfraDeepInfra is an official launch partner for NVIDIA Nemotron 3 Nano Omni, the first multimodal model in the Nemotron 3 family — a single open model that understands images, video, audio, documents, and text in one unified inference pass.
Kimi K3: 2.8T Open-Weight Multimodal ModelKimi K3: 2.8T Open-Weight Multimodal Model<p>Kimi K3, developed by Moonshot AI, represents a landmark achievement in open-source artificial intelligence. As a 2.8-trillion-parameter native multimodal Mixture-of-Experts (MoE) model, Kimi K3 is engineered to handle demanding computational tasks, from complex software engineering and long-horizon agentic workflows to deep scientific research. By combining a one-million-token context window with its architectural innovations, Kimi K3 [&hellip;]</p>