We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

A Milestone on Our Journey Building DeepInfra and Scaling Open Source AI Infrastructure
Published on 2025.04.22 by Yessen Kanapin, Co-Founder of DeepInfra
A Milestone on Our Journey Building DeepInfra and Scaling Open Source AI Infrastructure

Today we're excited to share that DeepInfra has raised $18 million in Series A funding, led by Felicis and our earliest believer and advisor Georges Harik.

When we founded DeepInfra in 2022, we saw a clear gap: while enormous resources were being poured into training AI models, the infrastructure needed to run these models in production was lagging behind.

The past two years have been a whirlwind. We've scaled our processing volume by over 8,000x since our seed stage. What started as a bet on AI infrastructure has quickly become a critical service for developers deploying increasingly sophisticated models.

Our growth accelerated following the emergence of "thinking models" like DeepSeek. These open source alternatives demonstrated that the innovation cycle in AI was becoming even more rapid than anticipated, requiring significantly more computation during inference.

The reality of deploying modern AI models is challenging for most organizations. Running these models requires significant compute resources, specialized hardware like GPUs that are difficult to acquire, and deep expertise in infrastructure optimization. Most companies simply can't afford the investment or overcome the supply chain challenges to build this infrastructure themselves.

This challenge has shaped our approach from day one. After years of scaling systems to hundreds of millions of users before founding this company, we've developed a set of core principles that guide how we build DeepInfra:

  1. We believe reliability is non-negotiable when your service powers critical applications. We design for zero downtime because AI infrastructure must be as dependable as the electricity powering your office.
  2. We've learned that performance creates competitive advantage. Our obsession with fast time-to-first-token and optimal GPU utilization isn't technical vanity – it directly impacts our customers' user experience and cost structure.
  3. We're convinced that privacy builds lasting trust. Our strict no-logging policy for user prompts isn't just a feature – it's a fundamental commitment to our customers' data sovereignty.
  4. And we know that deep infrastructure expertise matters at every layer. Understanding the full stack from hardware to application allows us to deliver superior performance while controlling costs.

These principles have guided our approach as we've expanded our computing capacity, recently receiving a large shipment of NVIDIA Blackwell GPUs with more on order to support our rapid growth. You can see how this funding injection will be put to good use.

To our customers who have trusted us with their production workloads: thank you. We're just getting started as we continue building the infrastructure that powers the next generation of AI applications.

Follow us on X (formerly Twitter) and LinkedIn to stay updated on our journey. We look forward to sharing more exciting developments in the coming months.

Servers

Related articles
Power the Next Era of Image Generation with FLUX.2 Visual Intelligence on DeepInfraPower the Next Era of Image Generation with FLUX.2 Visual Intelligence on DeepInfraDeepInfra is excited to support FLUX.2 from day zero, bringing the newest visual intelligence model from Black Forest Labs to our platform at launch. We make it straightforward for developers, creators, and enterprises to run the model with high performance, transparent pricing, and an API designed for productivity.
Qwen3.5 397B A17B API Benchmarks: Latency, Throughput & CostQwen3.5 397B A17B API Benchmarks: Latency, Throughput & Cost<p>About Qwen3.5 397B A17B Qwen3.5 397B A17B is Alibaba Cloud&#8217;s largest and most capable multimodal foundation model, released in February 2026. It features a hybrid Mixture-of-Experts (MoE) architecture with 397 billion total parameters and 17 billion active parameters per inference pass, utilizing 512 experts with a routing mechanism selecting a subset per token. This sparse [&hellip;]</p>
Qwen3.5 9B API Benchmarks: Latency, Throughput & CostQwen3.5 9B API Benchmarks: Latency, Throughput & Cost<p>About Qwen3.5 9B Qwen3.5 9B is the flagship of Alibaba&#8217;s Qwen3.5 Small Model Series, released on March 2, 2026. It is a dense multimodal model combining Gated Delta Networks (a form of linear attention) with a sparse Mixture-of-Experts system, enabling higher throughput and lower latency during inference compared to traditional dense architectures. The architecture utilizes [&hellip;]</p>