We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open-Weight AI Model ComparisonPublished on 2026.07.28 by DeepInfraKimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open-Weight AI Model Comparison

In the span of three months, three Chinese AI labs shipped open-weight models that individually would have rewritten the frontier story. Together, they signal something more structural: the open-weight tier is no longer a budget alternative to closed models. Kimi K3 (Moonshot AI, July 2026 — now also available through DeepInfra), DeepSeek V4 Pro (DeepSeek, […]

Hosted Agents: your own always-on AI agent, from $13/monthPublished on 2026.07.22 by DeepInfraHosted Agents: your own always-on AI agent, from $13/month

One click gives you a dedicated, isolated AI agent, pre-wired to fast inference and ready to work the moment it boots. No VMs, no SSH hardening, no patching. From $13/month, and idle is free.

We Benchmarked NVIDIA Vera, the CPU for Agents. Here's What We MeasuredPublished on 2026.07.21 by DeepInfraWe Benchmarked NVIDIA Vera, the CPU for Agents. Here's What We Measured

DeepInfra runs AI agents in production, so when NVIDIA built a CPU for agents, we measured it ourselves with our own harness, our own agent, and a methodology we locked before the hardware arrived.

DeepInfra Now Serves NVIDIA Nemotron 3 Embed: Frontier Retrieval for RAG and AgentsPublished on 2026.07.16 by Aray SultanbekovaDeepInfra Now Serves NVIDIA Nemotron 3 Embed: Frontier Retrieval for RAG and Agents

DeepInfra now serves NVIDIA Nemotron 3 Embed, the industry's leading open embedding model for enterprise search and agentic retrieval, available today in both 8B and 1B sizes.

Introducing the Flex Service Tier: Cheaper Inference When You Can WaitPublished on 2026.07.14 by DeepInfraIntroducing the Flex Service Tier: Cheaper Inference When You Can Wait

Run latency-tolerant work at 0.8× real-time — best-effort, sheddable, same OpenAI-compatible API.

Frontier-Level Agents on Open Models: LangChain Deep Agents + NVIDIA Nemotron 3 Ultra, Live on DeepInfraPublished on 2026.07.08 by Aray SultanbekovaFrontier-Level Agents on Open Models: LangChain Deep Agents + NVIDIA Nemotron 3 Ultra, Live on DeepInfra

Open models have reached frontier-level agent performance. Starting today, you can point LangChain Deep Agents at NVIDIA Nemotron 3 Ultra running on DeepInfra and get top-tier agent accuracy at roughly 10x lower cost than leading closed models.