DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Agents live or die on retrieval. Miss the right passage, code block, or document, and the agent reasons from the wrong context — wasting tokens and reducing answer quality. And agents retrieve constantly: decomposing tasks, rewriting queries, searching memory, inspecting code. For retrieval to be the default for every agent decision, the embedding model has to be accurate, fast, and cost-effective, all at once.
That's exactly what NVIDIA Nemotron 3 Embed delivers — and it's now available on DeepInfra.
Enterprise retrieval usually forces tradeoffs: better accuracy means paying more per query, larger models are harder to serve, and fast answers can mean missing the right content. Nemotron 3 Embed comes in two sizes so you don't have to pick one corner of that triangle:
Whether you're optimizing for maximum retrieval quality or production-scale throughput, NVIDIA Nemotron 3 Embed provides the right model.
Multi-turn agents retrieve repeatedly — for planning, long-term memory, code understanding, tool use, and multi-step reasoning. Weak retrieval means more turns, more tokens, and more hallucination risk. Strong retrieval reduces irrelevant context and keeps agents grounded. Nemotron 3 Embed is a stronger retrieval layer for agentic retrieval, query decomposition, query rewriting, code retrieval, enterprise search, and RAG applications.
Nemotron 3 Embed ships with open weights, datasets, and recipes, so you can inspect it, tune it, and fine-tune it for your domain. No black box, no lock-in. On DeepInfra, both sizes are available now through our standard OpenAI-compatible API, so you can move from experimentation to production without changing infrastructure.
Generating embeddings requires only a single OpenAI-compatible API call:
from openai import OpenAI
client = OpenAI(
base_url="https://api.deepinfra.com/v1/openai",
api_key="$DEEPINFRA_TOKEN",
)
resp = client.embeddings.create(
model="nvidia/Nemotron-3-Embed-8B",
input=["How do I reset my password?"],
)
print(resp.data[0].embedding[:8])
For high-throughput workloads, simply swap in nvidia/Nemotron-3-Embed-1B-BF16, or nvidia/Nemotron-3-Embed-1B-NVFP4 for maximum throughput on NVIDIA Blackwell GPUs. All three models are available on DeepInfra today.
Have questions or need help? Reach out at feedback@deepinfra.com, join our Discord, or connect with us on X (@DeepInfra) — we're happy to help.
MiMo-V2.5 Provider Pricing and Deployment Guide<p>MiMo-V2.5 is worth paying attention to because it puts three things developers usually have to trade off into the same conversation: open weights, a 1 million-token model design, and pricing that can be unusually low depending on where you buy it. On Xiaomi’s first-party API, Artificial Analysis lists MiMo-V2.5 at $0.14 per 1M input tokens […]</p>
Qwen3.5 9B API Benchmarks: Latency, Throughput & Cost<p>About Qwen3.5 9B Qwen3.5 9B is the flagship of Alibaba’s Qwen3.5 Small Model Series, released on March 2, 2026. It is a dense multimodal model combining Gated Delta Networks (a form of linear attention) with a sparse Mixture-of-Experts system, enabling higher throughput and lower latency during inference compared to traditional dense architectures. The architecture utilizes […]</p>
Open vs Closed Source AI Models: Intelligence, Price & Speed Compared<p>The LLM landscape in 2026 looks nothing like it did two years ago. Back then the assumption was simple: if you wanted the best model, you paid OpenAI or Anthropic, and that was that. Open source models were a respectable second tier, good for experimentation, fine-tuning, and budget workloads, but not quite there for serious […]</p>
© 2026 DeepInfra. All rights reserved.