DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

AI agents have grown up. People are running them as real assistants now: reading email, triaging Slack, kicking off cron jobs, calling tools, and remembering context for weeks at a time. The hard part isn't building one anymore. It's the jump from "I made something cool on my laptop" to "I have an agent that actually runs around the clock."
That jump is all infrastructure, and it's tedious. You rent a VM. You lock down SSH. You install a runtime, wire up model API keys, and set up TLS so the dashboard isn't wide open to the internet. Then you keep it patched, figure out how to update the framework without losing your data, and pay for the box even on the weekends when nobody's using it.
Deep Infra Hosted Agents takes all of that off your plate. One click gives you a dedicated, isolated agent that's pre-wired to fast inference and ready to work the moment it boots. Starts at $13/month.
We host two flavors, both built on the same proven runtime:
OpenClaw - the dashboard agent. OpenClaw comes with a full web dashboard. Chat with your agent, manage its skills, plug in MCP servers, schedule cron tasks, and peek at its memory and workspace, all from the browser. It hooks into the channels you already use: Slack, Telegram, WhatsApp, Signal, and more. If you want an assistant the whole team can talk to, this is the one.
Hermes - the self improving agent. Hermes is the lean, SSH-first sibling. Same engine, no dashboard. You ssh straight in with your own key and drive it from the command line. It's made for pros who'd rather live in the terminal: scriptable, quiet, no UI in the way.
Each agent type comes in two host tiers. There's a smaller one to keep costs low for light, always-listening assistants, and a larger one with more RAM for heavier workloads, bigger workspaces, and more tools running at once. Pick what fits and resize as you grow. The entry tier starts at $13/month.
🔌 Inference-ready from the first second. This is the part most setups get wrong. Every hosted agent boots already wired to Deep Infra's model APIs, the same platform serving inference at scale for thousands of customers. Nothing to paste, no endpoint to configure, no "why isn't my agent responding" rabbit hole. It can think the moment it starts, on top of fast, affordable, frontier-class models.
🖱️ One-click setup. No provisioning scripts, no Docker, no certificates. You click, and a fully configured agent comes up for you: runtime, dashboard, SSH, TLS, secrets, all of it.
🔄 One-click updates. Agent frameworks move fast. When a new version drops, you update with a single click, and your memory, workspace, conversations, and configuration come right along with it. The framework swaps underneath while your data stays put. No migrations, no fingers crossed.
💾 Automatic backups and point-in-time restore. Your agent's entire state is snapshotted automatically every day, and you can take a backup on demand any time you're about to try something risky. If something breaks or you just want to rewind, restore to any saved checkpoint in a click. It's an undo button for your whole agent. We keep a rolling history of recent snapshots and copy them across regions, so your data isn't riding on a single machine.
💤 Stopped instances cost nothing. Pause an agent and you pay $0 for compute while it's idle, but its disk, memory, and full state stick around. Start it again and it's back in seconds, right where you left off. Run it hard during the week, stop it for the weekend, and only pay for the time you actually used. Plenty of always-on services bill you flat whether the thing is working or asleep. This one doesn't.
🧠 Long-term memory that persists. Your agent doesn't get wiped between sessions. Files, conversations, learned context, and workspace state live on a durable disk that survives stops, starts, updates, and restores. The agent you talk to next month is the same one that remembers what you told it today.
🔗 Connect everything. Agents plug into the channels and tools you already use (Slack, Telegram, WhatsApp, Signal, and more) and extend through skills, MCP servers, and scheduled cron jobs. It can listen, act, and run on a schedule without you in the loop.
🔐 Truly isolated and private. Each agent runs in its own environment with its own kernel, not a shared container, so its world is genuinely yours. SSH access is gated by your key and lands only in your agent. The dashboard sits behind a private, per-instance URL over TLS. One key quietly unlocks all of your agents, each at its own address.
👥 Run more than one. Spin up several agents under a single account: a coding assistant here, a Slack concierge there, an experimental sandbox alongside. Each is isolated, each is billed on its own, and they're all reachable with the same key.
No surprise egress fees, no per-seat markup, no "enterprise, call us." Just the host, plus the inference you use.
Pick OpenClaw or Hermes, pick a size, and click create. A couple of minutes later you've got your own agent: isolated, inference-ready, reachable by dashboard or SSH, backed up automatically, and yours to keep, update, and pause whenever you like.
From $13/month. Idle is free. Inference runs on one of the most cost-effective AI platforms around.
👉 Get started at https://deepinfra.com/dash/agents
GLM-5.1 on DeepInfra: Z.AI’s Agentic Engineering Model<p>Z.AI’s GLM-5.1 scores 58.4 on SWE-Bench Pro — ahead of both Claude Opus 4.6 (57.3) and GPT-5.4 (57.7) on real-world software engineering tasks. It’s the direct successor to GLM-5, designed for agentic engineering: long-horizon coding tasks, terminal operations, and repository-level work. The core design premise is that previous models, including GLM-5, tend to plateau after […]</p>
Open vs Closed Source AI Models: Intelligence, Price & Speed Compared<p>The LLM landscape in 2026 looks nothing like it did two years ago. Back then the assumption was simple: if you wanted the best model, you paid OpenAI or Anthropic, and that was that. Open source models were a respectable second tier, good for experimentation, fine-tuning, and budget workloads, but not quite there for serious […]</p>
Gemma 4 Model Overview: Features, Architecture & Use Cases<p>Gemma 4 is Google DeepMind’s latest family of open-weight models, released on April 3, 2026 under the Apache 2.0 license. The family spans four model sizes — from edge-optimized variants for mobile devices to a 31B dense model for server-side deployments — with every model supporting multimodal input, built-in reasoning, and a context window of […]</p>
© 2026 DeepInfra. All rights reserved.