DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

AI agents have grown up. People are running them as real assistants now: reading email, triaging Slack, kicking off cron jobs, calling tools, and remembering context for weeks at a time. The hard part isn't building one anymore. It's the jump from "I made something cool on my laptop" to "I have an agent that actually runs around the clock."
That jump is all infrastructure, and it's tedious. You rent a VM. You lock down SSH. You install a runtime, wire up model API keys, and set up TLS so the dashboard isn't wide open to the internet. Then you keep it patched, figure out how to update the framework without losing your data, and pay for the box even on the weekends when nobody's using it.
Deep Infra Hosted Agents takes all of that off your plate. One click gives you a dedicated, isolated agent that's pre-wired to fast inference and ready to work the moment it boots. Starts at $13/month.
We host two flavors, both built on the same proven runtime:
OpenClaw - the dashboard agent. OpenClaw comes with a full web dashboard. Chat with your agent, manage its skills, plug in MCP servers, schedule cron tasks, and peek at its memory and workspace, all from the browser. It hooks into the channels you already use: Slack, Telegram, WhatsApp, Signal, and more. If you want an assistant the whole team can talk to, this is the one.
Hermes - the self improving agent. Hermes is the lean, SSH-first sibling. Same engine, no dashboard. You ssh straight in with your own key and drive it from the command line. It's made for pros who'd rather live in the terminal: scriptable, quiet, no UI in the way.
Each agent type comes in two host tiers. There's a smaller one to keep costs low for light, always-listening assistants, and a larger one with more RAM for heavier workloads, bigger workspaces, and more tools running at once. Pick what fits and resize as you grow. The entry tier starts at $13/month.
🔌 Inference-ready from the first second. This is the part most setups get wrong. Every hosted agent boots already wired to Deep Infra's model APIs, the same platform serving inference at scale for thousands of customers. Nothing to paste, no endpoint to configure, no "why isn't my agent responding" rabbit hole. It can think the moment it starts, on top of fast, affordable, frontier-class models.
🖱️ One-click setup. No provisioning scripts, no Docker, no certificates. You click, and a fully configured agent comes up for you: runtime, dashboard, SSH, TLS, secrets, all of it.
🔄 One-click updates. Agent frameworks move fast. When a new version drops, you update with a single click, and your memory, workspace, conversations, and configuration come right along with it. The framework swaps underneath while your data stays put. No migrations, no fingers crossed.
💾 Automatic backups and point-in-time restore. Your agent's entire state is snapshotted automatically every day, and you can take a backup on demand any time you're about to try something risky. If something breaks or you just want to rewind, restore to any saved checkpoint in a click. It's an undo button for your whole agent. We keep a rolling history of recent snapshots and copy them across regions, so your data isn't riding on a single machine.
💤 Stopped instances cost nothing. Pause an agent and you pay $0 for compute while it's idle, but its disk, memory, and full state stick around. Start it again and it's back in seconds, right where you left off. Run it hard during the week, stop it for the weekend, and only pay for the time you actually used. Plenty of always-on services bill you flat whether the thing is working or asleep. This one doesn't.
🧠 Long-term memory that persists. Your agent doesn't get wiped between sessions. Files, conversations, learned context, and workspace state live on a durable disk that survives stops, starts, updates, and restores. The agent you talk to next month is the same one that remembers what you told it today.
🔗 Connect everything. Agents plug into the channels and tools you already use (Slack, Telegram, WhatsApp, Signal, and more) and extend through skills, MCP servers, and scheduled cron jobs. It can listen, act, and run on a schedule without you in the loop.
🔐 Truly isolated and private. Each agent runs in its own environment with its own kernel, not a shared container, so its world is genuinely yours. SSH access is gated by your key and lands only in your agent. The dashboard sits behind a private, per-instance URL over TLS. One key quietly unlocks all of your agents, each at its own address.
👥 Run more than one. Spin up several agents under a single account: a coding assistant here, a Slack concierge there, an experimental sandbox alongside. Each is isolated, each is billed on its own, and they're all reachable with the same key.
No surprise egress fees, no per-seat markup, no "enterprise, call us." Just the host, plus the inference you use.
Pick OpenClaw or Hermes, pick a size, and click create. A couple of minutes later you've got your own agent: isolated, inference-ready, reachable by dashboard or SSH, backed up automatically, and yours to keep, update, and pause whenever you like.
From $13/month. Idle is free. Inference runs on one of the most cost-effective AI platforms around.
👉 Get started at https://deepinfra.com/dash/agents
Kimi K2 0905 API from Deepinfra: Practical Speed, Predictable Costs, Built for Devs - Deep Infra<p>Kimi K2 0905 is Moonshot’s long-context Mixture-of-Experts update designed for agentic and coding workflows. With a context window up to ~256K tokens, it can ingest large codebases, multi-file documents, or long conversations and still deliver structured, high-quality outputs. But real-world performance isn’t defined by the model alone—it’s determined by the inference provider that serves it: […]</p>
Best SaaS Platforms for Deploying Gemma 4 in 2026<p>Gemma 4 is available across a range of platforms — from fully managed API providers to local runners and no-code builders. The right choice depends on what you’re optimizing for: cost, latency, data privacy, local execution, or zero infrastructure overhead. This guide breaks down the top options by use case so you can match the […]</p>
Kimi K2 0905 API Benchmarks: Latency, Throughput & Cost<p>About Kimi K2 0905 Kimi K2 0905 is a state-of-the-art large language model developed by Moonshot AI, representing a significant advancement in open-weight AI capabilities. This Mixture-of-Experts (MoE) model features 1 trillion total parameters with 32 billion activated parameters per forward pass, making it highly efficient while maintaining frontier-level performance. The model supports a 256k […]</p>
© 2026 DeepInfra. All rights reserved.