DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

AI agents are moving beyond model calls to use tools, APIs, and execution environments. As that stack grows, the runtime becomes an important part of the architecture.
At DeepInfra, we work across that stack: inference, hosted agents, and sandboxes.
NVIDIA describes the distinction clearly: the harness guides what an agent tries; the infrastructure determines what it can do.
That maps directly onto how we think about agent infrastructure: model API, agent harness, and sandbox.
DeepInfra Sandboxes run untrusted agent code in hardware-isolated microVMs using Kata Containers, with each sandbox running its own kernel.
Network policy is enforced on the host, outside the VM. A sandboxed agent can reach the public internet while DeepInfra internal infrastructure and cloud metadata remain unreachable. Because these controls sit outside the VM, a confused or compromised agent cannot simply change its own boundaries.
Our Hosted Agents service applies the same principle to always-on agents such as OpenClaw and Hermes. Each agent runs in its own environment with its own kernel, SSH is gated by the user's key, and the dashboard is behind a private per-instance TLS URL.
This is why we see runtime isolation as a necessary complement to model-level safety.
DeepInfra is a first-party inference provider in NVIDIA OpenShell. We contributed to the DeepInfra provider profile and are listed in OpenShell's provider documentation.
OpenShell applies the same security principle one layer higher. Its gateway keeps credentials and policy outside the agent environment. For DeepInfra inference, the API key is bound to the approved DeepInfra endpoint, with network and L7 policy controls applied outside the agent.
This is exactly the kind of separation NVIDIA describes in the Open Agent Safety Platform: individual agent sandboxes, policy enforcement, and a gateway governing access to networks, credentials, tools, and model endpoints.
Our collaboration with NVIDIA on agent infrastructure also includes NemoClaw, where DeepInfra is one of the inference providers available through build.nvidia.com account linking, live since its GTC launch in March 2026.
We have participated in the OpenShell community through community calls, RFC discussions, and discussions on PRs and issues, and continue to stay engaged as the platform evolves. We're also evaluating where OpenShell could fit into DeepInfra's agent and sandbox infrastructure.
Runtime isolation has a cost, and that changes the compute profile of agent infrastructure.
In our published testing, NVIDIA Vera CPU was the fastest architecture tested in every workload category. At our QoS bar of p99 ≤ 6 seconds with zero errors, Vera supported 256 always-busy agents, each in its own microVM, on 20 physical cores, compared with 160–192 for the Intel incumbents.
One of the more important findings was that the isolation layer, not the CPU, accounts for the largest overhead in agent hosting.
For untrusted agents, that isolation is not something we want to remove. The CPU has to perform well with the isolation overhead included.
In our testing, Vera held up best. That's why we plan around Vera.
Our tested OpenShell integration guide for DeepInfra is available today:
DeepInfra + NVIDIA OpenShell integration guide
The guide shows how to run an agent in an OpenShell sandbox with DeepInfra as its inference provider. It covers a pinned OpenShell version, the maintained DeepInfra provider profile, configuring inference with an existing DeepInfra API key, defining a production variant sandbox policy that allows access to api.deepinfra.com while blocking other destinations, and also shows a full Hermes Agent walkthrough inside the sandbox.
Customers run OpenShell on their own hosts and configure their model and sandbox policies. DeepInfra provides the maintained provider profile, OpenAI-compatible chat and embeddings, our model catalog, and API keys.
As agent workloads grow, we'll continue evaluating and contributing to approaches that put security boundaries outside the agent itself.
For agents, guardrails matter. But the infrastructure still has to enforce the boundary.
DeepSeek V4 Pro: Model Overview, Features & Performance Guide<p>DeepSeek V4 Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model from DeepSeek, released on April 24, 2026 under the MIT license. It is designed for advanced reasoning, complex software engineering, and long-running agentic tasks, and arrives alongside DeepSeek-V4-Flash, a lighter 284B-parameter variant built for faster, lower-cost inference. The V4 series is DeepSeek’s first two-tier lineup […]</p>
Deploy Custom LLMs on DeepInfraDid you just finetune your favorite model and are wondering where to run it?
Well, we have you covered. Simple API and predictable pricing.
Put your model on huggingface
Use a private repo, if you wish, we don't mind. Create a hf access token just
for the repo for better security.
Create c...
Chat with books using DeepInfra and LlamaIndexAs DeepInfra, we are excited to announce our integration with LlamaIndex.
LlamaIndex is a powerful library that allows you to index and search documents
using various language models and embeddings. In this blog post, we will show
you how to chat with books using DeepInfra and LlamaIndex.
We will ...© 2026 DeepInfra. All rights reserved.