DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Agents live or die on retrieval. Miss the right passage, code block, or document, and the agent reasons from the wrong context — wasting tokens and reducing answer quality. And agents retrieve constantly: decomposing tasks, rewriting queries, searching memory, inspecting code. For retrieval to be the default for every agent decision, the embedding model has to be accurate, fast, and cost-effective, all at once.
That's exactly what NVIDIA Nemotron 3 Embed delivers — and it's now available on DeepInfra.
Enterprise retrieval usually forces tradeoffs: better accuracy means paying more per query, larger models are harder to serve, and fast answers can mean missing the right content. Nemotron 3 Embed comes in two sizes so you don't have to pick one corner of that triangle:
Whether you're optimizing for maximum retrieval quality or production-scale throughput, NVIDIA Nemotron 3 Embed provides the right model.
Multi-turn agents retrieve repeatedly — for planning, long-term memory, code understanding, tool use, and multi-step reasoning. Weak retrieval means more turns, more tokens, and more hallucination risk. Strong retrieval reduces irrelevant context and keeps agents grounded. Nemotron 3 Embed is a stronger retrieval layer for agentic retrieval, query decomposition, query rewriting, code retrieval, enterprise search, and RAG applications.
Nemotron 3 Embed ships with open weights, datasets, and recipes, so you can inspect it, tune it, and fine-tune it for your domain. No black box, no lock-in. On DeepInfra, both sizes are available now through our standard OpenAI-compatible API, so you can move from experimentation to production without changing infrastructure.
Generating embeddings requires only a single OpenAI-compatible API call:
from openai import OpenAI
client = OpenAI(
base_url="https://api.deepinfra.com/v1/openai",
api_key="$DEEPINFRA_TOKEN",
)
resp = client.embeddings.create(
model="nvidia/Nemotron-3-Embed-8B",
input=["How do I reset my password?"],
)
print(resp.data[0].embedding[:8])
For high-throughput workloads, simply swap in nvidia/Nemotron-3-Embed-1B-BF16, or nvidia/Nemotron-3-Embed-1B-NVFP4 for maximum throughput on NVIDIA Blackwell GPUs. All three models are available on DeepInfra today.
Have questions or need help? Reach out at feedback@deepinfra.com, join our Discord, or connect with us on X (@DeepInfra) — we're happy to help.
Nemotron 3 Nano Explained: NVIDIA’s Efficient Small LLM and Why It Matters<p>The open-source LLM space has exploded with models competing across size, efficiency, and reasoning capability. But while frontier models dominate headlines with enormous parameter counts, a different category has quietly become essential for real-world deployment: small yet high-performance models optimized for edge devices, private on-prem systems, and cost-sensitive applications. NVIDIA’s Nemotron family brings together open […]</p>
Build an OCR-Powered PDF Reader & Summarizer with DeepInfra (Kimi K2)<p>This guide walks you from zero to working: you’ll learn what OCR is (and why PDFs can be tricky), how to turn any PDF—including those with screenshots of tables—into text, and how to let an LLM do the heavy lifting to clean OCR noise, reconstruct tables, and summarize the document. We’ll use DeepInfra’s OpenAI-compatible API […]</p>
GLM-4.6 API: Get fast first tokens at the best $/M from Deepinfra's API - Deep Infra<p>GLM-4.6 is a high-capacity, “reasoning”-tuned model that shows up in coding copilots, long-context RAG, and multi-tool agent loops. With this class of workload, provider infrastructure determines perceived speed (first-token time), tail stability, and your unit economics. Using ArtificialAnalysis (AA) provider charts for GLM-4.6 (Reasoning), DeepInfra (FP8) pairs a sub-second Time-to-First-Token (TTFT) (0.51 s) with the […]</p>
© 2026 DeepInfra. All rights reserved.