DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Starting from langchain v0.0.322 you can make efficient async generation and streaming tokens with deepinfra.
The deepinfra wrapper now supports native async calls, so you can expect more performance (no more threads per invocation) from your async pipelines.
from langchain.llms.deepinfra import DeepInfra
async def async_predict():
llm = DeepInfra(model_id="meta-llama/Llama-2-7b-chat-hf")
output = await llm.apredict("What is 2 + 2?")
print(output)
Streaming lets you receive each token of the response as it gets generated. This is indispensable in user-facing applications.
def streaming():
llm = DeepInfra(model_id="meta-llama/Llama-2-7b-chat-hf")
for chunk in llm.stream("[INST] Hello [/INST] "):
print(chunk, end='', flush=True)
print()
You can also use the asynchronous streaming API, natively implemented underneath.
async def async_streaming():
llm = DeepInfra(model_id="meta-llama/Llama-2-7b-chat-hf")
async for chunk in llm.astream("[INST] Hello [/INST] "):
print(chunk, end='', flush=True)
print()
Best SaaS Tools and API Providers for MiMo-V2.5<p>As LLM architectures grow increasingly complex, the introduction of the MiMo-V2.5 series represents a significant step forward in multimodal capabilities and massive context handling. Integrating a model with a 1M-token context window and native multimodal support (image, video, audio, text) introduces substantial infrastructure considerations. For developers and enterprise architects, the priorities are clear: managing inference […]</p>
Best DeepSeek-V4.1-Flash API Providers in 2026<p>As LLM architectures grow increasingly sophisticated, deploying state-of-the-art models like DeepSeek-V4.1-Flash requires more than just a basic API wrapper. For engineering teams, the challenge lies in balancing time-to-first-token (TTFT), throughput, context caching, and enterprise-grade compliance. DeepSeek-V4.1-Flash offers remarkable capabilities—including a massive 1M+ token context window and Engram conditional memory—but unlocking its full potential depends heavily […]</p>
Kimi K3 vs Claude Opus 4.8 vs GPT-5.6 Sol: Practical AI Model Comparison<p>What’s changed since this comparison first ran Two of the three models compared here have since been succeeded, and GPT-5.6 Sol’s pricing has changed materially. Three of the strongest models available on DeepInfra right now don’t separate cleanly by capability tier. Kimi K3 (available via DeepInfra), Claude Opus 4.8, and GPT-5.6 Sol all score within […]</p>
© 2026 DeepInfra. All rights reserved.