DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Note: The list of supported base models is listed on the same page. If you need a base model that is not listed, please contact us at feedback@deepinfra.com
Rate limit will apply on combined traffic of all LoRA adapter models with the same base model. For example, if you have 2 LoRA adapter models with the same base model, and have rate limit of 200. Those 2 LoRA adapter models combined will have rate limit of 200.
Pricing is 50% higher than base model.
LoRA adapter model speed is lower than base model, because there is additional compute and memory overhead to apply the LoRA adapter. From our benchmarks, the LoRA adapter model speed is about 50-60% slower than base model.
You could merge the LoRA adapter with the base model to reduce the overhead. And use custom deployment, the speed will be close to the base model.
Getting StartedGetting an API Key
To use DeepInfra's services, you'll need an API key. You can get one by signing up on our platform.
Sign up or log in to your DeepInfra account at deepinfra.com
Navigate to the Dashboard and select API Keys
Create a new ...
GLM-5.3-Flash API Is Now on DeepInfra<p>Z.ai’s GLM-5.3-Flash activates just 18 billion of its 320 billion parameters at inference time, and on Z.ai’s reported results it outscores Claude Opus 4.8 on DeepSWE v1.1 (63.4 vs. 58.0), a demanding software engineering benchmark. It’s the first natively multimodal model in the GLM-5 series, combining text, image, and video input with a one-million-token context […]</p>
GLM-5.3-Flash API Providers: Speed & Cost<p>API Review Summary Metric Value Intelligence (Artificial Analysis Intelligence Index) 42 — well above the open-weight median (18) Speed 55.9 output tokens/sec — slower than the median (85.7 t/s) Latency (TTFT) 3.14s — higher than the median (2.05s) Cost (Z.ai first-party API) $0.15 / 1M input, $0.50 / 1M output; cache discount ~83% Cost efficiency […]</p>
© 2026 DeepInfra. All rights reserved.