DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Flan-UL2 is probably the best open source model available right now for chatbots. In this post we will show you how to get started with it very easily. Flan-UL2 is large - 20B parameters. It is fine tuned version of the UL2 model using Flan dataset. Because this is quite a large model it is not easy to deploy it on your own machine. If you rent a GPU in AWS, it will cost you around $1.5 per hour or $1080 per month. Using DeepInfra model deployments you only pay for the inference time, and we do not charge for cold starts. Our pricing is $0.0005 per second of running inference on Nvidia A100. Which translates to about $0.0001 per token generated by Flan-UL2.
Also check out the model page https://deepinfra.com/google/flan-ul2. You can
run inferences, check the docs/API for running inferences via curl.
First, you'll need to get an API key from the DeepInfra dashboard.
You can deploy the google/flan-ul2 model easily through the web dashboard or API. The model will be automatically deployed when you first make an inference request.
You can use it with our REST API. Here's how to use it with curl:
curl -X POST \
-d '{"prompt": "Hello, how are you?"}' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer YOUR_API_KEY" \
'https://api.deepinfra.com/v1/inference/google/flan-ul2'
To see the full documentation of how to call this model, check out the model page on the DeepInfra website or the API documentation.
If you want a list of all the models you can use on DeepInfra, you can visit the models page on our website or use the API to get a list of available models.
There is no easier way to get started with arguably one of the best open source LLM. This was quite easy right? You did not have to deal with docker, transformers, pytorch, etc. If you have any question, just reach out to us on our Discord server.
GLM-5.1 API Benchmarks: Latency, Throughput & Cost<p>Z.ai’s GLM-5.1 is an April 2026 open-weight reasoning model built for long-horizon agentic engineering — and accessing it effectively means navigating a real spread of provider options. Across 10 benchmarked API providers, blended pricing ranges from $0.74 to $1.70 per 1M tokens, output speed from 33.8 to 175.2 t/s, and the fastest provider is 5.2x […]</p>
Gemma 4 on DeepInfra: Fast & Scalable Open AI Models<p>Google DeepMind’s Gemma 4 scored 88.3% on AIME 2026 mathematics benchmarks in its 26B MoE variant — compared to 20.8% for its predecessor, Gemma 3 27B. That’s not an incremental update. The family spans four model sizes designed for hardware targets as different as a Raspberry Pi and a consumer GPU workstation, with every model […]</p>
Kimi K3 vs Claude Opus 4.8 vs GPT-5.6 Sol: Practical AI Model Comparison<p>The three strongest models available on DeepInfra right now don’t separate cleanly by capability tier. Kimi K3 (available via DeepInfra), Claude Opus 4.8, and GPT-5.6 Sol all score within 3 points of each other on the Artificial Analysis Intelligence Index. All three support one-million-token context windows. All three handle vision. And yet the right choice […]</p>
© 2026 DeepInfra. All rights reserved.