rate card
Models & pricing
The specialist models we've benchmarked, hosted and priced — with the long tail we're onboarding next below. Prices are in each model's native unit; realtime is the on-demand rate, batch is a discounted flexible tier (send X-Tier: batch).
3 matches in text generation · clear
| model | task | tier | realtime | batch |
|---|---|---|---|---|
| meta-llama/Llama-3.1-8B-Instruct | text generation | $0.08333/1M tok | $0.04167/1M tok | |
| Qwen/Qwen3-4B-Instruct-2507 | text generation | $0.05/1M tok | $0.025/1M tok | |
| deepreinforce-ai/Ornith-1.0-35B | text generation | $1.66667/1M tok | $0.83333/1M tok |