Predictable pricing.Transparent per model
Every serverless model priced at the credible market rate and repriced automatically as the market moves.
Every serverless model priced at the credible market rate and repriced automatically as the market moves.
Pay for what you use, one OpenAI-compatible key, every model in the catalog.
Pricing an agent workload means more than the sticker rate: average tokens per run, tool calls, fallback rate, and whether steady traffic should move to dedicated endpoints.
| Model | Input $/M | Output $/M | Context |
|---|---|---|---|
| Gemma 4 26B A4B | $0.060 | $0.300 | 256K |
| Mistral Nemo | $0.020 | $0.030 | 32K |
| Gemma 4 31B IT | $0.120 | $0.350 | 256K |
| Qwen3.6-35B-A3B | $0.100 | $0.150 | 256K |
| Nemotron 3 Super 120B A12B | $0.085 | $0.400 | 256K |
| GPT-OSS 20B | $0.030 | $0.130 | 128K |
| Qwen3 Coder 30B A3B | $0.070 | $0.260 | 256K |
| Llama 3.1 8B Instruct | $0.020 | $0.030 | 128K |
| Qwen3.5 9B | $0.100 | $0.150 | 250K |
| Step 3.7 Flash | $0.200 | $1.150 | 128K |
| GLM 4.5 Air | $0.125 | $0.850 | 128K |
| DeepSeek V4 Flash | $0.090 | $0.180 | 768K |
| GPT-OSS 120B | $0.039 | $0.100 | 128K |
| Llama 3.3 70B Instruct | $0.100 | $0.320 | 64K |
| MiniMax M2.7 | $0.180 | $0.720 | 200K |
| Qwen3 8B | $0.020 | $0.100 | 40K |
| GLM 5.2 | $0.909 | $2.856 | 1.0M |
| Qwen3 30B A3B Thinking 2507 | $0.080 | $0.280 | 256K |
| Model | Input $/M | Output $/M | Context |
|---|---|---|---|
| Qwen3.6-35B-A3B | $0.100 | $0.150 | 256K |
| Qwen3 Coder 30B A3B | $0.070 | $0.260 | 256K |
| Model | Input $/M | Output $/M | Context |
|---|---|---|---|
| PaddleOCR-VL 1.5 | $0.140 | $0.800 | 128K |
| Model | Input $/M | Output $/M | Context |
|---|---|---|---|
| Gemma 4 26B A4B | $0.060 | $0.300 | 256K |
| Gemma 4 31B IT | $0.120 | $0.350 | 256K |
| Step 3.7 Flash | $0.200 | $1.150 | 128K |
| Model | Input $/M | Context |
|---|---|---|
| BGE-M3 | $0.010 | 8K |
| Model | Per image | Context |
|---|---|---|
| FLUX.1 [schnell] | $0.0005 | — |
| Model | Rate | Context |
|---|---|---|
| Whisper Large V3 Turbo | $0.00067 / min | — |
| NVIDIA Parakeet TDT 0.6B v3 | $0.0015 / min | — |
| Model | Rate | Context |
|---|---|---|
| Kokoro-82M | $0.62 / M chars | — |
| Voxtral 4B TTS | Custom | — |
Voxtral 4B TTS is coming soon. Talk to us for early access.
Find your serverless to dedicated break-even point
$10/month in free credits for your first 3 months.
Your own endpoints, fine-tuned models, or whole clusters, priced per GPU-hour and metered per second, on FlexAI or your own infrastructure. On-demand and reserved options; managed fine-tuning available.
Compute hosted on FlexAI's NVIDIA and AMD fleets, with up to 99.9% SLA by tier.
*Contact sales for Essential and Custom package rates.
Private cloud on your hardware, with the same managed AI stack.
LegML fine-tuned a 32B legal LLM on FlexAI H100.
Estimate your savings against AWS Bedrock, Azure OpenAI, GCP Vertex, Together, Fireworks, OpenAI, or Anthropic, at FlexAI's published rates.
Leave your email and we'll send a model-by-model savings breakdown.
| Tier | Price | Includes | SLA |
|---|---|---|---|
| Starter | $10/month in free credits for your first 3 months, card required at signup, then pay-as-you-go at our published per-token rates | 2 workspace seats. OpenAI-compatible API. Dedicated endpoints on demand. Playground access. | 99% |
| Essential | $100 matched ($100 deposit → $200 credit). Contact us today; self-serve soon. | 8 seats. Concurrency + multi-fractional. Architecture call at $50 spend. SE access. HIPAA, DORA. | 99.5% |
| Custom | Custom pricing on your workload. Talk to us. | VPC, on-prem, air-gapped. Dedicated CSM. Geo redundancy. | 99.9% |
Building for production? See the FlexAI Startup Program. Apply-only, three stages, scaling on one account as your workload grows.
Models with royalty obligations or pricing floors (currently MiniMax M2.7 and FLUX.2) are excluded from Startup Program introductory discounts.
For each model we track the cheapest credible market rate (pricepertoken.com, Artificial Analysis, and provider pages) and price at that rate. When a source moves, a scheduled refresh reprices automatically.
Most token APIs reprice reactively when they lose share. A rate pinned to the public market rate is predictable: you can budget against it.
Coming from OpenAI or Claude? Read the migration guide
Comparing providers? See how FlexAI compares with OpenAI, Fireworks, AWS Bedrock, and CoreWeave.