Skip to content

    Predictable pricing.Transparent per model

    Every serverless model priced at the credible market rate and repriced automatically as the market moves.

    Serverless: Token Factory

    Pay for what you use, one OpenAI-compatible key, every model in the catalog.

    Pricing an agent workload means more than the sticker rate: average tokens per run, tool calls, fallback rate, and whether steady traffic should move to dedicated endpoints.

    ModelInput $/MOutput $/MContext
    Gemma 4 26B A4B$0.060$0.300256K
    Mistral Nemo$0.020$0.03032K
    Gemma 4 31B IT$0.120$0.350256K
    Qwen3.6-35B-A3B$0.100$0.150256K
    Nemotron 3 Super 120B A12B$0.085$0.400256K
    GPT-OSS 20B$0.030$0.130128K
    Qwen3 Coder 30B A3B$0.070$0.260256K
    Llama 3.1 8B Instruct$0.020$0.030128K
    Qwen3.5 9B$0.100$0.150250K
    Step 3.7 Flash$0.200$1.150128K
    GLM 4.5 Air$0.125$0.850128K
    DeepSeek V4 Flash$0.090$0.180768K
    GPT-OSS 120B$0.039$0.100128K
    Llama 3.3 70B Instruct$0.100$0.32064K
    MiniMax M2.7$0.180$0.720200K
    Qwen3 8B$0.020$0.10040K
    GLM 5.2$0.909$2.8561.0M
    Qwen3 30B A3B Thinking 2507$0.080$0.280256K
    ModelInput $/MOutput $/MContext
    Qwen3.6-35B-A3B$0.100$0.150256K
    Qwen3 Coder 30B A3B$0.070$0.260256K
    ModelInput $/MOutput $/MContext
    PaddleOCR-VL 1.5$0.140$0.800128K
    ModelInput $/MOutput $/MContext
    Gemma 4 26B A4B$0.060$0.300256K
    Gemma 4 31B IT$0.120$0.350256K
    Step 3.7 Flash$0.200$1.150128K
    ModelInput $/MContext
    BGE-M3$0.0108K
    ModelPer imageContext
    FLUX.1 [schnell]$0.0005
    ModelRateContext
    Whisper Large V3 Turbo$0.00067 / min
    NVIDIA Parakeet TDT 0.6B v3$0.0015 / min
    ModelRateContext
    Kokoro-82M$0.62 / M chars
    Voxtral 4B TTSCustom

    Voxtral 4B TTS is coming soon. Talk to us for early access.

    Find your serverless to dedicated break-even point

    Get an API key

    $10/month in free credits for your first 3 months.

    Dedicated GPU pricing

    Your own endpoints, fine-tuned models, or whole clusters, priced per GPU-hour and metered per second, on FlexAI or your own infrastructure. On-demand and reserved options; managed fine-tuning available.

    On FlexAI

    Compute hosted on FlexAI's NVIDIA and AMD fleets, with up to 99.9% SLA by tier.

    • State-of-the-art NVIDIA GPUs with NVLink and InfiniBand
    • On-demand and reserved pricing
    • Per-second metering
    • Region selection
    • Managed fine-tuning: per-M training tokens + storage $/GB-mo
    • One-click setup: get started in minutes
    Per GPU-hour, on-demand*
    B200$6.25/hr
    H200$3.15/hr
    H100$2.10/hr
    A100$1.80/hr
    L40S$1.50/hr

    *Contact sales for Essential and Custom package rates.

    Contact Us

    Custom: AI Factory & Enterprise

    Private cloud on your hardware, with the same managed AI stack.

    • Custom pricing on your workload. Talk to us.
    • VPC, on-prem, and air-gapped options
    • SLA 99.9% with geo redundancy
    • Dedicated CSM
    Proven on FlexAI
    • 75% lower compute cost
    • €22.5K total training cost

    LegML fine-tuned a 32B legal LLM on FlexAI H100.

    Cloud Savings Calculator

    Estimate your savings against AWS Bedrock, Azure OpenAI, GCP Vertex, Together, Fireworks, OpenAI, or Anthropic, at FlexAI's published rates.

    TierPriceIncludesSLA
    Starter$10/month in free credits for your first 3 months, card required at signup, then pay-as-you-go at our published per-token rates2 workspace seats. OpenAI-compatible API. Dedicated endpoints on demand. Playground access.99%
    Essential$100 matched ($100 deposit → $200 credit). Contact us today; self-serve soon.8 seats. Concurrency + multi-fractional. Architecture call at $50 spend. SE access. HIPAA, DORA.99.5%
    CustomCustom pricing on your workload. Talk to us.VPC, on-prem, air-gapped. Dedicated CSM. Geo redundancy.99.9%

    Building for production? See the FlexAI Startup Program. Apply-only, three stages, scaling on one account as your workload grows.

    Models with royalty obligations or pricing floors (currently MiniMax M2.7 and FLUX.2) are excluded from Startup Program introductory discounts.

    How our pricing works

    How the rate is set

    For each model we track the cheapest credible market rate (pricepertoken.com, Artificial Analysis, and provider pages) and price at that rate. When a source moves, a scheduled refresh reprices automatically.

    Why this matters

    Most token APIs reprice reactively when they lose share. A rate pinned to the public market rate is predictable: you can budget against it.

    Frequently Asked Questions

    Coming from OpenAI or Claude? Read the migration guide

    Comparing providers? See how FlexAI compares with OpenAI, Fireworks, AWS Bedrock, and CoreWeave.