Pricing calculator

Calculate the cheapest AI route.

Describe the workload, set the few numbers that move the bill, then compare API, batch, subscription, hybrid routing, local-model, and GPU-hosted options.

Workload calculator

Cost drivers

Examples
Advanced assumptions

These only affect GPU-hosted, local, storage, egress, and self-host break-even math.

Recommended route

Hybrid router is the safest default for this workload at $222/month.

Start with Hybrid router: Route routine inference to open models and escalate only the expensive judgment calls.

$177/mo over the cheapest lane buys routing controls and a small frontier reserve.

Estimate basis

Uses public list prices and the assumptions you supplied. Treat marketplace GPU, cloud procurement, and custom enterprise plans as manual-pricing checks before production.

Recommended

$222

Hybrid router

Cheapest

$45.05

Open-model API

Yearly

$2,663

Hybrid router

Why this works

  • Hybrid router gives the best balance of cost and quality for this workload.
  • Routine inference moves to cheaper open-model or batch lanes instead of paying frontier rates on every request.
  • The modeled cache assumption is 86%, so repeated context should reduce input cost if the app preserves cacheable prompts.

Avoid first

  • Frontier API only: $748/month. Too expensive unless every output directly changes a high-value decision.
  • Self-hosted H100: $692/month. Use only after utilization tests prove the GPU stays busy.

Keep these assumptions when you share this pricing run.

Compare routes

Monthly cost by strategy

Lower is better only after the quality bar is met.

Open-model API

Cheapest lane

Cheapest managed lane for bulk extraction, summarization, triage, and routine agent loops.

$45.05/mo

Weekly $10.40Yearly $541

Managed cloud AI

Manual pricing

AWS, Google Vertex AI, and Azure AI Foundry path for enterprise controls, region choices, and cloud procurement.

$76.07/mo

Weekly $17.55Yearly $913

Batch API

Usable lane

Non-urgent processing lane for documents, logs, offline analysis, and backfills.

$98.70/mo

Weekly $22.78Yearly $1,184

Subscription reserve

Human reserve

Human-in-the-loop coding and review budget before buying heavy API volume.

$165/mo

Weekly $38.09Yearly $1,981

Hybrid router

Best balance

Default production strategy: cheap routine lane plus small frontier reserve.

$222/mo

Weekly $51.21Yearly $2,663

AI Gateway router

Usable lane

Vercel AI Gateway control plane for provider routing, fallbacks, and observability with provider pass-through pricing.

$222/mo

Weekly $51.21Yearly $2,663

Local inference and fine-tuning

Benchmark local models before you buy dedicated GPU capacity.

These gates keep tuning and self-hosting tied to eval quality, accepted outputs, and real utilization.

Local inference

Benchmark local inference before renting dedicated GPUs.

Start with an open-model API baseline so local inference has a real cost and quality target.

Benchmark before GPU rental

Fine-tuning

Tune only after routing and retrieval fail

Create a golden eval set from repeated accepted and rejected examples before changing weights.

Eval set required

Self-hosting

$1,436/mo H100 baseline

Move steady traffic to rented GPUs only when utilization, storage, egress, and operations beat API pricing.

Idle time decides the bill

Inputs and break-even checkssoftware development

Workload assumptions

Type

software development

Input

650.0M tokens/mo

Output

32.0M tokens/mo

API/tools

0/mo

Cache

86%

GPU hrs

120/mo

Storage

40 GB

Egress

12 GB

Calculated from the fields on the left. Adjust the workload and recalculate to compare another route.

Self-hosting breakpoint

$1,436/month H100 baseline

Self-hosting uses the lowest verified H100 price currently loaded (RunPod H100 serverless worker). It usually beats frontier APIs before it beats low-cost open-model APIs; the GPU has to stay busy enough to overcome idle time, storage, egress, and operations overhead.

Cost componentsRecommended and self-hosted detail

Recommended components

Hybrid router

$222

Routine open-model lane

tokens

$40.55

90% routine share routed away from frontier models.

Official prices: deepseek-v4-flash

Frontier escalation lane

tokens

$61.38

10% hard-case share kept for frontier reasoning.

Official prices: anthropic-claude-sonnet-5

ChatGPT Plus with Codex

subscription

$20.00

Monthly subscription reserve; not a production API substitute.

Official prices: openai-chatgpt-plus

Claude Max 5x

subscription

$100

Monthly subscription reserve; not a production API substitute.

Official prices: anthropic-claude-max-5x

Self-hosted cost check

Self-hosted H100

$692

RunPod H100 serverless worker

gpu

$239

120 GPU-hours/month assumption.

Official prices: runpod-h100-serverless-hour

Storage

storage

$2.80

40 GB-month assumption.

Official prices: runpod-network-storage-standard

Network egress

egress

$0

12 GB egress assumption.

Official prices: vast-ai-internet-egress

Operations overhead

operations

$450

Conservative monthly allowance for monitoring, retries, orchestration, and maintenance.