Pricing calculator

Calculate the cheapest AI route.

Describe the workload, set the few numbers that move the bill, then compare API, batch, subscription, hybrid routing, local-model, and GPU-hosted options.

Workload calculator

Cost drivers

Examples
Advanced assumptions

These only affect GPU-hosted, local, storage, egress, and self-host break-even math.

Recommended route

Hybrid router is the safest default for this workload at $199/month.

Start with Hybrid router: Route routine inference to open models and escalate only the expensive judgment calls.

$98.74/mo over the cheapest lane buys routing controls and a small frontier reserve.

Estimate basis

Uses public list prices and the assumptions you supplied. Treat marketplace GPU, cloud procurement, and custom enterprise plans as manual-pricing checks before production.

Recommended

$199

Hybrid router

Cheapest

$99.95

Open-model API

Yearly

$2,384

Hybrid router

Why this works

  • Hybrid router gives the best balance of cost and quality for this workload.
  • Routine inference moves to cheaper open-model or batch lanes instead of paying frontier rates on every request.
  • The modeled cache assumption is 74%, so repeated context should reduce input cost if the app preserves cacheable prompts.

Avoid first

  • Frontier API only: $1,318/month. Too expensive unless every output directly changes a high-value decision.
  • Self-hosted H100: $932/month. Use only after utilization tests prove the GPU stays busy.

Keep these assumptions when you share this pricing run.

Compare routes

Monthly cost by strategy

Lower is better only after the quality bar is met.

Open-model API

Cheapest lane

Cheapest managed lane for bulk extraction, summarization, triage, and routine agent loops.

$99.95/mo

Weekly $23.07Yearly $1,199

Subscription reserve

Human reserve

Human-in-the-loop coding and review budget before buying heavy API volume.

$99.95/mo

Weekly $23.07Yearly $1,199

Managed cloud AI

Manual pricing

AWS, Google Vertex AI, and Azure AI Foundry path for enterprise controls, region choices, and cloud procurement.

$155/mo

Weekly $35.80Yearly $1,862

Hybrid router

Best balance

Default production strategy: cheap routine lane plus small frontier reserve.

$199/mo

Weekly $45.85Yearly $2,384

AI Gateway router

Usable lane

Vercel AI Gateway control plane for provider routing, fallbacks, and observability with provider pass-through pricing.

$199/mo

Weekly $45.85Yearly $2,384

Batch API

Usable lane

Non-urgent processing lane for documents, logs, offline analysis, and backfills.

$201/mo

Weekly $46.45Yearly $2,416

Local inference and fine-tuning

Benchmark local models before you buy dedicated GPU capacity.

These gates keep tuning and self-hosting tied to eval quality, accepted outputs, and real utilization.

Local inference

Benchmark local inference before fine-tuning.

Start with an open-model API baseline so local inference has a real cost and quality target.

Benchmark before GPU rental

Fine-tuning

$7,500 build cost, 8 mo break-even

Create a golden eval set from repeated accepted and rejected examples before changing weights.

$321/mo run cost

Self-hosting

$1,437/mo H100 baseline

Move steady traffic to rented GPUs only when utilization, storage, egress, and operations beat API pricing.

Idle time decides the bill

Inputs and break-even checkssupport

Workload assumptions

Type

support

Input

650.0M tokens/mo

Output

90.0M tokens/mo

API/tools

0/mo

Cache

74%

GPU hrs

240/mo

Storage

60 GB

Egress

30 GB

Calculated from the fields on the left. Adjust the workload and recalculate to compare another route.

Self-hosting breakpoint

$1,437/month H100 baseline

Self-hosting uses the lowest verified H100 price currently loaded (RunPod H100 serverless worker). It usually beats frontier APIs before it beats low-cost open-model APIs; the GPU has to stay busy enough to overcome idle time, storage, egress, and operations overhead.

Cost componentsRecommended and self-hosted detail

Recommended components

Hybrid router

$199

Routine open-model lane

tokens

$91.95

92% routine share routed away from frontier models.

Official prices: deepseek-v4-flash

Frontier escalation lane

tokens

$107

8% hard-case share kept for frontier reasoning.

Official prices: anthropic-claude-sonnet-5

Self-hosted cost check

Self-hosted H100

$932

RunPod H100 serverless worker

gpu

$478

240 GPU-hours/month assumption.

Official prices: runpod-h100-serverless-hour

Storage

storage

$4.20

60 GB-month assumption.

Official prices: runpod-network-storage-standard

Network egress

egress

$0

30 GB egress assumption.

Official prices: vast-ai-internet-egress

Operations overhead

operations

$450

Conservative monthly allowance for monitoring, retries, orchestration, and maintenance.