Workload calculator
LLM API Cost Calculator
Estimate monthly API spend from text-token workloads or native voice and realtime billing meters. Choose the workload type, then enter the usage you actually know.
Workload type
Estimate text workloads from input/output tokens, request volume, cache reuse, and processing-tier pricing.
Usage input
Estimate a repeated workload from tokens per request, request volume, cache reuse, and processing-tier pricing.
Choose a preset or scroll for custom input.
Explore benchmark scores to see which models perform best on specific tasks.
Input
Reasoning models can bill hidden reasoning/thinking tokens as output. For planning, include them in the output estimate; for actual usage, enter the provider-reported billed output total.
Estimate input tokens from text
Text is processed only in your browser and is not saved or sent anywhere. This is a reference count for o200k_base, not the exact billable token count for the selected model or provider. It also excludes API message, tool, schema, image, and other request overhead.
Advanced options
Voice / Realtime estimate
Estimate voice and realtime routes from the native billing unit stored in the catalog. Unlike billing units are never converted or added together.
Selected-meter estimate
This estimate calculates only the selected native meter. Text tokens, cached audio, tools, backend models, transcription, session context, and other provider-specific charges are not added automatically.
View model detail and data sourcesCost estimates use the generated model database last built on Sep 22, 2026. Pricing, lifecycle, and capability fields can be incomplete or provider-specific, so verify production decisions with the official provider.
How to use this page
Start with a preset for a forecast, then change tokens and request volume to match your product. If you already have token totals from a completed request or session, choose Use actual token usage and enter that breakdown directly.
How pricing is calculated
Costs are calculated as (tokens ÷ 1,000,000) × price per 1M tokens. For planned workloads, the cache hit rate splits input into uncached and cached portions. Only cached input uses the cached-input rate; output keeps the selected output rate. The formula is total = ((uncachedInputTokens × inputPrice) + (cachedInputTokens × cachedInputPrice) + (outputTokens × outputPrice)) × requests. If the selected pricing tier has no separate cached-input rate, cached input uses the input rate.
Read the calculator examples guide for chatbot, RAG, summarization, and coding-agent inputs before changing the fields.
Check when a subscription is enough and when usage-based API pricing matters for product work.
Compare OpenAI, Anthropic, Google, Mistral, and DeepSeek across cheapest chat, mid-range, and reasoning model pricing.
See MMLU, GPQA, HumanEval, and other benchmark scores across providers to understand model quality beyond pricing.
Put 2-3 models next to each other to compare pricing, context windows, modalities, and capabilities.