How to reduce LLM API costs: a task-level cost worksheet

Find what drives your API bill, calculate cost per accepted task, and compare caching, shorter context, model choice and retry limits with worked examples.

Published 2026-08-15 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

Multi-model API cost ledger showing tokens, cache, retries, fallbacks, and ownership
Cost control starts with attributable requests and reconciled units before model-policy changes.

To reduce an application's LLM API bill, first find which completed tasks consume the money. A lower token price can help, but repeated attempts, oversized context and unused output can outweigh it. Start with one workflow, calculate its cost per accepted result, then change one factor at a time. The worksheet below separates a cheaper request from a cheaper useful outcome.

Build the monthly estimate from a single task

For a simple text route without separate cache-write, tool or storage charges, use:

request_cost = (uncached_input × input_rate
              + cached_input × cache_read_rate
              + output × output_rate) / 1,000,000
monthly_inference = sum(cost of every attempt in the month)
cost_per_accepted_task = total_workflow_cost / accepted_tasks

Rates in this formula are per million tokens. Use the selected model's customer price and its actual billing unit; a video second or an image request cannot be inserted as a token count. Where cache writes, storage, tools or other services are billed separately, add them. Anthropic's token-counting documentation is useful for input estimation; a preflight estimate still does not tell you how much output or retry work the completed task will require.

Worked example: the cheaper request loses

These are hypothetical workloads and prices, not product benchmarks. Each option processes the same 10,000 tasks. A human-reviewed rubric determines which results are acceptable. All paid attempts are counted, including retries; the example excludes other workflow expenses to make the difference visible.

MeasureOption AOption B
Cost per attempt$0.010$0.008
Attempts per task, average1.01.4
Monthly inference cost$100$112
Accepted results9,0008,000
Inference cost per accepted result$0.0111$0.0140

Option B has a 20% lower request price yet a higher cost per accepted result. This does not mean a larger model always wins. It means the model decision needs task success and retry counts alongside rates. For a support workflow, the cost-per-resolution worksheet adds human escalation and reopened cases to the calculation.

Find the dominant cost before choosing a fix

Cost patternFirst experimentGuard against
Repeated long instructions or documentsMeasure eligible cache reuse at normal request intervalsAssuming a hot-loop hit rate lasts all day
Large retrieved contextReduce retrieved passages with a fixed relevance testRemoving evidence needed to answer correctly
Long generated answersAsk for a smaller deliverable and cap outputTruncating JSON or tool arguments
Many rejected answersImprove task instructions or test a stronger modelCounting API 200 responses as useful results
Repeated failuresBound retries and classify permanent errorsSeveral SDK and gateway layers retrying the same job
A few tenants dominate spendApply a per-tenant application budgetConfusing requests per minute with a spending cap

Measure the expensive pattern first. If output dominates, a cache experiment will have limited effect. If most spending comes from a few enormous jobs, a small average request size is a misleading budget input. Keep both median and upper-tail task costs, and investigate the most expensive completed jobs without collecting their private prompt text.

Put limits where a task can multiply work

Set an application-level maximum for tool rounds and total attempts, plus the route's supported output limit. Define what happens when the budget is reached: return a partial result, queue for review, or stop with an actionable error. A token limit alone does not cap a loop that sends another request after every tool result.

For example, an internal report assistant might allow one initial generation and one correction attempt. That is an illustrative policy, not a KeepRouter default. Keep the accumulated cost on the logical job, so a correction does not reset its budget. Do the same for retries after a timeout, where the earlier attempt may already have completed.

Turn the estimate into a paid-use decision

Use a fixed set of representative tasks, including a few long contexts and difficult cases. Record the model ID, endpoint, input/cache/output usage, attempts, accepted outcome, duration and charge. Compare one change against the current version. Keep the quality rubric and task set unchanged during that comparison, and inspect disagreements manually.

Open the API cost calculator to build low, expected and high usage scenarios. Then reconcile a small batch against your Usage records before increasing traffic. Calculator outputs are planning estimates; the paid result depends on actual quantities and supported operations. The cache cost guide covers the overlapping-token mistake, and the failover guide shows where attempt budgets can escape the application.

Frequently asked questions

Should we choose the lowest-priced model by default?

No. First require the model to pass the task's capability and quality contract, then compare measured total cost including retries and failures.

Are token estimates sufficient for billing reconciliation?

Use estimates for admission and planning, but reconcile with actual response usage and the billing ledger because tokenization and cache treatment can differ.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Anthropic token counting
  2. [2] OpenAI Responses API reference
  3. [3] Cloudflare AI Gateway caching
  4. [4] KeepRouter model catalog

Related guides

← All posts · Models & pricing · Get an API key