How to reduce LLM API costs: a task-level cost worksheet
Find what drives your API bill, calculate cost per accepted task, and compare caching, shorter context, model choice and retry limits with worked examples.
Published 2026-08-15 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

To reduce an application's LLM API bill, first find which completed tasks consume the money. A lower token price can help, but repeated attempts, oversized context and unused output can outweigh it. Start with one workflow, calculate its cost per accepted result, then change one factor at a time. The worksheet below separates a cheaper request from a cheaper useful outcome.
Build the monthly estimate from a single task
For a simple text route without separate cache-write, tool or storage charges, use:
request_cost = (uncached_input × input_rate
+ cached_input × cache_read_rate
+ output × output_rate) / 1,000,000
monthly_inference = sum(cost of every attempt in the month)
cost_per_accepted_task = total_workflow_cost / accepted_tasksRates in this formula are per million tokens. Use the selected model's customer price and its actual billing unit; a video second or an image request cannot be inserted as a token count. Where cache writes, storage, tools or other services are billed separately, add them. Anthropic's token-counting documentation is useful for input estimation; a preflight estimate still does not tell you how much output or retry work the completed task will require.
Worked example: the cheaper request loses
These are hypothetical workloads and prices, not product benchmarks. Each option processes the same 10,000 tasks. A human-reviewed rubric determines which results are acceptable. All paid attempts are counted, including retries; the example excludes other workflow expenses to make the difference visible.
| Measure | Option A | Option B |
|---|---|---|
| Cost per attempt | $0.010 | $0.008 |
| Attempts per task, average | 1.0 | 1.4 |
| Monthly inference cost | $100 | $112 |
| Accepted results | 9,000 | 8,000 |
| Inference cost per accepted result | $0.0111 | $0.0140 |
Option B has a 20% lower request price yet a higher cost per accepted result. This does not mean a larger model always wins. It means the model decision needs task success and retry counts alongside rates. For a support workflow, the cost-per-resolution worksheet adds human escalation and reopened cases to the calculation.
Find the dominant cost before choosing a fix
| Cost pattern | First experiment | Guard against |
|---|---|---|
| Repeated long instructions or documents | Measure eligible cache reuse at normal request intervals | Assuming a hot-loop hit rate lasts all day |
| Large retrieved context | Reduce retrieved passages with a fixed relevance test | Removing evidence needed to answer correctly |
| Long generated answers | Ask for a smaller deliverable and cap output | Truncating JSON or tool arguments |
| Many rejected answers | Improve task instructions or test a stronger model | Counting API 200 responses as useful results |
| Repeated failures | Bound retries and classify permanent errors | Several SDK and gateway layers retrying the same job |
| A few tenants dominate spend | Apply a per-tenant application budget | Confusing requests per minute with a spending cap |
Measure the expensive pattern first. If output dominates, a cache experiment will have limited effect. If most spending comes from a few enormous jobs, a small average request size is a misleading budget input. Keep both median and upper-tail task costs, and investigate the most expensive completed jobs without collecting their private prompt text.
Put limits where a task can multiply work
Set an application-level maximum for tool rounds and total attempts, plus the route's supported output limit. Define what happens when the budget is reached: return a partial result, queue for review, or stop with an actionable error. A token limit alone does not cap a loop that sends another request after every tool result.
For example, an internal report assistant might allow one initial generation and one correction attempt. That is an illustrative policy, not a KeepRouter default. Keep the accumulated cost on the logical job, so a correction does not reset its budget. Do the same for retries after a timeout, where the earlier attempt may already have completed.
Turn the estimate into a paid-use decision
Use a fixed set of representative tasks, including a few long contexts and difficult cases. Record the model ID, endpoint, input/cache/output usage, attempts, accepted outcome, duration and charge. Compare one change against the current version. Keep the quality rubric and task set unchanged during that comparison, and inspect disagreements manually.
Open the API cost calculator to build low, expected and high usage scenarios. Then reconcile a small batch against your Usage records before increasing traffic. Calculator outputs are planning estimates; the paid result depends on actual quantities and supported operations. The cache cost guide covers the overlapping-token mistake, and the failover guide shows where attempt budgets can escape the application.
Frequently asked questions
Should we choose the lowest-priced model by default?
No. First require the model to pass the task's capability and quality contract, then compare measured total cost including retries and failures.
Are token estimates sufficient for billing reconciliation?
Use estimates for admission and planning, but reconcile with actual response usage and the billing ledger because tokenization and cache treatment can differ.
Sources reviewed
Article last reviewed 2026-09-29
- [1] Anthropic token counting
- [2] OpenAI Responses API reference
- [3] Cloudflare AI Gateway caching
- [4] KeepRouter model catalog