Batch API vs Real-Time: When the Discount Is Worth It
Compare native batch APIs, application queues and real-time inference. Calculate savings, reconcile unordered results and account for deadlines and retries.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

A native Batch API can reduce inference charges for work that can wait. Putting ordinary requests into your own background queue does not, by itself, activate a provider's batch price. First establish which endpoint processes the work, what deadline it offers and how that endpoint bills supported models.
Batch processing is a good candidate for offline classification, evaluation and document enrichment. It is usually a poor fit when a person is waiting for the next sentence or when the next step depends immediately on a tool result.
Distinguish three execution choices
| Approach | Where work waits | What to verify |
|---|---|---|
| Real-time request | In the active request path | Interactive latency, limits and normal pricing |
| Application queue | In your jobs infrastructure | Worker capacity, retries and the API price actually used |
| Native Batch API | In the provider's batch system | Eligible operations, completion window and batch pricing |
As checked on September 29, 2026, the official OpenAI Batch guide and Claude batch guide describe 50% pricing relative to their corresponding standard API rates. Their documented batch windows extend up to 24 hours. These are native-provider terms, not a claim that KeepRouter exposes either batch endpoint or applies the same discount.
Check current model eligibility, endpoint availability and the applicable price before uploading work. An API-compatible chat route is not evidence of compatibility with batch creation, file upload, result retrieval or cancellation operations.
Calculate savings at the workload level
Suppose an offline classifier processes 100,000 independent records each month. At a hypothetical standard cost of $0.002 per record, model charges are $200. A qualifying 50% batch rate would make the same token workload $100, saving $100 before implementation and operations.
Now allocate $60 per month for the additional queue/result-reconciliation work. Net savings are $40. At only 10,000 records, gross savings would be $10, so that same allocation would exceed the benefit. These numbers are an illustrative worksheet, not measured operating costs or a provider quote.
standard_cost_per_record = 0.002
batch_fraction = 0.5
monthly_operations = 60
saving_per_record = standard_cost_per_record * (1 - batch_fraction)
break_even_records = monthly_operations / saving_per_record
print(round(break_even_records)) # 60000 records per monthReplace the operations allocation with your team's situation. An existing reliable job system changes the calculation. Include one-time engineering work over a stated planning period rather than pretending implementation is free. Also check whether different request modes change caching, token usage or other charges; a headline discount is not necessarily the final invoice reduction.
Give every record a durable identity
Prepare a small fixture before submitting a large file. Use three invented support tickets with stable IDs: billing-001, delivery-002 and account-003. Each result must join back to its original ticket even if the result order differs.
OpenAI's batch documentation explicitly warns that output ordering may differ from input ordering. Build the reconciliation around the request identifier rather than line numbers. The following local fixture demonstrates the join; it is not a provider response schema:
inputs = {"billing-001": "Duplicate charge",
"delivery-002": "Package late",
"account-003": "Cannot sign in"}
results = [("account-003", "account"), ("billing-001", "billing")]
joined = {record_id: label for record_id, label in results}
missing = set(inputs) - set(joined)
assert missing == {"delivery-002"}Reject duplicate result IDs and unknown IDs before committing changes. Keep the original input hash, model, prompt version and batch identifier with the job. A successful batch-level status does not eliminate the need to inspect each record's status and validate its output.
Design the deadline and recovery path
Set a business deadline separately from the provider's processing window. If results are needed at 09:00, submitting at 08:55 to a service that may take hours is not a viable plan. Allow time to retrieve results, validate them and handle failures.
Retry only records that need another attempt. Resubmitting all 100,000 records because 200 failed wastes work and can duplicate downstream updates. An application record ID should survive resubmission, while attempt IDs remain distinct. Apply the same principle when a cancellation leaves a mixture of completed and unfinished items.
If an urgent subset must use real-time inference, record that escalation separately and prevent a late batch result from overwriting a newer accepted result. Decide which outcome wins by version or job state, not by whichever response arrives last.
Choose the smallest useful experiment
Start with a representative set containing valid, invalid and ambiguous records. Measure acceptance, completion time, retrieval failures and total charges across attempts. Use the structured-output guide if downstream code expects a fixed schema, and the 429 guide for rate-limited real-time recovery.
Use the API cost calculator to estimate the supported real-time model baseline. Apply a batch discount only when the actual destination documents it for your operation. If the work must finish interactively or batch operations are unavailable on your route, keep real-time execution and optimize prompt size, model choice and retry behavior instead.
Frequently asked questions
Does my own queue qualify for batch pricing?
Not automatically. The request must use an eligible batch operation on a destination that documents batch pricing. Queued normal API calls retain their applicable normal rates.
Can batch results be joined by line number?
Do not assume matching order. Use stable request identifiers, inspect per-record status and reject missing, duplicate or unknown results before applying updates.
Does this article announce KeepRouter Batch API support?
No. Native-provider batch terms are discussed for comparison. Use KeepRouter's published API reference to determine available operations on its routes.
Sources reviewed
Article last reviewed 2026-09-29