# Batch API vs Real-Time: When the Discount Is Worth It

> Compare native batch APIs, application queues and real-time inference. Calculate savings, reconcile unordered results and account for deadlines and retries.

_Published 2026-09-29 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 4 minute read_

![Multi-model API cost ledger showing tokens, cache, retries, fallbacks, and ownership](https://keeprouter.com/editorial/blog/control-multi-model-api-costs.png)

_Compare the full cost of accepted tasks with explicit workload assumptions. Illustration, not a vendor quote._

A native Batch API can reduce inference charges for work that can wait. Putting ordinary requests into your own background queue does not, by itself, activate a provider's batch price. First establish which endpoint processes the work, what deadline it offers and how that endpoint bills supported models.

Batch processing is a good candidate for offline classification, evaluation and document enrichment. It is usually a poor fit when a person is waiting for the next sentence or when the next step depends immediately on a tool result.

## Distinguish three execution choices

| Approach | Where work waits | What to verify |
| --- | --- | --- |
| Real-time request | In the active request path | Interactive latency, limits and normal pricing |
| Application queue | In your jobs infrastructure | Worker capacity, retries and the API price actually used |
| Native Batch API | In the provider's batch system | Eligible operations, completion window and batch pricing |

As checked on September 29, 2026, the official [OpenAI Batch guide](https://developers.openai.com/api/docs/guides/batch) and [Claude batch guide](https://platform.claude.com/docs/en/build-with-claude/batch-processing) describe 50% pricing relative to their corresponding standard API rates. Their documented batch windows extend up to 24 hours. These are native-provider terms, not a claim that KeepRouter exposes either batch endpoint or applies the same discount.

Check current model eligibility, endpoint availability and the applicable price before uploading work. An API-compatible chat route is not evidence of compatibility with batch creation, file upload, result retrieval or cancellation operations.

## Calculate savings at the workload level

Suppose an offline classifier processes 100,000 independent records each month. At a hypothetical standard cost of $0.002 per record, model charges are $200. A qualifying 50% batch rate would make the same token workload $100, saving $100 before implementation and operations.

Now allocate $60 per month for the additional queue/result-reconciliation work. Net savings are $40. At only 10,000 records, gross savings would be $10, so that same allocation would exceed the benefit. These numbers are an illustrative worksheet, not measured operating costs or a provider quote.

```python
standard_cost_per_record = 0.002
batch_fraction = 0.5
monthly_operations = 60
saving_per_record = standard_cost_per_record * (1 - batch_fraction)
break_even_records = monthly_operations / saving_per_record
print(round(break_even_records))  # 60000 records per month
```

Replace the operations allocation with your team's situation. An existing reliable job system changes the calculation. Include one-time engineering work over a stated planning period rather than pretending implementation is free. Also check whether different request modes change caching, token usage or other charges; a headline discount is not necessarily the final invoice reduction.

## Give every record a durable identity

Prepare a small fixture before submitting a large file. Use three invented support tickets with stable IDs: billing-001, delivery-002 and account-003. Each result must join back to its original ticket even if the result order differs.

OpenAI's batch documentation explicitly warns that output ordering may differ from input ordering. Build the reconciliation around the request identifier rather than line numbers. The following local fixture demonstrates the join; it is not a provider response schema:

```python
inputs = {"billing-001": "Duplicate charge",
          "delivery-002": "Package late",
          "account-003": "Cannot sign in"}
results = [("account-003", "account"), ("billing-001", "billing")]
joined = {record_id: label for record_id, label in results}
missing = set(inputs) - set(joined)
assert missing == {"delivery-002"}
```

Reject duplicate result IDs and unknown IDs before committing changes. Keep the original input hash, model, prompt version and batch identifier with the job. A successful batch-level status does not eliminate the need to inspect each record's status and validate its output.

## Design the deadline and recovery path

Set a business deadline separately from the provider's processing window. If results are needed at 09:00, submitting at 08:55 to a service that may take hours is not a viable plan. Allow time to retrieve results, validate them and handle failures.

Retry only records that need another attempt. Resubmitting all 100,000 records because 200 failed wastes work and can duplicate downstream updates. An application record ID should survive resubmission, while attempt IDs remain distinct. Apply the same principle when a cancellation leaves a mixture of completed and unfinished items.

If an urgent subset must use real-time inference, record that escalation separately and prevent a late batch result from overwriting a newer accepted result. Decide which outcome wins by version or job state, not by whichever response arrives last.

## Choose the smallest useful experiment

Start with a representative set containing valid, invalid and ambiguous records. Measure acceptance, completion time, retrieval failures and total charges across attempts. Use the [structured-output guide](/blog/structured-outputs-json-mode) if downstream code expects a fixed schema, and the [429 guide](/blog/openai-compatible-api-429-errors) for rate-limited real-time recovery.

Use the [API cost calculator](/tools/api-cost-calculator) to estimate the supported real-time model baseline. Apply a batch discount only when the actual destination documents it for your operation. If the work must finish interactively or batch operations are unavailable on your route, keep real-time execution and optimize prompt size, model choice and retry behavior instead.

## Frequently asked questions

### Does my own queue qualify for batch pricing?

Not automatically. The request must use an eligible batch operation on a destination that documents batch pricing. Queued normal API calls retain their applicable normal rates.

### Can batch results be joined by line number?

Do not assume matching order. Use stable request identifiers, inspect per-record status and reject missing, duplicate or unknown results before applying updates.

### Does this article announce KeepRouter Batch API support?

No. Native-provider batch terms are discussed for comparison. Use KeepRouter's published API reference to determine available operations on its routes.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [OpenAI Batch API](https://developers.openai.com/api/docs/guides/batch)
2. [Claude Message Batches](https://platform.claude.com/docs/en/build-with-claude/batch-processing)

## Related guides

- [Structured Outputs vs JSON Mode: Validate Real Data](https://keeprouter.com/blog/structured-outputs-json-mode.md)
- [OpenAI-Compatible API 429 Errors: Limits, Quota and Retry](https://keeprouter.com/blog/openai-compatible-api-429-errors.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)

## Put your workload into the cost estimate

Choose a model and enter expected usage. Compare the estimate with a small real request before scaling.

[Estimate API costs](https://keeprouter.com/tools/api-cost-calculator)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
