RAG Cost per Query: A Worksheet Beyond Token Prices
Calculate RAG cost across indexing, retrieval, generation and retries. Use a worked example to compare cost per accepted answer instead of price per token.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

RAG cost per query is the allocated cost of preparing the knowledge base plus the cost of retrieving evidence and producing an answer. For a useful business comparison, divide that total by accepted answers, not just requests. A cheap answer that invents a policy still costs money to generate and repair.
This worksheet separates costs that change with document volume from costs that change with traffic. All prices and workloads in the example below are hypothetical planning inputs, not current vendor prices or measured KeepRouter results.
Separate indexing from answering
An embedding turns content into vectors for retrieval. The OpenAI embedding guide and Gemini embedding guide explain their respective APIs. Neither the existence of an embedding endpoint nor a low embedding price tells you the total cost of a retrieval application.
Build the ledger before choosing a generation model:
| Cost | What drives it | Where to record it |
|---|---|---|
| Parsing and OCR | Pages, file types and repeated processing | Ingestion job |
| Document embeddings | Tokens embedded, including changed documents | Index build or update |
| Vector storage | Stored vectors, replicas and service plan | Monthly infrastructure |
| Query embedding | Search requests, including retries | Query trace |
| Retrieval and reranking | Searches, candidates and reranker calls | Query trace |
| Answer generation | Input, output and applicable token categories | Each model attempt |
| Review and repair | Rejected answers and reviewer time | Outcome record |
Do not embed the entire knowledge base every time someone asks a question. Persist the index, identify changed documents and measure update work separately. If a model migration requires rebuilding embeddings, treat that rebuild as a migration cost instead of hiding it in ordinary query usage.
Work through a monthly example
Suppose a small support assistant handles 10,000 questions in a month. Allocate $20 for ingestion and document updates, $30 for vector infrastructure, $2 for query embeddings and $10 for retrieval plus reranking. Together these costs are $62.
Each generation attempt uses 3,000 input tokens and 400 billed output tokens. At illustrative prices of $1 per million input tokens and $5 per million output tokens, an attempt costs $0.005. If 1,000 questions each need exactly one extra generation attempt, there are 11,000 attempts and generation costs $55.
questions = 10_000
extra_attempts = 1_000
input_tokens = 3_000
output_tokens = 400
attempt_cost = (input_tokens * 1 + output_tokens * 5) / 1_000_000
generation = (questions + extra_attempts) * attempt_cost
total = 20 + 30 + 2 + 10 + generation
accepted = 9_000
print(round(total, 2)) # 117.0 dollars
print(round(total / questions, 4)) # 0.0117 per question
print(round(total / accepted, 4)) # 0.0130 per accepted answerThe $117 total excludes reviewer labor and other application infrastructure. Add those explicitly when comparing a production business case. Do not interpret the $0.013 figure as a quote for any product. Replace every input with your own measured workload and the destination's current billing rules.
Define an accepted answer before optimizing
For a policy assistant, use three short, fictional documents: a domestic return policy, an international return policy and an expired promotion. Ask which rule applies to an international order placed after the promotion ended. A passing answer must select the current international rule and cite the correct document.
Add an unanswerable question about a policy absent from the documents. The accepted outcome should acknowledge the missing information rather than invent a rule. Keep those two tasks in the evaluation set when changing chunk size, retrieval count or model.
This fixture exposes a common accounting mistake: reducing retrieved tokens can lower request cost while removing the evidence needed for a correct answer. Track evidence coverage and answer acceptance alongside cost. See the LlamaIndex integration guide for a generation-only experiment that preserves the retrieval setup.
Find the expensive stage with measurements
Attach one application query ID to retrieval, generation attempts and the final outcome. Record document IDs and token counts rather than copying private document text into general analytics. When an answer is rejected, label the cause: missing source, wrong retrieval, unsupported inference or formatting failure.
If generation dominates, shorten repetitive instructions or test a lower-cost model on the same evidence. If retrieval dominates, inspect candidate counts and reranking frequency. If updates dominate, check whether unchanged documents are being processed repeatedly. These interventions solve different problems; switching a chat model will not fix an unnecessarily rebuilt index.
Compare quiet and busy months separately. A $30 fixed service charge contributes $0.03 per query at 1,000 queries but $0.003 at 10,000. Traffic changes the allocation without changing the service's quality. Capacity upgrades can also create step changes, so a single linear estimate may understate the next tier.
Turn the worksheet into a model decision
Use the API cost calculator for the supported model-token portion, then add ingestion, storage, retrieval and review in your own ledger. The calculator is not a complete RAG infrastructure bill. Select candidates from the model catalog, keeping the same documents, prompts and acceptance rules.
Start with a bounded sample that includes failures, then compare total spend divided by accepted answers. Record latency and the percentage needing review as separate outcomes. If the cheaper model requires substantially more retries or manual correction, the lower token price has not delivered a cheaper working system.
Frequently asked questions
What belongs in RAG cost per query?
Include allocated ingestion and storage, query embeddings, retrieval, reranking, every generation attempt and any review cost relevant to the decision. Keep fixed and variable costs separate.
Does a cheaper generation model always reduce RAG cost?
No. Retries, rejected answers and manual correction can outweigh token savings. Compare the total cost per accepted answer on the same evaluation set.
Are the worksheet prices current vendor quotes?
No. They are explicit hypothetical inputs that make the calculation reproducible. Replace them with current prices and measured usage before budgeting.
Sources reviewed
Article last reviewed 2026-09-29