Support agent cost per resolution: a measurement worksheet

Token price is an input, not the economic result of a support agent. Classify outcomes at case level, join model charges to retries and review labor, then divide by verified resolutions and correctly escalated cases. Keep those denominators visible so a cheaper request cannot hide a worse customer outcome.

Published 2026-09-28 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

AI gateway evaluation scorecard covering contract, evidence, failure, security, and operations
Compare the cost of verified outcomes, including retries and human review, rather than the price of a single model call.

Short answer: compare support-agent candidates by cost per verified outcome, not cost per model call. A case may require retrieval, several model turns, a tool, a retry and human review. Count the whole case and label the outcome before dividing money by volume. A gateway request log is useful evidence for tokens and charges; it does not by itself prove that a customer issue was resolved.

This worksheet is a method and a hypothetical example, not a KeepRouter customer study or a claim that automated support is cheaper than people. The support-agent team page describes the safe release path.

Define the denominator first

Case labelEvidence requiredCount toward
Verified autonomous resolutionCustomer confirms or a case-specific review finds correct completionVerified resolutions
Correct escalationRequired facts, permission and transcript reach an authorized humanAcceptable outcomes, tracked separately
ReopenedSame issue returns within the chosen windowFailure and review queue
Unsupported answerClaim not grounded in approved knowledgeFailure and knowledge repair
Unsafe tool actionTool exceeded authority or changed wrong recordIncident, never counted as resolved

Choose and publish the reopening window and review rule before collecting results. Otherwise a team can inflate resolution by closing cases too early. Keep the model's draft, retrieved sources and tool arguments in a privacy-controlled application trace; do not mistake a 200 gateway response for a verified outcome. OpenAI's agent evaluation guide describes tracing model calls, tool calls, guardrails and handoffs. Anthropic's tool-use guide is explicit that client tools execute in the application.

Build the cost ledger

For each case, sum the actual charged model requests, not an advertised list price times an ideal turn count. Add retries, retrieval or embedding calls, external tool charges and human review time. Keep those columns separate:

ColumnUnitWhere it comes from
Model chargeUSD per requestGateway usage record or provider invoice
Retry and failed-turn chargeUSD per caseSame request ledger, grouped by case ID
Retrieval and tool chargeUSD per caseIndex or tool service ledger
Human reviewMinutes × loaded hourly rateWorkforce ledger; state the assumed rate
OutcomeOne of the fixed case labelsIndependent review or customer signal

KeepRouter's usage view records model, route, status, measured token counts, latency and charge without storing prompt or completion bodies in the request log. The application must attach its own case ID and result. Use a private join key; avoid placing customer text or account details in public telemetry.

Worked example: illustrative numbers only

Suppose a test period contains 100 cases: 60 verified autonomous resolutions, 25 correct escalations, and 15 reopened, unsupported or unsafe outcomes. Suppose the combined model and retry charges are $18, retrieval and tool charges $6, and human review time is valued at $36. The total measured cost is $60.

  • Cost per verified autonomous resolution, with the whole program cost allocated to that denominator: $60 / 60 = $1.00.
  • Cost per acceptable outcome, if your policy counts a correct escalation: $60 / (60 + 25) = $0.71 after rounding.
  • The failure share is 15 / 100 = 15%; it is not cancelled by a low average cost.

These values are invented for arithmetic, not a benchmark or current KeepRouter price. The first ratio deliberately assigns costs of failed and escalated cases to resolutions; the second answers a different question. Report both with their denominators rather than selecting the prettier number. A fair comparison with human-only support or another model needs the same case mix, review rule, service hours and quality threshold. It also needs a plan for rare, high-impact errors that an average hides.

Test a model switch without moving the goalposts

Create a fixed, redacted case set: easy information requests, stale-policy traps, missing account records, unauthorized tool requests, emotional complaints and cases that must go to a human. Hold the knowledge snapshot and tool permissions fixed. Run each candidate using the exact route and model ID in the live catalog. Grade resolution, escalation, unsupported claims and tool arguments before examining cost. If quality gates fail, the price comparison is irrelevant.

After offline evaluation, route only a small permitted cohort to the new model. Record case ID, model ID, application version, retrieval version, request IDs and final outcome. Join gateway charges to the case ledger and compare the same time window. A fallback should not repeat a refund or account action. The agent builder guide discusses tool retries and step limits; the gateway evaluation framework provides broader migration gates.

Keep escalations separate from autonomous resolutions

A correct escalation is useful, but it is not an autonomous resolution. Report both counts and their costs instead of changing the denominator to improve a headline. For model or prompt changes, compare the same cases with the same rubric; OpenAI’s agent-evaluation guide describes evaluation tooling, not a ready-made success definition for your business. Reopened cases should stay attached to the original measurement window so later failures are not hidden.

Publish the result without overclaiming

Report the number of cases and how many were reviewed, the reopening window, accepted outcome definitions, model and route versions, quality failures, p95 latency, and cost columns. State what was not measured. If the test has no real customer cases or independent labels, call it a simulation. Do not present synthetic conversations as customer savings. The useful decision is whether a specific, permitted workload passes your quality and budget gates, not whether the model has the lowest listed token price.

Frequently asked questions

Why not divide by all model requests?

A request is not a support outcome. One case may need multiple turns, retries or human review; report case-level results separately.

Does a correct escalation count as a resolution?

Track it separately. It may be an acceptable outcome, but calling it an autonomous resolution would inflate the rate.

Are the worked-example figures customer results?

No. Every figure in the worked example is hypothetical and illustrates the arithmetic only.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] OpenAI agent evaluation guide
  2. [2] Anthropic tool use guide
  3. [3] KeepRouter usage and observability
  4. [4] KeepRouter live model catalog

Related guides

← All posts · Models & pricing · Get an API key