How to evaluate an AI gateway: tests, costs and a filled scorecard
Run the same workflow on two gateway candidates, reject missing required features, and compare accepted tasks, total costs and latency with a practical scorecard.
Published 2026-08-15 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

Evaluate an AI gateway with one application workflow and the same test cases on each candidate. Reject a route that breaks a required feature before comparing its speed or price. For the remaining options, measure accepted tasks, complete task cost and user-visible latency. A small reproducible trial is more useful than scoring dozens of features your application will never use.
Choose two candidates and one workflow
For a streaming support assistant, the required behaviors might be: answer from supplied product information, return citations, call an order-lookup tool and stop cleanly on cancellation. Build a fixed set with ordinary questions, missing information, malformed tool inputs and longer conversations. Use synthetic records or appropriately de-identified examples. The sample size should reflect your risk; the 20-case example below is an early screening exercise, not a reliability benchmark.
Choose the product class before testing. OpenRouter documents caller-configurable provider routing; LiteLLM documents a proxy you can operate. KeepRouter provides managed catalog access through its documented API. These can serve similar requests while giving the customer different operating responsibilities. The gateway shortlist helps narrow the field.
Write pass conditions before seeing the answers
| Test | Pass condition | Why it matters |
|---|---|---|
| Text response | Required facts are present and unsupported claims are absent | A fluent answer may still be wrong |
| Streaming | First useful text appears and the stream reaches a valid end | An opened connection is not a completed answer |
| Tool round trip | Arguments validate and the matching result is consumed | Producing a tool name is only half the workflow |
| Missing knowledge | The assistant asks for clarification or declines to guess | A wrong confident answer can cost more than a retry |
| Cancellation | The UI stops and the application does not launch more work | Abandoned tasks can still create spending |
| Failure | Permanent errors stop; retryable errors stay within the task budget | Multiple retry layers can amplify an outage |
Set any thresholds from your application's requirements. Do not borrow a vendor's p95 figure as your latency target. Record time to first useful output and time to completed task separately, including retries. Keep the endpoint, model revision, output settings and test inputs consistent across candidates.
A filled-out scorecard, with illustrative results
The following numbers are invented examples to explain the decision, not measurements of KeepRouter or another service. Each candidate receives the same 20 tasks; the totals include failed and repeated attempts.
| Result | Candidate A | Candidate B | Decision implication |
|---|---|---|---|
| Accepted tasks | 18 of 20 | 16 of 20 | Inspect which tasks differ |
| Total inference cost | $0.36 | $0.24 | Raw spending favors B |
| Cost per accepted task | $0.020 | $0.015 | B remains cheaper on this small set |
| Required tool cases passed | 5 of 5 | 4 of 5 | B fails a required behavior |
| First useful output, median | 0.9 s | 0.7 s | Faster text does not repair the failed tool case |
If that tool behavior is mandatory, B should not receive production traffic until the failure is understood. If the application can keep tool tasks on another route, a split is possible. State that choice explicitly instead of averaging the failure into a single score. With only 20 tasks, report counts and examples; do not claim statistically established superiority or an uptime percentage.
Count the work that a request benchmark misses
Model charges are only one part of adoption. Include the time to map model IDs, remove product-specific fields, migrate monitoring, reconcile charges and operate the gateway. A self-hosted proxy also needs a named owner for updates, secrets, backups and incidents. Ask whether a missing feature can remain in your application without creating more maintenance than the gateway removes.
For managed routes, inspect the service's documented data handling and the exact settings you will use. A control listed in a product overview may depend on the chosen route or plan. KeepRouter's security page, model catalog and API reference answer different questions. Avoid scoring an undocumented feature as present merely because another gateway has it.
Move from the test to one controlled rollout
Start with the API migration checker if you already have an OpenAI-style configuration. It checks credential-free JSON locally; it does not execute inference or establish model quality. Then run the cases above using a restricted test key and small output limits. Use the migration guide for endpoint, stream and error checks.
Keep the previous configuration deployable. Expand traffic only after your required cases pass, costs reconcile and a rollback has been exercised. Record the test date, configuration, unresolved failures and the reason for choosing a candidate. Repeat the relevant subset when a model or SDK changes. This produces a decision your team can revisit, rather than a score that looks precise but has no operational meaning.
Frequently asked questions
Can a feature matrix choose an AI gateway?
It can shortlist candidates. Acceptance should come from reproducible tests of your exact contract, workload, data policy, and operating responsibilities.
Does a health check prove the generation route works?
No. It proves only the health surface at that time. A scoped generation test is needed for model eligibility, authentication, protocol, and usage evidence.
Sources reviewed
Article last reviewed 2026-09-29
- [1] Cloudflare AI Gateway documentation
- [2] Portkey AI Gateway documentation
- [3] LiteLLM documentation
- [4] OpenRouter quickstart
- [5] KeepRouter OpenAPI
- [6] OpenRouter documents caller-configurable provider routing
- [7] LiteLLM documents a proxy you can operate