Comparison
They block before the call. So do we.
This page claimed otherwise for a long time, and that was false. Since mid-2026, Cloudflare AI Gateway, LiteLLM and Portkey all refuse a request before calling the provider, and per-end-customer budgets exist there too. Here is what is still true, what they do better than us, and when you should pick them.
Two claims were removed from this page after the September 2026 audits: that we were alone in blocking before the call, and that per-end-customer budgets were an empty slot. Both were false.
The difference in one line
Everywhere else a ceiling is an amount somebody typed by hand, and margin is a report read the next morning. Here the ceiling is derived from the revenue you declare for that customer, and margin is a decision input during the request. The second place where we are alone is less noble and more profitable: the length of the answers.
Capability by capability
Fifteen rows, several of which we lose. A cell in pale ink is a cell lost, ours included.
| Capability | Outcap | Cloudflare AI Gateway | Portkey | LiteLLM |
|---|---|---|---|---|
| Refusal before the provider call | yes | yes | yes | yes |
| Budget per end customer | yes, project-scoped | yes, by metadata or Access identity | yes, Enterprise self-hosted | yes |
| Ceiling derived from customer revenue | yes | no | no | no |
| Margin decided during the request | yes | no | no | no |
| Degradation before cut-off | yes, model and length | no | no | no |
| Output cap learned per route | yes, percentile and margin | no | no | static, set by hand |
| Cut at a sentence boundary, JSON repaired | yes | no | no | no |
| Routing to a cheaper model | yes, saving priced per request | yes | yes | yes |
| Response cache | yes, exact | yes | yes | yes |
| OpenTelemetry export | yes, no content | yes, prompts included | yes, Enterprise | yes |
| Observability platform, evals, replay | no | basic | yes | partial |
| Providers covered | OpenAI and Anthropic | multiple | over 1,600 | over 100 |
| Counters shared across instances | no, a single instance | hosted service | hosted service | yes, with Redis |
| Setup | two lines, zero infrastructure | Cloudflare account | hosted service | self-hosted, Postgres and Redis |
| Entry price | free beta | free | free tier, then paid | open source, free |
Checked against each product's public documentation in September 2026. A note on the degradation row: Cloudflare can fall back to another model when the provider refuses, which is a failure mechanism and not the one described here; Outcap's degradation is decided by the end customer's margin, before the call, on a request the provider would have accepted. According to its documentation (https://developers.cloudflare.com/ai-gateway/observability/otel-integration/), Cloudflare's OpenTelemetry export sends prompts and responses, with no documented option to strip them; Portkey's export to your own tool is reserved for Enterprise and self-hosted plans. Other products do compute margin per end customer, Amberflo, Paid.ai and Weflayr among them: they display it and they alert on it, none of them uses it to decide during the request. Competitors move fast: tell us about a stale cell and it gets fixed.
Plainly
When to choose another tool
Three situations where we would tell you to walk away, because a comparison that never sends anyone elsewhere is not a comparison.
Cloudflare AI Gateway
If you want a free spend ceiling and your infrastructure is already there. It is free, it is proven, and it covers more providers than we do. What you will not find: a ceiling that knows what your customer pays you, and control over answer length. Check what its trace export carries in the way of content before wiring it into your observability tool.
Portkey
If you are a large team that needs governance, guardrails and compliance at scale, with a very wide provider catalogue and a real observability platform. Worth knowing before signing: per-end-customer budgets and trace export to your own tool sit on the Enterprise and self-hosted plans.
LiteLLM
If you want a hundred providers, open code, a gateway running on your own machines, and you know how to operate Postgres and Redis. It reserves the worst case before the call like us, and its counters are shared across instances, which ours cannot do. Read its open tickets on budget enforcement before relying on it, including issue 28750 of 24 May 2026, still without a maintainer answer.
The other case
When to choose us
When you resell AI to your own customers and a single one of them can cost you more than they pay you. The margin platforms will tell you which one, next month, in a report. Outcap derives that customer's ceiling from the revenue you declare for them, compares it to their spend this month on every request, and serves them on a cheaper model with a shorter answer before cutting them off. With no declared revenue it decides nothing: no billing connector, no invented ceiling.
What Outcap does not do
The six lines our competitors leave off their marketing pages, and that you would otherwise meet in production.
No observability platform
No evals, no replay, no prompt comparison. The OpenTelemetry export sends content-free traces to the tool you already use, it does not replace it.
Two providers
OpenAI and Anthropic, in their own formats. No translation from one format to the other, no local model, no four-figure catalogue.
A single instance
The counters live in memory, which is what makes the decision free in latency and also why scaling out is forbidden in our deployment file. The hosted services opposite do not have this problem.
An exact cache, nothing more
Exact match on the normalised body, never semantic, never on streaming. A fallback response, or one our cap truncated, never enters the cache.
Input guardrails are incomplete
They recognise verifiable formats: API keys, cards with their Luhn check, IBANs, social security numbers. Not names, not postal addresses, not images. Prompt injection attempts are flagged and never blocked.
A young beta
No certification, no status page, no production history to show. In exchange, the code runs on your machines and leaving takes one line of configuration.
Check it on your own traffic
The default mode changes no request. It is the only way to find out which of these rows matters to you.