Questions

Objections first

Nine questions an engineer asks before putting a proxy in front of production. The answers reflect what is true in September 2026, including where it does not flatter us.

A question that is not here? Write to hello@outcap.tech. If it matters to someone else, the answer lands on this page.

What it is worth

Does this actually buy me anything

01I set max_tokens myself, in my own code.

And that is the right practice. If you already set it route by route from your real answer lengths, our cap will buy you close to nothing: it will confirm your number. What it adds is measuring instead of guessing, raising itself when it truncates more than 2% of answers, and standing down where a cap does damage: never on a reasoning model, never on a request that declares tools, never before ten samples on the route. And if it is your bill that worries you, this is not where it goes down anyway.

The cap we send is never higher than the max_tokens you wrote yourself.

02How much does the output cap really save?

Little, and we would rather say it on the pricing page than in a support ticket. It sits above the 99th percentile of your answer lengths, with a margin on top: by construction it catches outliers, not everyday traffic, and a saving is only counted when the answer exceeds that cap. We publish no percentage for it, because the only honest one would be measured on your traffic. The single exact saving in the product is routing: the gap between two public prices, computed request by request. The cap does not lower an average bill, it stops one runaway answer from raising it.

The Analyze page prices routing from your provider's usage export, with no account, and the file never leaves your browser.

03A cap will truncate my answers mid-sentence.

That is exactly the defect the clean cut handles. When our cap is the one cutting, the answer is pulled back to the last sentence boundary, provided at least 60% of it survives, and a damaged JSON body is repaired before it leaves: the invariant is tested, our cut never produces invalid JSON. In streaming, repairing on the fly would mean buffering the whole answer, so adding latency: we refuse, and the cap takes a wider margin instead. Tool calls and multi-choice responses are never rewritten. And if it is your own max_tokens doing the truncating, we touch nothing: we do not repair a cut we did not cause.

A cap that truncates more than 2% of a route's answers is raised automatically, with nothing for you to watch.

The others

Why not them

01Cloudflare AI Gateway already caps my spend, for free.

It does, and if that is all you need, take it: it is free, it is proven, and your infrastructure may already be there. We do not claim to be the only ones refusing a request before the call, this page used to say so and it was false. Two things are not there: a ceiling derived from what your customer pays you, and control over answer length. One thing to check before wiring their trace export into your observability tool: according to their documentation it carries prompts and responses, with no documented option to strip them.

https://developers.cloudflare.com/ai-gateway/observability/otel-integration/

02LiteLLM is open source and covers a hundred providers.

And it reserves the worst case before the call, like us: that mechanism does not set us apart, contrary to what this page claimed for a long time. If you want a hundred providers, open code, and you know how to run Postgres and Redis, it is probably the right choice, all the more so since its counters are shared across instances and ours are not. Our argument is not that the mechanism is exclusive, it is that it is verified: every guarantee published here has its test, and every fix is checked by putting the defect back into the code to watch a test fail. Before relying on its budget enforcement in production, read its open tickets on the subject, including issue 28750 of 24 May 2026, which asks for a project-scoped end-customer budget and has no maintainer answer.

Outcap speaks two formats only, OpenAI and Anthropic. That is a limit, not a stance.

The risk

Putting it in front of production

01Do you log my prompts?

No. What gets written to the database is metadata: request id, route, requested and served model, status, tokens, cost, overhead in milliseconds, and the decisions taken. The route fingerprint is a hash of the system prompt, not the prompt. Two things we would rather say ourselves: the x-outcap-user header is stored exactly as you send it, so if you put an email in there, we store an email; and when you turn on the response cache, response bodies sit in RAM for their time to live, swept every thirty seconds. The right way to believe us is not to believe us: send a recognisable sentence, export your logs as CSV, search for it.

The Security page lists what is stored, what never is, and how to check it yourself.

02How much latency does the proxy add?

We publish no number, and that is deliberate. Almost every number published in this category is measured against a fake upstream, which removes the dominant variable, the model's generation time: those numbers are true and useless. What we publish instead is the method and the script that reruns it: percentile against percentile at the same concurrency, alternating series so a provider drift hits both, warm-up excluded, no request discarded. On your traffic the measurement already runs: every request logs its own overhead in milliseconds, the overview shows the median and the 95th percentile, and the column is in the CSV export. Until there is a status page with several months of history, a number printed here would read as bluff.

Full method in docs/BENCHMARK.md, script pnpm bench. It prints by itself that a result against a fake upstream flatters us.

03What happens if Outcap goes down?

Case by case, including the case that does not suit us. Database unavailable: traffic flows, already-known keys are served from cache even when stale, unknown routes get temporary in-memory state; you lose statistics, and during that window a budget scoped to a route does not apply to a route that has not been warmed yet. Proxy unavailable: your traffic stops, exactly as with any proxy, and that is the sentence most often missing on the other side. Leaving takes one line: you put the provider's base URL back, and your keys never left your code. The least comfortable point stays scaling: the counters live in memory in a single process, so with two instances real spend can reach the limit multiplied by the number of instances. The proxy detects its peers, writes it to its logs, and shows it on its health endpoint.

04Will you still be around in eighteen months?

Nobody can promise you that, and a promise would be the worst possible argument. The facts: Outcap is written by one person, with no funding, and no paying customer to date. What is under your control: the code runs on your machines with a compose file and a guide, your provider keys are never stored or logged, your metadata exports to CSV whenever you want, and rolling back is one environment variable. After the acquisitions and the maintenance-mode announcements of 2026 in this category, an exit clause is worth more than a guarantee clause.

A question that is not on this list

The most useful answers on this page came from questions asked by people who did not have an account yet.

FAQ · Outcap