LLM cost-control proxy · OpenAI and Anthropic formats
A ceiling you typed by hand knows nothing about what the customer pays you.
Outcap derives each end customer's ceiling from the revenue you declare for them, compares it to their spend this month on every request, and serves them on a cheaper model before cutting them off. The decision happens during the request, not in a report the next morning.
Said up front: if Outcap goes down entirely, your traffic stops, like with any proxy. This page also says the rest of what the product does not do.
- key
- kill switch
- input
- marginderived ceiling reached
- budgets
- rate
- cache
- dispatch
- served model
- gpt-4o-mini
- x-outcap-margin
- degraded
- cap sent
- 480 tokens
- request cost
- $0.0009
The money
Where it goes, and what we take back
Three mechanisms, ordered by what they actually return. Each with its limit, because a number you cannot recompute is worth nothing.
Routing, the only exact saving
Route by route, you choose to serve a cheaper model from the same family. Outcap changes the model field and nothing else, then computes the saving to the cent: requested price minus served price, request by request. The gaps below are the providers' public prices, so you can check them yourself.
You decide whether a route can live with the smaller model. Outcap does not judge answer quality, and never routes to another provider: your key opens one world only.
Price it on my own bill| Requested | Served | Output, per million tokens | Gap |
|---|---|---|---|
| gpt-4o | gpt-4o-mini | $10.00 → $0.60 | −94% |
| gpt-5 | gpt-5-mini | $10.00 → $2.00 | −80% |
| claude-sonnet-4-6 | claude-haiku-4-5 | $15.00 → $5.00 | −67% |
| claude-opus-4-5 | claude-sonnet-4-6 | $25.00 → $15.00 | −40% |
Margin per end customer
You declare what each customer pays you per month. The ceiling stops being a number someone guessed: it is that revenue times the share you accept to spend on it. When a customer gets close, they are served on the cheaper model with a shorter output cap. Blocking comes last, and only if you asked for it.
With no declared revenue, Outcap decides nothing. No billing connector, no invented ceiling, and a customer's first request always goes through.
The learned output cap
Outcap measures how long your answers really are, route by route, and sets a cap above the 99th percentile. A runaway answer is stopped at a sentence boundary, and a cut JSON body is repaired before it leaves.
Worth knowing: this cap does not lower your average bill. It sits above the 99th percentile, so it catches outliers, not everyday traffic. Anyone promising tens of percent from a cap is selling you an average they never measured.
The mechanism
Eight checks before your money leaves
In this order, in the same synchronous turn, with no database access at all. A refusal costs zero tokens.
| Check | What it looks at | What it returns if it refuses |
|---|---|---|
| Outcap key | The project key, its state, its owner. | 401 |
| Kill switch | The project switch, before any work proportional to the body. | 429 outcap_kill_switch |
| Input guardrails | API keys, personal data, invisible text in the prompt. | 400, or redaction before sending |
| Margin | The end customer's spend this month against the ceiling derived from their revenue. | Cheaper model, then 429 if you chose it |
| Budgets | Counters per project, key, route or end customer, in memory. | 429 outcap_budget_exceeded, not retryable |
| Rate limits | A token bucket, in requests or in tokens. | 429 retryable, with retry-after |
| Cache | Exact body match, never semantic. | The answer you already paid for, at 0 tokens |
| Dispatch | Served model, applied cap, backup key if yours is refused. | The provider's answer |
The dashboard
What you look at on Monday morning
Screenshots of the product as it runs, taken on a demonstration project. The amounts are those of that demo traffic, not a projection, and no customer was invented for the picture.


Thirty seconds
A real call, on us
No account, no key. The response below comes out of the production proxy, cap applied and clean cut included.
curl -s -X POST https://proxy.outcap.tech/try -H "content-type: application/json" -d '{"prompt": "Why do model answers so often run long?"}'proxy response
{
"mode": "live",
"model": "gemini-2.5-flash",
"response": "LLM responses running too long is a common phenomenon, and it's not
due to a single factor, but rather a complex interplay of how these models are
designed, trained, and used. […] During training, they are optimized to
maximize the likelihood of the training data.",
"outcap": {
"cap_applied": 120,
"output_tokens": 120,
"finish_reason": "length",
"clean_cut": "sentence_boundary",
"cost_usd": 0.0003084
}
}A real call, fixed model, fixed cap, paid by us. The demo has a daily budget: past it the answer is labelled as simulated, never dressed up as a real call.
What Outcap does not do
Ten competitors out of twelve state no limitation anywhere on their commercial pages. Here are ours, before you find them in production.
One instance only
The counters live in memory: that is what makes the decision free in latency, and it is also why two instances would let twice the ceiling through. The proxy detects it, writes it in its logs and shows it in its health endpoint.
If we go down, you go down
A database outage does not cut you off: known keys keep being served and traffic flows. But if the proxy itself is unavailable, your traffic stops. That is true of any proxy, including the ones that do not write it down.
Two providers
OpenAI and Anthropic, each in its own format. No house provider, no local model, no translation from one format to the other.
Input guardrails are incomplete
They recognise checkable formats: API keys, cards with their Luhn digit, IBANs, social security numbers. Not names, not postal addresses. Prompt injection attempts are flagged, never blocked: a new wording gets through, and claiming otherwise would be a lie.
No evals, no observability
Outcap does not compare your prompts and does not score your answers. It sends its traces to the tool you already use.
Young beta
No certification, no status page, no production history yet. Leaving takes one line of configuration, and the code can run on your own machines.
The proof
Why believe a guardrail
A guardrail you believe is active and is not is worse than no guardrail. Here is what you can check from your side.
361
defects put back
After each fix, the defect is put back into the code to check that at least one test fails. A passing test proves nothing until you have seen what makes it fail.
1,155
tests, all green
Every guarantee published here has its test: budgets held under concurrent bursts, JSON never broken by a cut, no prompt ever logged.
9
direct dependencies
The command to recount them is on the security page. After the LiteLLM supply chain attack of March 2026, installed surface is an argument.
1
line to leave
You changed one base URL to come in. You change it back to leave, and your keys stay yours, at the provider.
The test and defect figures are from 16 September 2026 and can be recounted with pnpm test and node scripts/mutation.mjs.
Plug it in observation mode, decide afterwards
The default mode changes no request at all. You see what Outcap would have done, then you switch on what you want, route by route.