LLM cost-control proxy · OpenAI and Anthropic formats

A ceiling you typed by hand knows nothing about what the customer pays you.

Outcap derives each end customer's ceiling from the revenue you declare for them, compares it to their spend this month on every request, and serves them on a cheaper model before cutting them off. The decision happens during the request, not in a report the next morning.

Said up front: if Outcap goes down entirely, your traffic stops, like with any proxy. This page also says the rest of what the product does not do.

POST /v1/chat/completionsgpt-4ox-outcap-user: cust_8f2route: assistant
  1. key
  2. kill switch
  3. input
  4. marginderived ceiling reached
  5. budgets
  6. rate
  7. cache
  8. dispatch
200 · response served
served model
gpt-4o-mini
x-outcap-margin
degraded
cap sent
480 tokens
request cost
$0.0009
One request from an end customer approaching the ceiling derived from their revenue. The order of the checks and the headers are the ones in the code.

The money

Where it goes, and what we take back

Three mechanisms, ordered by what they actually return. Each with its limit, because a number you cannot recompute is worth nothing.

Routing, the only exact saving

Route by route, you choose to serve a cheaper model from the same family. Outcap changes the model field and nothing else, then computes the saving to the cent: requested price minus served price, request by request. The gaps below are the providers' public prices, so you can check them yourself.

You decide whether a route can live with the smaller model. Outcap does not judge answer quality, and never routes to another provider: your key opens one world only.

Price it on my own bill
RequestedServedOutput, per million tokensGap
gpt-4ogpt-4o-mini$10.00 → $0.60−94%
gpt-5gpt-5-mini$10.00 → $2.00−80%
claude-sonnet-4-6claude-haiku-4-5$15.00 → $5.00−67%
claude-opus-4-5claude-sonnet-4-6$25.00 → $15.00−40%

Margin per end customer

REVENU DÉCLARÉ DU CLIENT100 %PLAFOND DÉRIVÉ = REVENU × PART ACCEPTÉEblocagedégradationDÉPENSE DU MOISzone laiton : modèle moins cher + plafond de sortie plus court, la requête passe

You declare what each customer pays you per month. The ceiling stops being a number someone guessed: it is that revenue times the share you accept to spend on it. When a customer gets close, they are served on the cheaper model with a shorter output cap. Blocking comes last, and only if you asked for it.

With no declared revenue, Outcap decides nothing. No billing connector, no invented ceiling, and a customer's first request always goes through.

The learned output cap

LONGUEUR DES RÉPONSES D'UNE ROUTEp99plafond = p99 × 1,3coupé proprement99 réponses sur 100 passent sans être touchéesschéma de mécanisme : la forme dépend de votre trafic, pas la règle

Outcap measures how long your answers really are, route by route, and sets a cap above the 99th percentile. A runaway answer is stopped at a sentence boundary, and a cut JSON body is repaired before it leaves.

Worth knowing: this cap does not lower your average bill. It sits above the 99th percentile, so it catches outliers, not everyday traffic. Anyone promising tens of percent from a cap is selling you an average they never measured.

The mechanism

Eight checks before your money leaves

In this order, in the same synchronous turn, with no database access at all. A refusal costs zero tokens.

CheckWhat it looks atWhat it returns if it refuses
Outcap keyThe project key, its state, its owner.401
Kill switchThe project switch, before any work proportional to the body.429 outcap_kill_switch
Input guardrailsAPI keys, personal data, invisible text in the prompt.400, or redaction before sending
MarginThe end customer's spend this month against the ceiling derived from their revenue.Cheaper model, then 429 if you chose it
BudgetsCounters per project, key, route or end customer, in memory.429 outcap_budget_exceeded, not retryable
Rate limitsA token bucket, in requests or in tokens.429 retryable, with retry-after
CacheExact body match, never semantic.The answer you already paid for, at 0 tokens
DispatchServed model, applied cap, backup key if yours is refused.The provider's answer
Read the mechanism in detail

The dashboard

What you look at on Monday morning

Screenshots of the product as it runs, taken on a demonstration project. The amounts are those of that demo traffic, not a projection, and no customer was invented for the picture.

Dashboard overview: amount saved, recommended actions, requests, tokens, cost, simulated and realised savings, cost per day and top routes.
Real screenshotWhat you spent, what was taken back, and the actions your own data justifies.
Logs page: one row per request with time, route, model, status, tokens, cost, latency and what was particular about the request.
Real screenshotOne row per request: served model, cost, decision. Here two refusals, one on budget and one on rate, each tied to the end customer.

Thirty seconds

A real call, on us

No account, no key. The response below comes out of the production proxy, cap applied and clean cut included.

curl -s -X POST https://proxy.outcap.tech/try -H "content-type: application/json" -d '{"prompt": "Why do model answers so often run long?"}'

proxy response

{
  "mode": "live",
  "model": "gemini-2.5-flash",
  "response": "LLM responses running too long is a common phenomenon, and it's not
    due to a single factor, but rather a complex interplay of how these models are
    designed, trained, and used. […] During training, they are optimized to
    maximize the likelihood of the training data.",
  "outcap": {
    "cap_applied": 120,
    "output_tokens": 120,
    "finish_reason": "length",
    "clean_cut": "sentence_boundary",
    "cost_usd": 0.0003084
  }
}

A real call, fixed model, fixed cap, paid by us. The demo has a daily budget: past it the answer is labelled as simulated, never dressed up as a real call.

What Outcap does not do

Ten competitors out of twelve state no limitation anywhere on their commercial pages. Here are ours, before you find them in production.

One instance only

The counters live in memory: that is what makes the decision free in latency, and it is also why two instances would let twice the ceiling through. The proxy detects it, writes it in its logs and shows it in its health endpoint.

If we go down, you go down

A database outage does not cut you off: known keys keep being served and traffic flows. But if the proxy itself is unavailable, your traffic stops. That is true of any proxy, including the ones that do not write it down.

Two providers

OpenAI and Anthropic, each in its own format. No house provider, no local model, no translation from one format to the other.

Input guardrails are incomplete

They recognise checkable formats: API keys, cards with their Luhn digit, IBANs, social security numbers. Not names, not postal addresses. Prompt injection attempts are flagged, never blocked: a new wording gets through, and claiming otherwise would be a lie.

No evals, no observability

Outcap does not compare your prompts and does not score your answers. It sends its traces to the tool you already use.

Young beta

No certification, no status page, no production history yet. Leaving takes one line of configuration, and the code can run on your own machines.

The proof

Why believe a guardrail

A guardrail you believe is active and is not is worse than no guardrail. Here is what you can check from your side.

361

defects put back

After each fix, the defect is put back into the code to check that at least one test fails. A passing test proves nothing until you have seen what makes it fail.

1,155

tests, all green

Every guarantee published here has its test: budgets held under concurrent bursts, JSON never broken by a cut, no prompt ever logged.

9

direct dependencies

The command to recount them is on the security page. After the LiteLLM supply chain attack of March 2026, installed surface is an argument.

1

line to leave

You changed one base URL to come in. You change it back to leave, and your keys stay yours, at the provider.

The test and defect figures are from 16 September 2026 and can be recounted with pnpm test and node scripts/mutation.mjs.

Plug it in observation mode, decide afterwards

The default mode changes no request at all. You see what Outcap would have done, then you switch on what you want, route by route.

Outcap, the proxy that decides before your money leaves