llm-guard
A self-hosted gateway that sits in front of your OpenAI, Anthropic or Gemini calls. It records what each request actually cost, attributes that cost to the API key, project and end user behind it, detects runaway agent spend while it is still happening, and enforces hard budgets.
It runs on the Python standard library. No pip install, no virtualenv, no third-party package anywhere in the request path.
Source on GitHub → · Read the notes →
The problem it solves
A budget that checks between requests cannot stop the request that is crossing the line. For a short chat call that hardly matters. For a 200-step agent loop it is the entire problem: OWASP records that a 200-step loop costs more than 100x a single call, and that 62% of agent bills come from context being re-sent on every step.
llm-guard watches for three shapes of runaway spend:
- Context loops — a sustained pathological input-to-output token ratio. Normal traffic runs 5:1 to 15:1. The documented production incident hit 74:1 and 175:1.
- Velocity bursts — spend rate far above your own account's baseline, not above some absolute number that would be wrong for everyone.
- Retry storms — bursts of failing calls, which still cost latency and often tokens.
Detection only ever reports. Enforcement is a separate, explicit decision: a per-key budget with action=block returns HTTP 429, and optional per-stream caps abort a single streaming response mid-flight.
What makes it different
| Zero runtime dependencies | Standard library only. Auditable in one sitting, which matters because a gateway on the request path is a high-value target. |
| Cost from billing truth | Prices come from the usage block the provider returned, never from a local token estimate. |
| Per-model cache economics | Cache reads are billed at 2.5% to 50% of input depending on the model. Treating that as a constant is a silent 4-5x error. |
| No data leaves your network | No telemetry, no CDN, no phone-home. The dashboard is server-rendered SVG. |
| Unknown is a valid answer | A model with no price on file is reported as unpriced, not estimated. |
Try it in thirty seconds
No API key, no network access, no signup:
git clone https://github.com/leyao-daily/llm-guard.git
cd llm-guard
python3 -m llmguard seed --reset --compare-days 30
python3 -m llmguard anomalies
python3 -m llmguard dashboard --out dash.html
The demo dataset contains a real unbounded-context loop, so the detector has something to find on the first run.
Measured overhead
Benchmarked against a zero-latency local upstream, which is the worst case for a ratio — real LLM calls take 400 to 4000ms:
| Scenario | req/s | mean | p95 |
|---|---|---|---|
| Direct to upstream | 10,136 | 2.87 ms | 3.20 ms |
| Through llm-guard | 3,688 | 8.12 ms | 11.81 ms |
| Through llm-guard (SSE streaming) | 3,851 | 7.70 ms | 8.33 ms |
About 5 ms per request, and streaming responses are metered correctly, including usage that only arrives in the final frame.
What it does not do
It does not terminate inbound TLS — put a load balancer in front of it. It does not cache responses, retry requests, or store prompts and completions. Those are deliberate omissions rather than a roadmap.
Licensed MIT. Copyright LYE LABS LIMITED.