Open source · MIT · Zero dependencies

Stop runaway AI agent spend before the invoice arrives.

llm-guard sits in front of your model API calls, attributes every dollar to the key that spent it, and detects a runaway loop while it is still running. Python standard library only — nothing to install, and nothing to audit but our code.

llm-guard anomalies
  Runaway-spend detection
  window: last 15 min   ·   io ratio > 30:1   ·   velocity > 25x baseline
  CRITICAL io_ratio  [key: prod-agent]
    prod-agent is running a 74:1 input-to-output ratio across 67 calls
    in the last 15 min (12,006,400 in / 162,216 out, $38.45). Normal
    traffic sits at 5:1-15:1.
    → what to do
      The prompt is being re-sent in full on every step. Cap the context
      (summarise or truncate history), and enable prompt caching so the
      repeated prefix is billed at the cache rate. OWASP attributes ~62%
      of agent bills to re-sent context.
  WARN     retry_storm  [key: prod-batch]
    prod-batch had 42 failed calls out of 56 (75%) in the last 15 min.
Runtime dependencies
0
Standard library only
Added latency
4.6ms
Worst case, local upstream
Throughput
4.1k/s
vs 10.6k/s with no proxy in the path
Test suite
214
No network, no API keys

Where the money actually goes

A budget dashboard tells you what you already spent. The failure that produces five-figure bills is a single agent chain looping, retrying or fanning out — and a daily cap fires hours after the money is gone.

Your app
Agent / SDK
Unchanged. Point the base URL at the gateway.
llm-guard
Meter, attribute, cap
Reads usage from the response, prices it per model, enforces the budget, aborts a stream mid-flight.
Provider
OpenAI / Anthropic
TLS terminated in the gateway using the standard library.
You get
Cost per key
SQLite, a self-contained dashboard, alerts before the ceiling.

Why it is built this way

Three decisions that shape everything else, including what the tool deliberately does not do.

0

No dependencies

A gateway sits on the request path and holds your API keys. No FastAPI, no httpx, no uvicorn — about 7,900 lines you can read in an afternoon, and a surface small enough to get past a security review.

?

Unknown over wrong

Cache reads cost between 2.5% and 50% of input depending on the model. Treating that as one constant is a silent four-to-five-times error. If a model has no price on file, the cost is recorded as NULL and reported as unpriced. A plausibly wrong number is worse than an obvious gap.

!

Advisory, not destructive

Detectors never block traffic. A detector that silently kills production is worse than the bill it prevents. The one exception is the optional per-stream cap, and the number it acts on is a local estimate, never used for billing.

What it catches

The failure patterns are not our invention. They come from OWASP AISVS C9.1 and a published catalogue of 63 confirmed production incidents across 21 orchestration frameworks. Each detection cites the pattern it matched, the number of recorded incidents, and the largest reported loss.

PatternObservable signalRecorded incidents
Unbounded context loopInput/output ratio above 30:111
Delegation fan-out raceConcurrent burst across many keys11
Retry stormFailure rate above 5%9
Premium model as defaultOne model at 8x the average cost per call—
Uncached repeated prefixCache hit rate below 35%—
Velocity spike without a ceilingPeak day at 8x the median day—
No cost attributionAll spend on a single key—
Unpriced trafficSuccessful calls with no rate on file—

Run llm-guard patterns to see the signals, thresholds and citations in full.

Start in one command

No API key, no network, no signup. The demo data ships with a real 74:1 loop so the detector has something to find on the first run.

bash
$ pip install git+https://github.com/leyao-daily/llm-guard.git
$ python3 -m llmguard seed --reset --compare-days 30
Seeded 13,254 requests spanning 60 days ($227.10 of simulated spend).
  Last 30 days is the reporting window; the preceding 30 days gives the
  period-over-period delta.
$ python3 -m llmguard anomalies
$ python3 -m llmguard diagnose --client "Acme Corp" --out report.html

If your AI bill is unpredictable

We find out why. A fixed-price diagnosis delivered as a written report: what is driving it, what to change first, and what each change is worth per month. No infrastructure to deploy, read-only access to usage data, and no prompts or completions ever leave your account.

$1,500 fixed 5 working days Refunded if nothing is actionable

If you would rather run it yourself

llm-guard is MIT licensed and complete. Zero runtime dependencies, standard library only, budgets and runaway-loop detection included. There is no crippled edition and no feature held back for paying customers.

MIT licensed 15 CLI commands Air-gap friendly

We are a small team in Hong Kong and would rather talk to people running this in production than ship a newsletter. hello@lye-labs.com