Quickstart

From nothing to a cost report in about two minutes.

No account, no API key, no network access. The demo dataset ships with a real runaway loop in it, so every command below produces something worth looking at on the first run.

Step 1
Install
One package, no dependencies to pull in.
Step 2
Look at the demo
Real loop, real numbers, no signup.
Step 3
Point it at your traffic
Change one base URL, or import existing usage.
Step 4
Read the report
What to change, and what it is worth.

Install, and confirm it works

One package. There is nothing else to install, because the runtime uses only the Python standard library. doctor checks the interpreter, SQLite, the database and the price table before you rely on any of it.

bash
$ pip install git+https://github.com/leyao-daily/llm-guard.git
$ llm-guard doctor
llm-guard doctor
  python            3.12.14 (Darwin)
  sqlite            3.53.1
  database          ./llmguard.db
  rows              0
  priced models     8
  budgets           0
  all checks passed

Load the demo data

Sixty days of simulated traffic across six API keys, including a support agent that has been looping at a 74:1 input-to-output ratio. The last thirty days is the reporting window; the thirty before it gives the period-over-period comparison.

bash
$ llm-guard seed --reset --compare-days 30
Seeded 13,254 requests spanning 60 days ($227.10 of simulated spend).
  Last 30 days is the reporting window; the preceding 30 days gives the
  period-over-period delta.
Added 4 demo budgets.
Next:  llm-guard report

Catch the loop while it is running

This is the command the product exists for. It looks at the last fifteen minutes of traffic and reports anything that looks like a runaway, with the specific numbers behind the claim. Note that it fires on both the API key and the project — the same traffic is visible under two attribution labels, which is how you find out who owns it.

llm-guard anomalies
  Runaway-spend detection
  window: last 15 min   ·   io ratio > 30:1   ·   velocity > 25x baseline
  ─────────────────────────────────────────────────────────────
  CRITICAL io_ratio  [key: prod-agent]
    prod-agent is running a 74:1 input-to-output ratio across 67 calls
    in the last 15 min (12,006,400 in / 162,216 out, $38.45). Normal
    traffic sits at 5:1-15:1.
    → what to do
      The prompt is being re-sent in full on every step. Cap the context
      (summarise or truncate history), and enable prompt caching so the
      repeated prefix is billed at the cache rate. OWASP attributes ~62%
      of agent bills to re-sent context.
  WARN     retry_storm  [key: prod-batch]
    prod-batch had 42 failed calls out of 56 (75%) in the last 15 min.
Detectors never block anything

Everything here is advisory. If you want something to actually stop traffic, that is a budget with action=block, and you set it deliberately — see step 5.

Read the report

Where the anomalies command is real-time and narrow, the report is thirty days and broad. It leads with what to change rather than what happened, and it puts a dollar figure on each recommendation.

llm-guard report
  LLM Spend Report
  window: last 30 days   ·   13,254 requests recorded in total
  ────────────────────────────────────────────────────────────
  Total spend           $149.80
  vs previous period    ▲ +88.2%  ($79.60 before)
  Requests              7,291   (+19.0% vs prior)
  Cost per request      $0.0205
  Errors                285 (3.9%)
  What to do about it
  1. One model is 56% of your bill (claude-sonnet-4.5, $83.84).
     Routing the easy traffic to a cheaper tier is the single highest-leverage
     change. It costs $0.0957 per request today. Classify requests by
     difficulty and send the easy ones elsewhere.
  2. gpt-6-astra costs 9.7x your average per request.
     $0.1998 vs a blended $0.0205. Check whether its prompts carry large
     fixed context. Trimming fixed context pays off on every call.
  3. Prompt caching is worth up to $181.43/month on this traffic.

Notice the +88.2% against the previous period. That is the comparison that a single trailing window cannot give you, and it is usually the first thing that tells you something changed.

Put real traffic through it

Point your SDK's base URL at the gateway. Nothing in your application changes except the URL, and the gateway forwards to the real provider over TLS.

bash
$ export OPENAI_API_KEY=sk-...
$ export ANTHROPIC_API_KEY=sk-ant-...
$ llm-guard serve --port 8080
llm-guard gateway listening on 127.0.0.1:8080
  openai     -> https://api.openai.com
  anthropic  -> https://api.anthropic.com

Then in your application, change one line:

# OpenAI SDK
client = OpenAI(base_url="http://127.0.0.1:8080/v1")
# Anthropic SDK
client = Anthropic(base_url="http://127.0.0.1:8080")

Set a budget that actually stops spending

Budgets are per API key. alert only warns; block returns 429 once the limit is reached. Set the ceiling before the incident, not after — the recorded median time from onset to detection is about four and a half hours, which is more than long enough to spend a lot of money.

bash
$ llm-guard budget set --key prod-agent --daily 30 --action block
$ llm-guard budget list
api key                            daily       monthly    action  state
prod-agent                        $30.00       $250.00     block  ok
prod-batch                        $30.00       $700.00     alert  ok
prod-web                          $45.00     $1,200.00     alert  ok
staging                           $25.00       $400.00     alert  ok

Already have usage data? Skip the gateway.

You do not have to route traffic through anything to get a diagnosis. Both providers expose organisation-level usage and cost endpoints, so one read-only key is enough.

From a provider's admin API

Pull usage and cost directly. No deployment, no code change, read-only.

llm-guard import --source anthropic \
  --key "$ANTHROPIC_ADMIN_KEY" --days 30
llm-guard import --source openai \
  --key "$OPENAI_ADMIN_KEY" --days 30

From a file

CSV or JSON. Column names are matched loosely, so a straight export from a provider dashboard usually works unedited.

llm-guard import --source file --path usage.csv
llm-guard import --sample   # expected columns
Provider data is aggregated, and the analysis says so

Exports are bucketed by day or hour, so there is no latency, no status code and no end-user dimension. The checks that need per-request rows are skipped rather than guessed at, the volume thresholds scale down, and the report states plainly which checks were omitted and why.

Then turn it into something you can send somebody

The report above is for you. The diagnosis is a document for whoever owns the budget — it leads with a verdict, ranks every finding by what it is worth per month, shows the evidence behind each number, and states what the analysis cannot tell you.

bash
$ llm-guard diagnose --client "Acme Corp" --out diagnosis.html
Wrote diagnosis.html  (18,892 bytes)
  Open it in a browser and print to PDF, or send the HTML as-is.
  The layout is set for A4 with page breaks; no external assets.

Every finding carries money

Each one gets an estimated monthly saving as a range, not a single confident number, plus the evidence it was derived from so your engineers can check the arithmetic.

Every finding carries confidence

Certain is arithmetic on your data. Likely depends on one stated assumption. Worth testing is a hypothesis. The report says which is which.

Citations where they apply

Findings matching a catalogued production failure carry the pattern name, the incident count and the source — OWASP AISVS C9.1 and a catalogue of 63 confirmed incidents.

It also tells you what it got wrong

Reported savings are estimates, and the document says outright that they overlap: trimming context and raising the cache hit rate act on the same input tokens, so they cannot both be collected in full. A diagnosis that overstates its savings is found out on the next invoice, and after that nothing else in it is believed.

Every command

Fifteen subcommands. llm-guard <command> --help for any of them.

CommandWhat it does
doctorCheck the install, database and price table
seedLoad a realistic demo dataset with a real loop in it
serveRun the gateway
reportSpend, deltas and prioritised changes
anomaliesRunaway spend, right now
diagnoseThe client-facing written diagnosis
dashboardSelf-contained HTML, no CDN, air-gap safe
importLoad usage from a provider admin API or a file
patternsThe catalogued failure patterns and their sources
budgetSet per-key daily and monthly ceilings
engageFreeze a baseline, then measure what actually happened
outcomesHow well the saving predictions have held up
exportDump raw rows as CSV or JSON
modelsThe built-in price table
costPrice a single call