LLM Cost Diagnosis

We find out where a company's AI spend is going, and what to change. Fixed price, fixed scope, no infrastructure to deploy.


The problem we see most often

An agent workload starts re-sending its entire conversation history on every step. Input tokens grow while output stays flat. Nothing crashes, nothing alerts, and the first signal is a bill.

OWASP's research into agent execution budgets catalogues 63 confirmed production overruns across 21 orchestration frameworks, and puts Fortune 500 leakage at roughly $400M in unbudgeted spend. It also records that runtime visibility exists in only about a fifth of organisations.

The recorded failure patterns are not mysterious. They are just not visible until somebody looks at the token counts.

How a diagnosis works

No deployment. No agents in your VPC. No code changes.

1. You give us read access, not infrastructure

Two ways, whichever you prefer:

Either way it is read-only. We do not ask for a write key, and we do not need access to prompts or completions.

2. We analyse metadata, never content

Token counts, model ids, timestamps, attribution labels, and what the provider actually billed. That is the whole input.

We do not read your prompts. We do not read your completions. We could not tell you what your application does if you asked.

3. You get a written report

A short document, A4, the kind of thing you can forward to a CTO without rewriting it. It contains:

What it costs

PriceWhat you get
Scoping callFree, 15 minutesWe confirm the data is available and tell you honestly whether a diagnosis is worth doing yet. Sometimes the answer is no.
Diagnosis$1,500The written report, delivered within 5 working days of receiving data.
Deployment and monitoring$2,000 setup, then $1,000/monthWe install guardrails in your infrastructure, wire up alerting, keep the price table current, and answer when something looks wrong.
Ongoing cost governanceFrom $2,500/monthWe keep watching: routing changes, cache policy, context budgets, a quarterly review. The alternative is hiring for it.

If the diagnosis finds nothing worth acting on, we refund it. We would rather lose the fee than defend a report that says "your spend is fine but here are some charts".

That is not a marketing promise. It is a consequence of the data model: we can see before we quote whether there is anything in there, because the scoping call already showed us the shape of the usage.

What we do not do

Why the tooling is free

llm-guard is MIT licensed and complete. Zero runtime dependencies, runs on the standard library, enforces budgets and detects runaway loops. If your team wants to run it themselves, do that. It is genuinely the whole thing, not a crippled edition.

We charge for the part a tool cannot do: knowing which patterns matter, calibrated against what we have seen, and being accountable for the answer.

Get started

Email hello@lye-labs.com with:

  1. Roughly what you spend per month on model APIs.
  2. Which providers.
  3. Whether you have had a cost surprise in the last year. If yes, roughly how much.

That is enough for a scoping call. If it looks like there is something to find, we will say so. If it does not, we will say that too.


LYE LABS LIMITED (灵野科技有限公司) · Hong Kong