£8,000 fixed fee
A two-week fixed-fee diagnostic for teams running LLM features in production. Eval coverage, CI integration, production sampling, alerting and calibration — scored, evidenced, and ranked by severity. No discovery dance.
What you get
Scope fixed in writing before kickoff — five audit areas, the report, the exec summary, the review call. If the work runs long, that's on us.
Week one: kickoff, read-only access, system map. Week two: analysis, draft, internal review.
Findings ranked by severity — failure mode named, evidence attached, fix sized in engineer-days.
Risk level, top three findings, next step, rough cost. No code, no eval jargon — the version that travels to a board.
The engineer who wrote the report walks you through it. Not an account manager.
How it runs
The same four stages run every engagement. The service decides the scale, not the method.
01
Planner
NDA signed, read-only access granted, system map drawn before analysis starts.
02
Hub
Golden set, CI integration, production sampling, alerting and calibration — each scored against evidence.
Pricing
Two-week diagnostic, 30-page report, one-page exec summary, 60-minute review call. 50% on signature, 50% on delivery.
£8,000Audit findings become the scope — no second discovery. Four to eight weeks, fixed-fee, same engineer leads.
£15k–£30kScore review, swap-risk runs, quarterly golden-set refresh. Same engineer.
From £3,500/moProof in production
Categorisation, summarisation, and triage automated across an operations team's daily workload. Graceful degradation when the…
View projectA documentation platform that reads code and emits prose a non-engineer can act on. Used internally on every Canarlo build.
The 7 mistakes businesses make with AI — buying tech before naming the problem, leaking client…
Read guideRAG, fine-tuning and prompting explained in plain business terms — what each does, when to use it,…
Read guideQuestions
Yes. The methodology is provider-agnostic. Anthropic, Google, open-weight, hosted, or self-hosted — the audit reads golden sets, CI configs, sampling logic, alert routing.
Read-only. CI logs, prompt registry, existing eval configs, a sample of production traces. No direct database access, no source-code commits. NDA signed before kickoff.
Four thousand on signature, four thousand on report delivery. No day-rate clock. Scope fixed in writing — five audit areas, the report, the exec summary, the review call.
What we look at
01
Golden set audit
Breadth, labelling consistency, staleness, size against system surface.
02
CI integration
Do evals run on every PR, or just on a laptop before release?
03
Production sampling
Real inputs feeding back into the suite, or a golden set frozen for months.
04
Alerting
Score regression, cost spike, latency cliff — routed to a channel nobody mutes.
05
Calibration
Does reported confidence match real accuracy? Without it, human-in-the-loop routes on a lie.
Talk to us
Scoped and priced before anything is built.
03
Forge
Findings ranked by severity, fixes sized in engineer-days, raw eval output in the appendix.
04
CanOS
DIY fix, scoped remediation, or an ongoing retainer to keep the harness alive — your call.
When is AI the wrong tool? Six honest signals AI is the wrong answer for a UK business task — and…
Common starting point — its absence is usually the top finding. We sample production traces, cluster by intent, and bootstrap a candidate set as part of the audit.
Three paths, your call. Most buyers fix the gaps themselves. If you want us to do the work, audit findings become the scope — four to eight weeks, fixed-fee. Or a monthly retainer from £3,500.