01 Services

Three divisions.
One standard of evidence.

The benchmarking is not marketing for the consulting — it is the reason the consulting is any good. Pick a division to see what the work actually involves.

Division 02 · The main event

Production AI for messy, regulated work

The systems that remove the repetitive reading, typing and checking filling your team's week. Built for real conditions — bad scans, inconsistent layouts, personal data, and an auditor who will eventually ask how it decided.

What we build

  • Document automationReads scanned PDFs and correspondence, classifies them, pulls out the fields you need, and checks its own output against a fixed format.
  • Meaning-based searchFinds the right record by what it means, not by keyword — across documents nobody has tagged.
  • Timeline reconstructionAssembles a defensible chronology from fragmented records that arrived in no order.
  • Personal data handlingDetects and consistently masks personal details across a whole bundle, with an audit trail.

What you get

  • Accuracy you can readMeasured against a held-out sample, not asserted. You see what it got right and what it referred to a human.
  • A model that can be swappedRankings move every few months. Nothing is welded to one provider.
  • Your team on the toolingHandover from the first sprint, not a black box you rent forever.
  • The numbers for a board packVolume, accuracy, hours returned, cost per item.
The shape of an engagement
Week 1 — DiscoveryA week inside your process, following the paperwork. You get the jobs ranked by what they actually cost in hours and salary, and a written business case. Useful even if you stop here.
Week 2 — MeasurementFour to six models run the same task set against your documents, including the bad scans. You see the scoreboard and the cost per document.
Weeks 3–4 — PilotThe smallest thing that proves it, on real data, that your team can put their hands on. Checking designed in from the start.
ThenEither the numbers justify going further, or they do not. If they do not, we will be the ones to say so.
FeesDiscovery is a small fixed fee. A pilot is fixed-scope. Ongoing support is monthly. We will not quote a figure before seeing the process — that is guesswork dressed as a price list.
Division 01 · The public one

We measure the models in the open

Every month, the major AI models run identical real tasks and the results are published in full — including the ones that make popular models look bad. No sponsored verdicts, ever.

What gets published

  • The Prametriq IndexScores across reasoning, document reading, value for money and speed.
  • The methodThe task set, the prompts, the scoring. Disagree with a reading and you can rerun it.
  • The raw outputsNot just the score — what the model actually produced.
  • Plain-English teardownsWhat changed in each release, and whether it matters to anyone outside a research lab.

Why it exists

  • Model choice is a decision, not a preferenceWhen we recommend one for your job, it is a measurement we can show you.
  • Nothing is tuned per modelSame prompts, same schema, same retry policy. Tuning per model is how leaderboards lie.
  • Rankings move constantlyWhat was right six months ago frequently is not now.
  • Nobody pays for placementThere is no commercial relationship with any model provider.
Status
Current stateThe site shows illustrative placeholder figures, clearly labelled. The first live measurements publish at launch, with the harness and raw outputs open-sourced alongside.
Cost to youFree, and always will be. It is published work, not a product.
Division 03 · The front of the house

Websites that are measured too

The same habit applied to what your customers actually see: fast, accessible, and measurable, with AI tooling built in rather than bolted on afterwards.

What we build

  • WebsitesFast, accessible, and readable on a phone in daylight. Scored, not guessed at.
  • Content systemsSo the site does not go stale the week after launch.
  • Social managementDelivered with the AI tooling we benchmark, with analytics that mean something.
  • Accessibility workTreated as bugs, because that is what they are.

How it differs

  • Built by the engineerNot templated by someone who has never had to maintain it.
  • You own itStandard tooling, your repository, no proprietary lock-in.
  • Performance is a numberLoad time and accessibility are measured before and after.
  • A landing pad for smaller workWhere a full AI engagement is not justified yet.
The shape of an engagement
BuildFixed scope, fixed fee, agreed before anything starts.
OngoingMonthly, if you want content and social handled. Cancel whenever — no minimum term.

Not sure which one you need?

Most people are not. Describe what is slow and we will tell you which — or that none of it applies yet.

Tell us what is slow