Three divisions.
One standard of evidence.
The benchmarking is not marketing for the consulting — it is the reason the consulting is any good. Pick a division to see what the work actually involves.
Production AI for messy, regulated work
The systems that remove the repetitive reading, typing and checking filling your team's week. Built for real conditions — bad scans, inconsistent layouts, personal data, and an auditor who will eventually ask how it decided.
What we build
- Document automationReads scanned PDFs and correspondence, classifies them, pulls out the fields you need, and checks its own output against a fixed format.
- Meaning-based searchFinds the right record by what it means, not by keyword — across documents nobody has tagged.
- Timeline reconstructionAssembles a defensible chronology from fragmented records that arrived in no order.
- Personal data handlingDetects and consistently masks personal details across a whole bundle, with an audit trail.
What you get
- Accuracy you can readMeasured against a held-out sample, not asserted. You see what it got right and what it referred to a human.
- A model that can be swappedRankings move every few months. Nothing is welded to one provider.
- Your team on the toolingHandover from the first sprint, not a black box you rent forever.
- The numbers for a board packVolume, accuracy, hours returned, cost per item.
We measure the models in the open
Every month, the major AI models run identical real tasks and the results are published in full — including the ones that make popular models look bad. No sponsored verdicts, ever.
What gets published
- The Prametriq IndexScores across reasoning, document reading, value for money and speed.
- The methodThe task set, the prompts, the scoring. Disagree with a reading and you can rerun it.
- The raw outputsNot just the score — what the model actually produced.
- Plain-English teardownsWhat changed in each release, and whether it matters to anyone outside a research lab.
Why it exists
- Model choice is a decision, not a preferenceWhen we recommend one for your job, it is a measurement we can show you.
- Nothing is tuned per modelSame prompts, same schema, same retry policy. Tuning per model is how leaderboards lie.
- Rankings move constantlyWhat was right six months ago frequently is not now.
- Nobody pays for placementThere is no commercial relationship with any model provider.
Websites that are measured too
The same habit applied to what your customers actually see: fast, accessible, and measurable, with AI tooling built in rather than bolted on afterwards.
What we build
- WebsitesFast, accessible, and readable on a phone in daylight. Scored, not guessed at.
- Content systemsSo the site does not go stale the week after launch.
- Social managementDelivered with the AI tooling we benchmark, with analytics that mean something.
- Accessibility workTreated as bugs, because that is what they are.
How it differs
- Built by the engineerNot templated by someone who has never had to maintain it.
- You own itStandard tooling, your repository, no proprietary lock-in.
- Performance is a numberLoad time and accessibility are measured before and after.
- A landing pad for smaller workWhere a full AI engagement is not justified yet.
Not sure which one you need?
Most people are not. Describe what is slow and we will tell you which — or that none of it applies yet.
Tell us what is slow