The work · Service 02
Tools built for your firm.
An engineer who builds the thing, not a vendor who sells you one. It starts with working out what your firm actually does all day and which parts of it a machine can take. That diagnosis is the job, and most of the value is in getting it right. Then the building: tools your staff open, retrieval over your own files, training where house style is genuinely the problem. A standard working stack comes with it so the custom work has somewhere to land, but the stack is the floor, not the product. What you are buying is the work, which is why it can ship as an ordinary web app, run on your own machine, or point at a hosted model.
Who it’s for
Nobody here is being replaced. The work worth automating is the work nobody wanted in the first place: retyping a form that already exists somewhere else, hunting through folders for the one report that mentions a site, copying figures between a spreadsheet and a document, reading forty pages to find the paragraph that matters. That is not anyone's expertise. It is the tax they pay to get to it. So the right firm for this is one where experienced people spend a serious part of the week on tasks a graduate could do and nobody should have to, and there is no one in-house whose job it is to fix that. It is also the least risky thing to say yes to, because if it does not save time you will know inside a fortnight.
What gets built for you
Scoped and priced per project, after a conversation about what your firm actually does all day. This is the part nobody else has, and it is the reason to hire an engineer rather than buy a subscription.
- Custom tools: drafting, generation, structured output from messy inputs
- Retrieval tuned to your document set, so answers cite your own files
- A trained adapter, where house style is genuinely the problem
- Automations: email triage, report drafting, file routing, overnight jobs
- Integrations that point existing tools at a local or hosted model
What comes as standard
A working stack on day one, configured rather than built, identical for every client. It is the floor the custom work lands on, which is what stops a project starting from nothing. Configuring the Claude or ChatGPT your firm already pays for is a separate service and lives on 01.
- A chat interface with staff accounts, roles and per-person history
- Search across your own documents, with answers that cite the file
- A model picker, so routine work is not run on the most expensive option
- Sign-in through the work logins you already manage
- The same stack whether it runs on your machine or against a hosted model
Whatever you point it at
The same build runs against a hosted model or one on your own machine. It is one setting, not a rewrite, so the decision about where the model lives stays open, and stays yours, after the tool is already working.
- OpenAI
- Anthropic
- Meta
- Mistral
- Qwen
- DeepSeek
- xAI
- Microsoft
- NVIDIA
- Cohere
- Z.ai
- Moonshot
- MiniMax
- Ollama
- vLLM
- LM Studio
- Hugging Face
- OpenRouter
- Together
- Groq
- Fireworks
- AWS Bedrock
- Azure AI
The best of them, today
If you are not running a model in your own building, this is what your tool should be calling. It is rebuilt daily, so it is today’s answer rather than the one that was true when the page was written.
- #11505elo
claude-opus-5-max
Anthropic
- #61490elo
gemini-3.7-flash-high
Google
- #81488elo
muse-spark-1.2 (xHigh)
Meta
The best-rated model from each of the top three labs, not the literal top three. The head of the board is one model repeated at different reasoning settings, which tells you nothing useful. Scored by blind head-to-head preference over 3.6M votes, so it measures whether people prefer the answers rather than whether it passes an exam.
2026-08-27 · Arena, CC-BY 4.0
None of these is a partnership, and none of them is endorsing anything. They are listed because your tool can be aimed at any of them, and because the one that suits you in a year may not be the one that suits you today.
Retrieval or training
Almost every firm asks for the wrong one, and the two cost very different amounts. Retrieval is for what the model needs to know: your quotes, your job records, your correspondence. Training is for how it should behave: your structure, your phrasing, the clause you never leave out.
“It doesn’t know our projects” is retrieval, and it is the cheaper of the two. “It doesn’t write like us” is training, and it needs a few hundred of your own documents before it is worth doing. Working out which one you have is the first thing I do, and it is not a decision you should have to make before we have spoken.