How it works
What actually runs on the machine.
One repeatable system, installed the same way each time, then pointed at your data. No part of it is exotic. The value is that it's assembled, maintained and tuned for your firm rather than left as a pile of parts.
The layers
Private chat
A familiar chat interface for staff, with accounts, roles and per-person history. It can sign in with your existing work logins rather than another password to manage.
Retrieval over your documents
Your file share is parsed, indexed and searchable, so answers come from your own reports and correspondence, and cite them. This is what firms usually mean when they say they want AI that “knows our work”.
Automation
Scheduled and triggered jobs: email triage, drafting, file routing, overnight processing. Because the machine is yours, running work around the clock costs electricity rather than per-token fees.
Your own tools
Bespoke applications running on the same box, pointed at the local model. Anything already built against a standard model API moves across by changing one setting.
What arrives, in order



Which machine
Five candidates, from the cheapest thing worth buying to the one almost nobody needs. Pick one and it tells you what it costs landed in New Zealand, what it runs, and how fast, on the same models, so the numbers are comparable rather than cherry-picked.

Mac mini M6
32 GB unified memory
- Memory
- 32 GB unified
- Bandwidth
- 170 GB/s
- gpt-oss-120bdoes not fit
The workhorse. Mixture-of-experts, ~63 GB, roughly o4-mini class.
- Qwen3-30B-A3Bnot measured
The fast one. ~18 GB, for extraction and answering from documents.
- A dense 70Bdoes not fit
The trap. Loads on almost everything here and is useless on most of it.
The cheapest honest starting point, and for a lot of firms the right one. It runs the small and mid-sized models well and will not touch a 120B. Buy this if the work is drafting, summarising and search over your own documents, and you would rather find that out for under two thousand dollars than eleven.
Snapshot · 2026-07-29 · published figures, not measured by me
What leaves, and when
A local model handles ordinary office work well. On the hardest questions a frontier model is still better, and pretending otherwise would not survive your first difficult afternoon. So the question is not whether to use one. It is who decides, and when. The machine ships sealed; opening a route out is a deliberate change, scoped with you, and only two shapes of it make sense.
Sealed
The default
Everything runs on the machine and nothing goes out. There is no outbound path to any AI provider. Not a setting, an absence. It is the mode every install ships in, and the one to stay in for anything privileged. If your compliance officer wants proof rather than a promise, the firewall rule is the proof, and you are welcome to watch it being written.
Redacted hybrid
Narrow, and caveated
Names, addresses and identifiers are stripped before a prompt leaves, the frontier model reasons over placeholders, and the answer is reassembled here. The honest limit: the best available redaction misses roughly three names in ten. That makes it usable for incidental personal detail in an otherwise general question, and unusable for privileged client matter. I will not sell it as more than that.
Open hybrid
Where the work isn't privileged
Staff choose per message: keep this one local, or send this one out to a frontier model. No automatic routing on difficulty, because difficulty and sensitivity are the same axis and the machine cannot tell them apart. The person who wrote the question can. Sensible for engineering and survey work, where the constraint is expertise rather than confidentiality.
Two things hold in either hybrid mode, or I would not build it. The cloud account runs on terms that forbid training on your content, and every call that leaves is logged where you can read it: what went, when, and who sent it. A promise you cannot audit is not a control.
What actually fits, today
A new open model lands most weeks, so this is rebuilt from the Hugging Face and OpenRouter APIs rather than written down once. Pick a machine and it tells you what will physically hold, and where I have an opinion, I have said so.
26 GB usable for weights, after the operating system and the runtime take theirs.
gpt-oss-120b
116.8B · 75.4 GBmeasuredToo largeopenai/gpt-oss-120bapache-2.0Published 2025-08-045.3M downloads131k contextMixture-of-experts · 5.8B active per tokenRunning thisIn use on a machine I look after.
The current default, and the one every price on this site was built around. Mixture-of-experts, so it is far faster than its 117B suggests. Benchmarks roughly a generation behind frontier and comfortably ahead of what drafting and document questions need.
gpt-oss-20b
20.9B · 13.9 GBmeasured~45 tok/sestimatedFitsopenai/gpt-oss-20bapache-2.0Published 2025-08-046.5M downloads131k contextMixture-of-experts · 4.2B active per tokenWould run thisI'd put this in an office tomorrow.
What I reach for when responsiveness matters more than depth: extraction, classification, answering from a document that is already in front of it. Leaves enough memory free to run something else beside it.
DeepSeek-V4-Pro
1598.8B · 959.3 GBestimatedToo largedeepseek-ai/DeepSeek-V4-PromitPublished 2026-04-22809k downloads1.0M contextMixture-of-experts · 75.6B active per tokenI'd argue againstFits, and still the wrong choice.
The strongest open model on this list and it needs roughly 960 GB at 4-bit, which is nine of the machine I recommend. Worth knowing about so you can say why it is not on the table, rather than being surprised by it in a meeting.
Open weights and runs on your machine are different claims, and the gap between them is the whole point of this table. Where Unsloth publish a 4-bit build, the figure is the size of the file you would actually download; everywhere else it is calculated from the parameter count. Both add 20% for the cache and runtime. Nothing here estimates speed. That has to be measured on the machine, and a mixture-of-experts model will beat a dense one of the same size by a distance.
Rebuilt 2026-08-31 fromHugging Face Hub APIOpenRouterOllama libraryUnsloth GGUFsllmfit (speed formula, MIT)LobeHub icons
How far behind is it, really
The best model that fits one of these machines, against the best that exists, benchmark by benchmark.
See the measured gapEvery model, side by side
The three above are the shortlist. The full database is public, filterable, and rebuilt from the source APIs rather than written down once.
Open the full databaseThe rated index
Several hundred models, filterable to what fits a 128 GB box. Licensed for internal use, so it sits behind a client sign-in rather than on the public site.
Open the index · clients