The frontier–local gap, in numbers
How far behind is the best model you can run in your own office, and is that gap growing or shrinking? The evidence disagrees with itself, so both sides are here.
Zavier Taylor
The one question this answers
Everything else rests on this. If the model you can run locally is hopeless, the whole idea fails and you should stop reading. If it is close enough on the work that matters, the idea holds.
So rather than assert it, here is the evidence, including the parts that cut against me.
The two open frontiers people confuse
There is not one open-source frontier. There are two, and conflating them is where most arguments about this go wrong.
The first is what you can host in an office: models that fit in roughly 128GB of memory on a quiet desktop machine. That is the honest ceiling for on-premise AI.
The second is what exists as open weights: near-frontier models of several hundred billion parameters that are genuinely open, and genuinely need a datacentre.
A private box lives in the first category. So the relevant question is never "how good is the best open model". It is how good is the best open model that fits on the hardware I would actually put in your building. Those are very different numbers, and vendors quote the flattering one.
Where local reaches parity, and where it does not
The comparison worth making is not a single leaderboard score. It is per task, because the answer changes completely depending on what you are asking for.
| Local reaches parity | The frontier is still clearly ahead |
|---|---|
| Summarising documents | Long-context reasoning, 100k tokens and beyond |
| Extracting structured data | Long-horizon, multi-step autonomous work |
| Answering questions about your files | The hardest reasoning, and multi-file coding |
| Drafting in a set style | |
| Classification and translation |
The load-bearing fact is that the left column is roughly eighty percent of what a professional office actually asks of these systems. That is why a private box is viable at all. Not because it matches the frontier everywhere, but because it matches it on the work that fills the day.
The corollary matters just as much: if your bottleneck sits in the right-hand column, a box is the wrong purchase and I will say so.
Is the gap closing or widening? The honest answer is that it is contested
Two credible trackers reach different conclusions, and presenting only one of them would be sales copy.
Narrowing, or at least stable. One major aggregator has tracked a consistent three to six month gap for over eighteen months, with the head-to-head Arena gap roughly halving year on year. (OpenRouter, open-weight models analysis)
Widening slightly. An independent tracker finds open models lagging by around four months, or eight index points, which is slightly larger than the year before. (Epoch AI, the open/closed gap)
The cleanest way to reconcile them: static benchmark scores are narrowing, while real-world agentic capability, the long multi-step work where a model has to not drift, is widening. Both measurements are correct. They are measuring different things.
Which is precisely why the defensible architecture is local for the bounded majority, with an optional and controlled bridge to a frontier model for the hard remainder. Not local replaces everything. Anyone promising you that is selling.
The number nobody quotes, and the one you should demand
Benchmark scores are the wrong question for a machine in your office anyway. The right one is measured tokens per second, on the specific model you will be given, on the specific hardware you are buying.
A dense 70-billion-parameter model will load onto a machine that cannot usefully run it, and produce text more slowly than you can read. It fits, and it is useless. The reason is memory bandwidth rather than capacity, and no specification sheet will tell you. What actually runs on the box covers how that trap is set, and which models actually fit has the measurements.
Why publish the numbers that do not flatter the pitch
Because the buyer is a sceptic, and should be.
A vendor who shows only the winning figures is concealing the losing ones. Showing the whole board, including where the cloud is comfortably ahead, is what lets you make the right call, and sometimes the right call is that you do not need me. A claim you can trace to its source is worth more than one you have to take on trust.
If the evidence ever turns against the case for on-premise AI, if the gap stops closing on the work that matters, this page will say so. That is not a principle I expect to be tested soon, but it is the reason the citations are here rather than a summary of them.
Two other pieces bear on the same decision from different angles: why keep it in-house makes the confidentiality argument, and falling prices, rising bills deals with the cost objection, which is the one most people raise first and is mostly wrong in an interesting way.