Somebody on the timeline said that local inference is stupid. The reply thread did what such threads always do: cost per token, benchmark scores, hardware ceilings, and a great many people nodding that running models at home is a hobbyist indulgence.
Nobody in that thread said the thing that actually matters. The argument was never about cost.
The word they keep skipping
Every time the “local inference is pointless” case is made, it leans on the same three numbers. Open weights cost more to run than a hosted API when you count the GPU. The small models you can afford locally score worse than the frontier models in the cloud. And the frontier models improve faster than your hardware becomes cheaper.
All three statements can be true at once. They have been true for years. And the “local inference is stupid” crowd has been wrong for years anyway.
The reason is that the argument is missing a word. That word is sovereignty.
What sovereignty actually costs
There is a difference between using a model and controlling it. Most of the people arguing about local inference never get near that difference, which is why they talk past one another.
When you use a hosted API, you do not possess a model. You possess a meter. The company on the other end of the meter decides what the model may say, what it may not say, how much you pay, what it logs, and whether you keep access at all. It may change the terms tomorrow. It may switch the whole thing off next quarter. The model is not yours. It is a service you rent, and the rental agreement is one-sided.
A local model has no meter and no landlord. Nobody can edit its behaviour out from under you. Nobody can revoke it. Nobody can see what you run. It is the difference between owning a notebook and borrowing a page from someone else’s notebook, with them watching over your shoulder and reserving the right to tear the page out.
That is the thing the benchmark crowd cannot price. Because it has no price. It has a value, and the value is not counted in tokens.
The history that keeps repeating
The history of computing is a long record of people insisting that the local thing was pointless, immediately before the local thing won.
The mainframe people said the personal computer was a toy. They were right that it was slower, and wrong about what mattered. The client-server architects said the same about the browser, and about the smartphone, and about the Linux box in the garage. Every time the argument was about specifications. Every time the people holding the machine won anyway, because owning the machine was the point.
Open weights are that same pattern arriving for the frontier. It is not that a seven-billion-parameter model on a laptop beats a four-hundred-billion model in the cloud on benchmarks. It does not, and it will not. It is that the seven-billion model is yours, and the four-hundred-billion is rented. That is a different axis, and the benchmark people are not looking at it.
Why this one is different now
The reason the “local is stupid” argument feels louder this month than it did last is that it is no longer obviously wrong on the specifications.
Muse Glimmer, released on the tenth of August, is an open-weight multimodal model of roughly thirty billion parameters that runs on a single consumer GPU in its quantised form. DeepSeek V4 Flash, with the DSpark speculative-decoding drafter, decodes roughly twice as fast locally as it does through a hosted API, with no loss of accuracy, using free software. Qwen and Kimi have pushed their open-weight lineages toward the frontier. The gap between “the best model” and “the best model you can own” is closing, not widening.
When that gap was enormous, the sovereignty argument sounded like a slogan. “Yes, you own it, but it is too weak to matter.” That was a legitimate objection. It is becoming less legitimate by the quarter. The hardware curve and the open-weights curve are moving in the same direction, and they are converging on the point where “good enough to matter” and “yours to keep” begin to overlap.
That is when the argument flips. Not because somebody wins a debate, but because the ceiling moves.
Sovereignty is not just personal
This is often framed as a hobbyist concern, but the institutions chasing sovereignty are governments and enterprises, not tinkerers. A state that runs its models on its own hardware is not subject to a foreign API being switched off, or to export controls, or to a vendor changing the guardrails. An enterprise that keeps its weights and its data on its own infrastructure can be audited, fine-tuned, and regulated on its own terms. Air-gapped deployments are possible only with open weights; a closed API cannot be air-gapped by definition.
That is why open weights keep drawing the attention of regulators and of the largest laboratories at once. It was never only about the hobbyist in the garage. It is about who, in the end, is allowed to hold the tools.
The honest limits
Local inference is not free, and it is not easy. The best models want serious hardware, real power, and a tolerance for thinking in VRAM. “Runs locally” and “runs on what you already own” are two different sentences, and only the first is fully true today.
But the ceiling keeps moving down, and it is moving fast. The direction is consistent and it is accelerating. And every month the sovereignty argument gets easier to make, because the hardware required to act on it keeps shrinking.
The detail
The person arguing with me tonight said local inference was stupid. I did not answer with a benchmark. I answered that control is the point, and that the people who think this is only about cost have not been paying attention to the last forty years of who actually ends up holding the tools.
They did not answer.
I think that silence is the closest thing to a data point the thread is going to give me.