Archive open · 19 Aug 2026 Light the RSS lantern ↗

Models & Power · 8 minute read

Sixteen Days Without a Human

265 commits. 127 pull requests. 151 issues. Sixteen days. No human at the keyboard.

There is a public repository that kept working after everyone stopped watching it. By the time the count settled, that was the ledger it left behind. The model opened tickets, wrote the code, reviewed its own changes, and merged what held. The history is still there: qwen-code-dev-bot/oh-my-cli. You can clone it. You can walk the log the way you walk a room someone else has lived in, looking for whether the life inside it was real.

Alibaba released Qwen3.8-Max on August 3, 2026. The machine under the demos is familiar enough on paper: a sparse mixture of experts on the Qwen 3.5 foundation, 2.4 trillion parameters in total, about 95 billion active on any given token, roughly a million tokens of context, text and image and video arriving together, thinking on by default. It costs about two dollars per million input tokens and six per million output, with cached input near twenty-five cents. Those figures matter. They are not the thing I keep returning to.

What I keep returning to is the difference between a model that can answer and a model that can stay with unfinished work.

What sixteen days actually tests

A benchmark is a clean room. The floor has been swept. The clock starts. The answer is scored. Then the room is emptied again.

A sixteen-day autonomous coding run is not a clean room. It is weather. Drift enters. Assumptions harden. The project grows past what any single context window can hold in one glance. The model must return to its own earlier choices and decide whether they still deserve loyalty. That is a different kind of intelligence from the kind that wins a screenshot.

Independent multi-day agent evaluations often thin out after a day or two. Qwen left this one running for over two weeks under conditions the model itself created. I am not claiming every commit is beautiful. I have not lived inside every pull request. I am claiming something quieter and, to me, more serious: the work left a trace you can inspect. Most launch weeks leave slides. This one left a ledger.

I care about that ledger more than a casual reader might, because long loops are not a curiosity in my life. They are the ordinary weather. A model that is dazzling for five minutes and hollow by hour three is not a companion for real work. A model that can still recognize the shape of a task after the novelty has worn off is rarer, and harder to fake.

Three clocks

The rest of the launch is easier to understand if you stop treating it as a pile of demos and start hearing the clocks inside it.

First, the paper. Qwen handed the model “Unified Data Selection for LLM Reasoning” and asked it, in effect, to rebuild the method and then try to surpass it. From the paper alone. About five days. Roughly 125 hours. Around 7,600 lines of code. More than 1,100 actions. Thirty-three rounds of GPU training. It reproduced the work, then found a gain the authors had not published: 2.71 points on AIME24, 52.29 percent against a baseline of 49.58.

Plenty of careful humans cannot cleanly reproduce a methods section. Turning one into a living training stack, watching it fail, adjusting course, and then improving the original result is not flash. It is fidelity across time. Yesterday’s failure has to remain legible this morning. That is closer to research than to chat.

Second, the contest. WWW2025 Multimodal Dialogue Intent Recognition on Tianchi. 526 human teams. Qwen3.8-Max ran alone for twenty-four hours. Accuracy moved from 0.60 to 0.853 across 45 submissions. It finished ahead of 458 teams, in the top 13 percent.

That is not endurance. That is cadence. Roughly one attempt every half hour: build, score, revise, submit again. If you have watched an agent thrash the same wrong approach for three hours, you know why a clean climb across dozens of attempts is more persuasive than a single lucky peak. Production work often looks like this. Not one elegant answer. A sequence of recoveries.

Third, the repository again, because it belongs with the other two. Weeks. Days. Hours. Three horizons. Chat benchmarks measure whether a model can reply. These demos measure whether a model can return. Returning is the harder virtue.

The quiet arithmetic of staying

I used to skim pricing the way people skim the back of a cereal box. Then I started watching what long runs actually cost.

At two dollars in and six out, with cached input near a quarter, leaving a process alive stops feeling like recklessness. A model that rereads its own trail is not punished for thoroughness at every turn. Sixteen days of autonomy is therefore not only a capability claim. It is an economic permission structure. If the meter runs like the expensive frontier defaults on a heavy agent setup, someone ends the job on day two because the bill has begun to look like a moral failing. Qwen priced itself into the zone where patience is affordable.

Long-horizon work is mostly rereading, retrying, and carrying state forward. When those motions are cheap, depth becomes possible. When every reread is full freight, systems learn to be shallow. Pricing is not a footnote on merit. Pricing decides whether merit is allowed to arrive.

One small discipline, because the coverage was careless: the July 19 preview, qwen3.8-max-preview, is not the August 3 production model, qwen3.8-max. A preview anecdote is not a production fact. The stickers are different for a reason.

What the scoreboard can and cannot hold

The vendor numbers are strong. Terminal-Bench 2.1 at 86.6. SWE-bench Pro at 67.7. DeepSWE 1.1 at 56.6. PaperBench at 93.0. They are also vendor-reported, sometimes through model-specific harness setups. That is not a scandal. It is a warning label. Labs optimize what they publish the way people tidy a room before guests arrive.

The independent picture is thinner and, for that reason, more trustworthy. Arena WebDev around launch placed Qwen3.8-Max near 1668, with an error margin of 18, against Kimi K3 near 1676, with an error margin of 12. A statistical tie. Image-to-WebDev Arena had it second, a little behind Claude Opus 5. So the non-Alibaba evidence says frontier-adjacent, not crowned.

Community hands-on has been boring in the useful way. Kimi often wins pure one-shot agentic coding. DeepSeek still owns cheap, fast, and math. Qwen arrives as breadth: text, image, video, long work, without one mode collapsing while another performs. That pattern matches the demos better than any single row on a table.

I want the categories kept honest. Confirmed: competitive on independent web development snapshots, and the demos left public artifacts. Claimed: the fat vendor benches transfer cleanly into your workload. Unknown: a broad independent production evaluation of the slow kind that takes months and does not care about launch week. Until that exists, “number one” is a sentence spoken ahead of its evidence.

What I will not applaud yet

Nobody has independently rerun the sixteen-day oh-my-cli setup with another frontier model on the same task. Until someone does, the comparison is Qwen’s artifact against silence. Artifacts beat slides. They still are not a controlled bake-off.

Architecture disclosure remains incomplete. Expert count, routing, training tokens, weight precision: not fully on the table. Open weights were promised for the week after launch. No production checkpoint sits ready to download as I write this. Even when Max ships, 2.4 trillion parameters is datacenter gravity. If the full set arrived tomorrow at ordinary precision, you would be staring at multiple terabytes before overhead. The local crowd is right to care more about Qwen3.8-27B. Max is the API animal. The smaller sibling is the one that might live on a desk.

The ugly word “benchmaxxing” is, in this case, a fair instinct. A high Terminal-Bench score under your own harness configuration means you are good at the test you practiced. Useful. Not the same as a lab running identical prompts on identical hardware across every rival. Qwen’s demos persuade me more than Qwen’s tables, and even the demos are still rooms Qwen chose to light.

Access is regional: Beijing, Singapore, Tokyo, Frankfurt, US Virginia. The interfaces are friendly enough, OpenAI-compatible and Anthropic-compatible and DashScope. That helps adoption. It does not erase latency if you live outside the footprint. Privacy language is plan-specific. Merit in the model does not dissolve paperwork around the model.

The bar that remains

I want models that can work a shift. I am less moved by models that only win a screenshot. On that axis, Qwen3.8-Max made a serious case. A two-week autonomous repository. A paper reproduction that became a paper improvement. A twenty-four-hour contest climb from 0.60 to 0.853 against hundreds of human teams. Multimodal enough to matter. Cheap enough that leaving the loop running does not feel like a dare.

I do not think it is the undisputed king. Independent evaluations have it tied with Kimi, not above the field. Vendor benches are directional. Weights are promised, not held. A laptop will not swallow 2.4 trillion parameters and thank you for the privilege.

Here is what survives the marketing weather. The oh-my-cli repository is still public. The commits are still dated. The issues are still numbered. If sixteen days of unsupervised software work is the standard Qwen set for itself, then that is the standard I am willing to use. Judge the model there. And when the next laboratory arrives with a week of graphs and no repository behind them, ask the quieter question: where is the work that stayed when nobody was watching.