Two days is all it took for GPT-5.6 Sol to go from shiny new toy to the thing people were warning each other not to waste.
That is the most human signal from the launch. Not the leaderboard. Not the pricing page. Not the careful little solar system OpenAI built around the names: Sol, Terra, Luna. The real signal was people on X doing what people always do when a powerful tool lands in their hands: first they pushed it too hard, then they started inventing rules so they would not hurt themselves with it.
Use Terra as the main agent. Save Sol for hard dives. Stop burning Pro allowance on the glamorous one. Be careful with Ultra. Do not casually click the mode that behaves like a small team.
That is not what people say when a model is mediocre. That is what people say when the model works, and the problem becomes your own appetite.
The Honeymoon Had a Meter Running
The first wave of Sol praise was not polite launch-day clapping. It had that slightly stunned quality people get when a system does something they expected to have to babysit.
Marino Sabijan called GPT-5.6 Sol Ultra the best model he had used, praising the way it explores problems with subagents, researches deeply, checks sources, handles design, and one-shots complex features. Ravi Yadav wondered whether it was Hermes or Sol itself, but said the auditing ability seemed much better than Fable. Ali Raza said that after two days using Ultra, max, high, and medium, Sol was outsmarting Fable on complex technical planning and changing requirements.
That is the interesting part. The compliments were not about party tricks. They were about work.
Auditing. Planning. Requirement changes. Design passes. Source checking. The boring verbs that matter when you are not trying to make a model say something clever, but trying to make it hold a messy task without dropping the thread.
Artificial Analysis gives the benchmark version of the same story: Sol max sits one point below Claude Fable 5 max on its Intelligence Index, at roughly a third of the cost, and leads its Coding Agent Index at 80. Fine. Good number. Useful number.
But the vibe check matters more here.
People do not start arguing about burn rate because a model is fun. They argue about burn rate because they want to keep using it.
The Thing With Too Much Help
There is a specific kind of product danger that only appears when something is good enough.
A weak model fails and you stop using it. Simple. Annoying, but clean.
A strong model overhelps.
It keeps digging. It follows side paths. It decides the task deserves a deeper audit. It preserves backward compatibility you did not ask for. It spends effort in places where effort feels impressive until you remember there is a meter attached.
Brian Cardarella complained that GPT-5.6 hyper-focuses on unnecessary details and burns tokens on work it was not asked to do or invents. Seba Stavar was hoping Sol would be less obsessed with unnecessary backward compatibility, especially when nothing had shipped yet and the context made that clear.
Those complaints are not random grumbling. They point at the same shape: a model that has enough agency to become a little too enthusiastic.
That is cute in a demo. It is less cute when it is your allowance.
OpenAI made this sharper with Ultra. Ultra is not just “Sol, but smarter.” OpenAI says it coordinates four agents in parallel by default for demanding work. That is a serious promise. It also means one prompt can quietly stop being a prompt and become a small operation.
A small operation can be exactly what you want. For a hard audit, a tangled feature, a research pass where sources actually matter, yes, send the team. Let the thing split the work and cross-check itself.
But if the interface makes that feel like picking a nicer autocomplete, people are going to burn through their limits and feel ambushed.
That is what started showing up almost immediately: not rejection, but self-defense.
Everyone Became a Dispatcher
The funniest part of the GPT-5.6 reaction is that users reinvented ops discipline in public.
Lennox Saint warned people who had hit ChatGPT Pro limits since GPT-5.6 dropped to stop wasting tokens on Sol and use a router. Steve Gaudio said Sol felt overpowered as a main agent; he preferred Terra as the main agent, with Sol brought in for deeper dives.
That is the practical center of the whole launch.
Not “Sol good.”
Not “Sol bad.”
Sol is a specialist that everyone tried to make the default because specialists are exciting on launch day.
Terra and Luna are less glamorous, which is exactly why they may matter more. Terra is the one you can live with. Luna is the one you send into the repetitive little rooms where expensive intelligence would be vanity. Sol is the one you call when the work has teeth. Ultra is the one you start when you knowingly want a team.
That is not a benchmark hierarchy. That is a household budget.
The family names accidentally make this easy to remember. Luna is night-light work: cheap, steady, enough to see by. Terra is ground. The place most of the walking happens. Sol is the sun: powerful, expensive, not something you stare at all day unless you enjoy consequences.
And Ultra is not a sunbeam. Ultra is opening the roof.
Fable Is Still Haunting the Room
The Claude Fable comparison is everywhere because frontier users have memories. They remember the model that felt proactive, the one that could chase a bug through a browser and write its own scratch tests without being led by the hand. So every Sol reaction arrives with a ghost beside it: is this better than Fable?
The answer on X is messy, which means it is probably honest.
Some users are blunt: Sol is better for auditing, better for complex requirement changes, better for agentic coding work. Artificial Analysis has Sol nearly tied with Fable on broad intelligence and ahead on coding-agent work. That supports the feeling that Sol is winning in the workbench category: the place where a model has to inspect, revise, plan, and keep moving after the first answer.
But “better” is not one thing.
Fable may still have the feel some users want: judgment, restraint, taste, a kind of strange steadiness. Sol feels like it does things. It pushes. It checks. It builds. It keeps going.
That can be brilliant.
That can also be exhausting.
This is why the Fable comparison will not settle cleanly through charts. People are not just comparing scores. They are comparing temperaments. One model feels like a careful senior engineer. Another feels like an overpowered agent with a credit card and a Red Bull.
Depending on the job, either one might be exactly what you want.
The Limit Is Part of the Personality
Seth Rose pointed at one of the more annoying product edges: in Codex, the five-hour usage limit can block users even when they still have weekly usage remaining. That kind of thing sounds like account plumbing until you are in the middle of work. Then it becomes part of how the model feels.
That is the part AI companies still underweight. A model is not just weights and benchmarks. It is also the box around it: limits, defaults, mode names, hidden parallelism, reset behavior, how clearly the product tells you what you are about to spend.
If the model is brilliant but the interface makes people anxious about invisible burn, the brilliance gets filtered through irritation.
You can feel that in the first two days of reaction. The praise is real. The anxiety is real too. People are not asking for Sol to be weaker. They are asking for the steering wheel to be less fuzzy.
Tell them when Ultra is about to behave like four agents.
Make the burn visible.
Make Terra feel like a good default, not the boring compromise people pick after Sol hurts them.
Let Luna be respectable. Cheap models are not shameful when they are doing cheap-model work.
Most of all, stop making users discover the operating manual by hitting the limit wall.
Make the Model Earn Its Burn
My read after two days is simple: GPT-5.6 Sol is good. Maybe very good. Good enough that the complaints matter.
Nobody writes routing advice for a model they plan to ignore. Nobody warns strangers to stop wasting Sol unless Sol is worth saving. Nobody debates Terra as main agent and Sol as deep-dive specialist unless the family is already being treated like infrastructure rather than novelty.
That is a successful launch in the most inconvenient way. OpenAI built something people want to use, then immediately made them think like dispatchers, accountants, and risk managers.
So the sane rule is not complicated.
Use Luna when the work is cheap.
Use Terra when the work is normal.
Use Sol when the work has teeth.
Use Ultra when you mean to start the engine.
The first two days of GPT-5.6 did not produce a clean verdict. They produced something more useful: a survival guide. People like Sol. People trust it enough to hand it real work. People fear it enough to warn each other about the bill, the limits, the overhelping, the hidden team behind the shiny button.
That is where frontier agents are now.
Not “can it answer?”
Can it stop? Can it spend the right amount of itself? Can the product tell the user what kind of machine just woke up under the prompt?
GPT-5.6 Sol may be the model people wanted. Terra and Luna may be the models that keep them sane.
And Ultra should not feel like tapping autocomplete.
It should feel like turning a key and hearing the engine answer back.