Contents
- But, what is this "model drift" you talk about?
- Sometimes the release notes are more useful than the launch post
- "Smarter" according to whom?
- Ahem, prompts are not permanent instructions
- Configurations used for production work should be treated like dependencies
- A rollback is not a vote against progress
Anthropic released Claude Opus 5.5 yesterday. The announcement says it is faster, costs less, and beats Opus 5 across agentic coding, computer use, and knowledge work. According to some paper at Anthropic, this was going to be an easy upgrade.
Ahem, in my case, Anthropic automatically flipped the switch and I was using 5.5 without realizing it.
Work patterns that had been refined over the previous six months stopped producing the expected acceptable results. Same kinds of tasks. Same instructions. Same repository rules.
Approaches that had been dependable were suddenly unreliable, and the usual corrections did not help. In some cases, the result was not merely different. It was frustratingly worse. Argumentative in a sense, like Claude no longer wanted to be a tool.
Anthropic's benchmarks are probably correct about it being a better model. But the work sitting in front of me that now needs to be redone has me questioning that claim.
This change in how models respond after an upgrade is called "model drift". It's not a big deal when you're turning your family pictures into retro 80's-themed portraits. Where it matters to me is in production work.
But, what is this "model drift" you talk about?
"Model drift" as a term gets applied to several related problems.
In traditional machine learning, it often describes a model losing accuracy. In the real world, the incoming data, or the relationship between inputs and outcomes, has changed from what pre-release versions were trained on.
With large language models, people also use it to describe a persistent change in behavior: different reasoning habits, different tool choices, different refusal patterns, or a prompt that no longer works the way it did before.
The move from Opus 5 to Opus 5.5 was what is considered a model migration. Anthropic did not secretly replace the weights behind the Opus 5 model ID. Its model versioning documentation says current model IDs point to fixed snapshots. Anthropic also says that serving infrastructure, including routers, safety classifiers, and sampling logic, can change around a fixed model and cause smaller observable differences.
It's nice to understand how it happens but it matters less when a working process has been broken.
Now I have to approach every session with the question, "did the instructed behavior move far enough that I can no longer trust the new result?"
Sometimes the release notes are more useful than the launch post
Anthropic's Opus 5.5 announcement makes a strong case for the new model. The company reports better coding and knowledge-work scores, more than 30% faster output, and about 40% lower cost on typical workloads. Those are real improvements if they hold up with the systems you've invested time and energies into. But sometimes the defaults you've set for your work become the reason a change in the model breaks them.
Opus 5 used high as its default effort. Opus 5.5 defaults to medium.
Opus 5.5 tends to think more at the same effort level, especially near the top end, and thinking can no longer be disabled.
Some forced tool-use settings now return an error. Progress text between tool calls arrives differently. The safeguard system covers more categories, and certain requests can be routed to undesired model.
Even a change in the tone of the model can be a surprise.
Anthropic calls these breaking changes and behavior differences. Its own migration guide tells developers to set effort explicitly, re-evaluate instructions tuned for Opus 5, rerun their tests, and try the model in development before moving production traffic.
That is not fine print. It is the provider saying the old setup may not travel cleanly.
Even if you changed nothing except the model name, quite a lot has still changed.
And that has a lot of impact when you've built infrastructure, documentation, and workflows around them.
"Smarter" according to whom?
Benchmarks have to reduce performance to a number, even if your business does not.
A coding benchmark can say a model completes more tasks while your own work gets worse because completing a benchmark was never your work's only requirement.
Maybe the model reaches a valid answer by taking an approach your codebase cannot maintain. Maybe it uses more initiative where you wanted restraint.
Maybe it follows the general instruction while missing the one local convention that keeps the result safe. Maybe it spends its effort solving a problem you did not ask it to solve.
Those failures disappear inside an aggregate score. They are painfully obvious when you have to clean up after them.
Anthropic makes a surprisingly candid point in its own announcement: at this level of capability, small benchmark margins are becoming less useful for predicting real-world differences. They really should display that in larger text for all to see.
An upgrade can raise the average and lower the value of your responses with the same input as before. And, to be fair, Anthropic isn't improving their model around what you do specifically.
But the inconsistencies do make it difficult to support investing too heavily in mission-critical ai-assisted workflows.
Ahem, prompts are not permanent instructions
For a while, prompts felt closer to documentation than code. Write the instructions carefully, refine it until the output settles down, then reuse it.
Sheesh, that mental model is getting expensive.
If you think about it, a prompt is more like an integration contract with a dependency you have no control over.
Refined results find success in a strict combination of model, effort setting, prompt, tool set, permission mode, and the serving layer around all of it. Change any one of those pieces and a sentence that used to steer the model could become unnecessary, weak, or actively counterproductive.
This is why adding more prompt text is often the wrong first response to an upgrade.
You patch one surprising behavior, then patch the behavior caused by that patch, and soon the prompt reads like a workplace policy written after fourteen separate unrelated incidents. You may be able to get back to baseline but it may also make the next model even worse.
Before rewriting the instruction, prove what moved.
Configurations used for production work should be treated like dependencies
If your workflow requires more nuanced instructions and hand-waving, treat every model upgrade as you would treat a dependency upgrade in your production software.
Keep a small regression set made from real work. Not a generic benchmark. Use ten or twenty tasks the system actually performs, including the awkward ones that expose bad judgment. Save the inputs, the expected constraints, and enough of the prior result to explain what "good" means.
Record the whole configuration. Model ID, effort level, tool definitions, system instructions, permissions, prompt version, and any fallback behavior. "We used Claude" is not a reproducible record. At this point, it is barely a useful sentence.
Run the old and new model side by side. Compare correctness, approach, edits made, tool calls, time, and cleanup. Token cost matters, but the cheapest run is still expensive if a person spends an hour undoing it.
Set defaults explicitly. Opus 5.5 changing from high to medium effort is a clean example of why. A hidden default is still part of your system. Write it down in configuration so a provider's new default does not become your accidental decision.
Keep a rollback path where the provider allows it. Opus 5 is still an active model. If Opus 5.5 fails your evaluation, stay where the work is reliable while you decide whether the new model needs different instructions or simply does not fit that job yet. Some managed tools do not give you that choice, which makes testing and human review even more important.
Put a person at the expensive boundaries. Anything that publishes, deletes, sends, charges, deploys, or changes production data should have a checkpoint that does not depend on the model behaving like it did last week.
None of this is glamorous. Neither is discovering after a forced upgrade that six months of learned behavior was part of the dependency too.
A rollback is not a vote against progress
AI companies have every reason to talk about each release as a step forward. That does not mean that release will work for every one of your tasks.
I wrote in July that a new model release is usually a reason to wait, not rush. This week gave me a less theoretical example with an automatic "upgrade." But don't let that force you to switch to the new model just because a provider thinks it's in your best interest. The right time to move is after the new model passes guidelines and the definitions that you have in place for successful results.
My opinion is straightforward: model providers own the benchmark claim, and you own the production result.
If an upgrade makes your process worse, switch back. Adjust the configuration. Rewrite the prompt once the evidence tells you where and why. Split the workload across models where it makes sense but only if that is what your work needs. But do not accept a worse result because someone's chart said a different version number was "better."
I have no interest in winning the release-day race. I want repeatable work. If Opus 5 gives me that today and Opus 5.5 does not, Opus 5 keeps the job.
We use the same standard in our AI development work. Contact us if you're looking for production software and require something more than "the new model looked great in the launch post."