← All Posts

Opus 5 and the Effort Dial

Technology

Last week Anthropic released Claude Opus 5, and the headline was not intelligence. The intelligence is there, near the top of what exists. The headline was the price: close to flagship capability at roughly half the flagship cost, with a dial that lets you choose how hard the model thinks on any given task. It is the fourth model Anthropic has shipped in its Claude 5 family in under two months. I run my work and a good part of my life on these systems, so a release like this is not news I read. It is news that rearranges my week.

The Bottleneck Moved

For the past few years, the model was the constraint. You designed around its blind spots, kept humans on everything that mattered, and waited for the next release the way farmers wait for rain. That era is ending, quietly and fast. When near-flagship intelligence costs half as much, and a dial lets you spend less on easy work and more on hard work, raw capability stops being the scarce input.

What becomes scarce is everything around the model: the specification that tells it what to build, the evaluation that tells you whether it did, and the judgment to know which tasks deserve maximum effort and which do not. The effort dial makes this literal. Choosing how much thinking to buy per task is now an actual parameter, and someone has to set it well. That is a product decision, not a model feature, and it lands on whoever owns the work.

Do Not Marry a Model

Four releases in two months also settles an argument I used to have with myself about switching costs. There is no blockbuster-migration rhythm anymore, no annual upgrade you plan a quarter around. Improvements arrive like weather now. The only sane response is to make your own system model-agnostic: hold your specs, your tests, and your evaluation sets as the permanent assets, and treat the model as an engine you can swap when a better one shows up.

This is where the discipline pays twice. I already believe evaluation is the hard skill of working with artificial intelligence (AI), because you cannot delegate work you cannot check. A new release adds the second payoff: an evaluation set you trust is also how you adopt a new model in a day instead of a month. Run the new engine against the same checks, compare, promote or decline. No vibes, no vendor benchmarks, your own work as the test.

What I Am Actually Doing About It

So this week looks like this. First, the new model gets a tryout on low-stakes work, the same way I test everything: small blast radius first, my own website before anything that matters more. Second, it runs against the checks I already own, and it gets promoted only where it wins. Third, I take the effort dial seriously as a habit, not just a setting: most tasks do not need maximum thinking, and learning which ones do is exactly the right-sized-rigor muscle I try to apply everywhere else.

The models are becoming abundant and cheap. Your specs, your checks, and your judgment are not. Invest accordingly.

0 Comments