"This is the worst version of the car you'll own."
That was a quote from the delivery manager when my wife and I picked up our Lucid Gravity. He said that since the car is very software-oriented and gets regular over-the-air updates, the car keeps improving. The car I bought is the worst version because it keeps getting better. His experience over the last few years with the company has shown that the changes made in software result in customers growing more fond of their vehicles.
Or maybe they're less annoyed by things that don't work well. Ah, the joys of software.
I was reminded of this interaction when I saw someone send a note at Redgate about AI with this quote: “Today's AI is the worst AI you will ever use.”
Is that true? I guess, to the extent that we think the AI model you use today will improve and its capabilities will grow, that might be true. Certainly, the move from Sonnet 4 to 4.5 to 4.6 to 5 seems to be, well ...
The same to me. Has the model improved? By benchmarks and measures, it has. The same could be said for the GPT models from OpenAI. However, I don't know that I have found much difference. Depending on my guidance, I think I've found that I'm as likely to get great results as poor ones from different models. I've tended to stick with Sonnet models, though often for the focused tasks I pick, the Haiku models work, as do the Opus ones.
I'm cheap, so I don't often see the point in using Opus. I haven't found it to be clearly better, though; to be fair, I'm not letting it loose on large tasks, and I am conscious of costs. I'm not paying, but I tend to treat all things I do as if I were paying the bill myself. That helps me to be aware of when I am being effective and efficient.
When I've seen my cars upgrade (first a BMW, then a Tesla, now a Lucid), there have been clear changes that I've appreciated. There are sometimes annoying upgrades, which always remind me of Jeff Moden: Change is inevitable... change for the better is not. That can be true, but I find that the changes made in software often make sense, and I can see how they an improvement. Even when I don't like them.
For a non-deterministic model, however, I don't know how easy it is to measure change, and if the model is "smarter", but I don't use the extra smart-ness, is it better? I have no idea how to measure these improvements, and I am certainly skeptical that these benchmarks and the frontier model hype is driven more by the need for profit than by a better AI model.