Erik Istre
@eistre91
Software engineer mostly talking about LLMs and video games. My opinions are my own, etc. etc.
I've seen people complaining about Opus 5 performing in ways that I haven't encountered. This is probably workflow specific in a lot of ways. But also I wonder if people are using different effort levels like xhigh/max and underestimating the impact?
Hot takes on Opus 5. Promising release. - So Sonnet 5 is just completely pointless now isn't it? Opus 4.8 was already better on a cost per task/performance ratio and this looks better. - What's up with the notable quality decrease on coding benchmarks moving from xhigh -> max?
That hasn't been my experience but it's plausible that there's hard to define stuff that we do differently which results in different experiences. I will say Luna is HIGHLY variable based on effort. Its effort curve is insane on benchmarks like Luna and genuinely abnormal from previous models.
The agent steps graph really tells the story pretty well doesn't it. This feels like they tried to give it Fable-like behavior with relentlessly pursuing a goal, but it fails to do so well because it doesn't have enough latent knowledge due to a smaller parameter count.
About right. I think AI can be super useful and that it's also arriving in a time that society isn't ready to absorb it, no political will to deal with the real downsides like electricity consumption, and true transformative integration into knowledge work will take place over ~5 years.