Rendered at 19:19:46 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
itzikkatz 2 hours ago [-]
If its price is almost the same as Opus 5.5 but it’s ultimately not as good, I don’t understand the use case for this model. Perhaps Haiku 5.5 could fill a real need.
simianwords 13 hours ago [-]
It’s better than both fable and Astra? I’m going to start trusting this index less and less.
jug 5 hours ago [-]
It's apparently great at Max but extremely costly. Looks best to me at Medium and High. Over that and I'd go Opus. I consider Fable largely obsolete.
ramon156 10 hours ago [-]
the issue is, what will we use? companies are benchmaxxing, so how do you make a bench that cannot be cheated?
Human curated benches aren't accurate enough
epolanski 6 hours ago [-]
The problem with benchmarks is that they are end to end.
You go from one prompt to the final solution, whereas for many of us it is about the experience of iterative, multi turn working.
Human curated benches aren't accurate enough
You go from one prompt to the final solution, whereas for many of us it is about the experience of iterative, multi turn working.