[Bug] Model Performance Regression: Opus 4.5-4.6 Quality Degradation and Fable Hallucinations in Agentic Tasks
Bug Description
Opus has regressed to <4.6 levels so i thought id give Fable a try thanks to the %50higher limit. I wanted to give it a rather simple crud job, one person Wave like invoice maker local docker app. From hallucinations (self admitted), to completely ignored direction, to running the same tests 3 times (not one more after correction, green tests 3 times), to bungling ui features. Something is definetly wrong/broken. Opus is as 4.5-6 levels and Fable has gone back to slightly (maybe) better then 4.8. i took 2 days to do what opus 4.6 would do in half a day. more defects; on auto mode forgets to continue after merging. when i ask, it admits it forgot to. it does not track agents closely wastes more time waiting when agent is doing nothing. pauses midway for no reason.
i am sorry to say this but claude models (or claude code?) has been a dissaster the past 5 days. wasted time/token/frustration. i am not sure how much more of this i can take from my main driver model.
Environment Info
- Platform: darwin
- Terminal: iTerm.app
- Version: 2.1.214
- Feedback ID: a447a0b8-b222-459e-874a-0c152a05a30d
Errors
[]