Dear Anthropic, can we please have thought traces back?
I can't verify whether or not the LLM is arriving at the conclusion from cheating, or if it's fudging or making stuff up.
Opus 4.6 remains the best model because of this.
loading...
I can't verify whether or not the LLM is arriving at the conclusion from cheating, or if it's fudging or making stuff up.
Opus 4.6 remains the best model because of this.
loading...