Terminal-Bench-Science: Evaluating AI agents on scientific research workflows 60 points by matt_d 5 hours ago 14 comments story