Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

(terminal-bench-science.ai)

34 points | by matt_d 2 hours ago

5 comments

  • jerpint 23 minutes ago
    The fact that opus 5 is outperforming fable is odd to me

    From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)

    • jhbadger 3 minutes ago
      This is measuring on scientific tasks though. I haven't used Fable lately but when I was playing with it when it was new its "safety" features made it practically impossible to use for biomedical science -- it once refused to work on a pipeline of mine that was analyzing pathogenicity islands in bacteria (presumably because it had guardrails to that effect to stop potential bioterrorists and the like)
  • akshay_akula 1 hour ago
    Evals on actual research workflows is the right direction, most agent benches are toy tasks.
  • rubslopes 2 hours ago
    I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
  • mlmonkey 53 minutes ago
    Sad to see no mention of Gemini ...
  • vatsachak 1 hour ago
    Damn. These things aren't AGI... but I don't care.

    Luna is good enough for me to give a parser spec and have it write one.