Add missing eval-results file for Terminal-Bench 2.1

#41
by SaylorTwift HF Staff - opened

The model card reports Terminal-Bench 2.1: 70.2%, but the repo's .eval_results/ only covers DeepSWE, SWE-bench Multilingual, and SWE-bench Pro. This adds the missing Terminal-Bench 2.1 entry (harborframework/terminal-bench-2.1, task terminalbench_2_1). Note: the card also reports SWE Atlas (Codebase QnA) and Toolathlon Verified, but neither has a registered eval.yaml on the Hub, so they can't be added here.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment