Terminal-Bench Science (Harbor Framework, 2026)
github.com/harbor-framework/terminal-bench-scienceBenchmark evaluating AI agents on complex real-world scientific workflows in terminal environments across life, physical, earth, and mathematical sciences; featured on model cards for Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro (200+ stars, Apache 2.0)
Sourced from
- Awesome AI for Science — github.com/harbor-framework/terminal-bench-science
- GitHub — github.com/harbor-framework/terminal-bench-science
Related resources
Benchmark evaluating AI agents for end-to-end automated research from re-discovery to new-discovery, with 40 real-science tasks across 10 disciplines, curated datasets from published papers, and expert-curated multimodal rubrics (170+ stars, MIT License)
102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
Multimodal deep learning framework integrating peptide-MHC protein sequence, structure, and biochemical properties to predict class-I immunogenicity for infectious disease epitopes and cancer neoepitopes with cancer-wildtype contrastive learning, enabling personalized vaccine design (Krishnaswamy Lab, Yale University)
Parallel symbolic regression network evaluating millions of expressions on GPU with automated subtree reuse, Nature Computational Science cover article (MIT, 2026)
Advanced OCR with PP-StructureV3 document parsing, 13% accuracy improvement, supports 80+ languages
SOTA multimodal document parsing with 1.2B parameters outperforming GPT-4o, converts PDFs to LLM-ready Markdown/JSON