ResearchClawBench (InternScience, arXiv 2026)

github.com/internscience/researchclawbench
Active264updated 3 weeks ago
Jupyter Notebook
MIT

Benchmark evaluating AI agents for end-to-end automated research from re-discovery to new-discovery, with 40 real-science tasks across 10 disciplines, curated datasets from published papers, and expert-curated multimodal rubrics (170+ stars, MIT License)

Sourced from

  • GitHub — github.com/internscience/researchclawbench
  • Awesome AI for Science — github.com/internscience/researchclawbench

Related resources

First agentic LLM for autonomous data science with end-to-end pipeline from data to analyst-grade reports

Active4.7K2 weeks ago
Python
MIT

Open-source AI workbench for scientific research that automates the full research loop — literature review, hypothesis generation, code writing, experiment execution, database querying, and report writing — with 290+ skills, specialized research agents, and a browser-based workspace (1453+ stars, Apache 2.0, 2026)

Active3.3K1 month ago
TypeScript
Apache-2.0

Offline-first scientific writing workspace powered by Claude, integrating LaTeX, Python, and 100+ scientific skills with local execution, Zotero integration, and privacy-focused design (2026)

Active1.8K2 months ago
TypeScript
MIT

End-to-end autonomous AI research engine that turns an idea into a complete LaTeX paper by dispatching real computational experiments to local GPUs or SLURM clusters, collecting actual results, generating figures/tables, and writing a data-grounded manuscript rather than LLM hallucinations (OpenRaiser, 1.5K+ stars, MIT License, 2026)

Active1.3K1 month ago
Python
MIT

Agent skills (SKILL.md + deterministic tools) for the AI4S workflow — topic exploration, literature survey, runnable experiments, publication-grade papers, and integrity audit, with every citation and number traceable to its source (by ai4s-research, maintainers of this list; MIT, 2026)

Active2362 months ago
Python
MIT

Research coding benchmark curated by scientists with 338 subproblems across 16 subdomains (physics, math, materials, biology, chemistry), evaluating LLMs on realistic scientific programming tasks with gold-standard solutions (NeurIPS 2024)

Active2312 weeks ago
Python
Apache-2.0