REFUTE

github.com/connerlambden/refute-inspect
Active2updated 1 month ago
Python
MIT

REFUTE is an open benchmark for scientific critique honesty and epistemic calibration on recent life-science and biomedical literature. It tests whether models keep claims inside what the evidence allows (overclaim / planted-flaw / falsifier selection) and whether stated confidence is calibrated, with judge-free MCQ axes plus open-ended critique scoring.

Sourced from

  • GitHubgithub.com/connerlambden/refute-inspect
  • bio.toolsbgpt-refute

Related resources

A benchmarking platform for molecular generation models.

Stale9882 years ago
Python
MIT

Weather prediction benchmark

Stale8322 years ago
Jupyter Notebook
MIT

Large-scale benchmark suite for protein fitness prediction and design, aggregating 200+ deep mutational scanning assays and clinical variant datasets across diverse protein families and taxa, with standardized zero-shot and supervised leaderboards for variant effect prediction, mutation effect prediction, and protein language model evaluation (OATML & Marks Lab, NeurIPS 2023 Spotlight, Datasets & Benchmarks)

Active4645 months ago
HTML
MIT

Scientific machine learning benchmarks & differential equation solvers

Active3441 month ago
MATLAB
MIT

Unified benchmarking framework for protein representation learning, providing standardized interfaces for pre-training and diverse downstream tasks including structure prediction, fitness prediction, and property prediction across multiple protein datasets and model architectures (ICLR 2024, 273+ stars, MIT License)

Idle2751 year ago
Python
MIT

Benchmark evaluating AI agents for end-to-end automated research from re-discovery to new-discovery, with 40 real-science tasks across 10 disciplines, curated datasets from published papers, and expert-curated multimodal rubrics (170+ stars, MIT License)

Active2311 month ago
Jupyter Notebook
MIT