Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

641 of 6,511 resources

Showing 150

Run locally - Benchmarks - Whitepaper - Model details - Responsible use

Active622 hours ago

BondShift: Organic Mechanism Reasoning

Active123 days ago
Python

Vision-language model for dermatology, pretrained with MAGEN (Multi-Agent data GENeration) and O-MAKE (Ontology-based Multi-Aspect Knowledge-Enhanced pretraining).

Active434 days ago

A clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.

Active05 days ago
Python

Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active276 days ago
Python

SERAPH is a deep learning model designed for 3-state (Q3) protein secondary structure prediction. It processes raw single amino acid sequences and predicts residue-level secondary structure states: Alpha Helix (H), Beta Sheet (E), or Coil/Loop (C).

Active01 week ago

!TVBP Architecture !Parameters !Trainable !Brownian Reservoir !Framework !Biology

Active1201 week ago

Technical Report 🧬

Active5.2K1 week ago
Python

Nesso-1 is a fast, structure-based protein–ligand binding-affinity model. Given a protein sequence and a ligand (SMILES / CCD code / SDF), it predicts a binding affinity scalar along with a binder/non-binder score.

Active61.9K1 week ago

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2131 week ago
Python

MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…

Active1.6K1 week ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5781 week ago
Python

💻 GitHub Repo • 🤗 Model Collections • 📖 MemSFT Paper

Active281 week ago

HealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.

Active221 week ago
Python

ProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.

Active191 week ago
Python
Active1341 week ago
Active1751 week ago

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).

Active01 week ago
Python

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.

Active01 week ago
Python
Active1972 weeks ago

Contrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active222 weeks ago
Python

🩺 HealthGPT-Pro: A High-Performance Multimodal Large Language Model for Medical Understanding and Analysis

Active2.1K2 weeks ago

A binary healthy hard coral vs bleached hard coral classifier built on top of the ReefNet species LoRA model BobDerBaum/bioclip-2.5-vith14-reefnet-lora, which provides the fine-tuned vision-encoder LoRA adapters. Only a small linear head is trained on top (frozen backbone + LoRA + 2-way linear…

Active02 weeks ago

A LoRA fine-tune of imageomics/bioclip-2.5-vith14 trained contrastively on the ReefNet 1.0 coral-reef species dataset (ReefNet/ReefNet-1.0), 92-class global curated split (train 48,312 / image-val 32,792 / image-test 33,090 / source-val 8,074; split cache 56ea94e36f9f).

Active42 weeks ago
Active02 weeks ago
Python

A lightweight language model designed to answer science questions clearly and accurately in English.

Active2202 weeks ago

!Nucleus Resynthesis

Active1.6K2 weeks ago

Winnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a calibrator trained on the HeLa Single Shot dataset as referenced in our paper: De novo peptide sequencing rescoring and FDR estimation with Winnow.

Active272 weeks ago

Winnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a pretrained, general-purpose calibrator that maps raw InstaNovo model confidences and complementary features (mass error, retention time, beam features, fragment…

Active332 weeks ago

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active468.3K2 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active58.5K2 weeks ago
Python

464 fine-tuned DNABERT models for regulatory variant effect prediction

Active242 weeks ago
Python
Active42 weeks ago
Python

GB.DNA-7B is DNA foundation model trained on 10.6 billion nucleotides from 796 species, enabling genome mining, in silico mutagenesis studies, gene expression prediction, and directed sequence generation.

Active873 weeks ago

Trinity-Mini-AI-Scientist

Active133 weeks ago
Python

This is the first release of the Cherimoya Accessibility aTlas (CATv1): a collection of over 7,500 Cherimoya models trained on DNase-seq and ATAC-seq experiments from the ENCODE Project. Cherimoya models are state-of-the-art predictors of local chromatin accessibility, mapping a DNA sequence to a…

Active03 weeks ago

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active520.7K3 weeks ago
Python

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active5123 weeks ago
Python

In retrieval systems, embedding models determine the quality of your search.

Active8.8K3 weeks ago
Python

MoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.

Active269.8K3 weeks ago
Python

The commands below use the glyph package. Install it from the code repository:

Active03 weeks ago

The commands below use the glyph package. Install it from the code repository:

Active2853 weeks ago
Python

scUNVEIL is a pretrained foundation model for human single-cell RNA sequencing (scRNA-seq).

Active1483 weeks ago

This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.

Active143 weeks ago
Python