Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

477 of 7,078 resources

Showing 51–100

> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…

Active221 month ago
Python

知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题

Active01 month ago
Python

See the upstream model card for full details, training data and citation.

Active2151 month ago
Python

L1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.

Active3871 month ago
Python

💻 Github | 📄 Report | 🚀 Demo

Active2661 month ago
Python

For a convenient overview and download list, visit our model page for this model.

Active1K1 month ago
Python

SmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.

Active581 month ago
Python
Active241 month ago
Python

Minimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.

Active541 month ago
Python
Active9.9K1 month ago
Python

!Aerova

Active1201 month ago
Python

CENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.

Active4931 month ago
Python

!Aerova

Active251 month ago
Python
Active7611 month ago
Python

OmniTCR is a component-aware autoregressive foundation model for learning relationships among peptide epitopes, major histocompatibility complex (MHC) molecules, T-cell receptor alpha chains (TRA) and T-cell receptor beta chains (TRB).

Active01 month ago
Python

PertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.

Active5031 month ago
Python

A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.

Active821 month ago
Python

Eleuthia is a binary protein variant classifier that predicts whether a mutated protein sequence is likely Pathogenic or Benign. It is fine-tuned from ESM-2 and intended for research support in variant prioritization workflows.

Active671 month ago
Python

A 350M encoder that finds nine types of personally identifiable information across 17 languages and returns exact character spans for review and redaction.

Active8341 month ago
Python

BondShift: Organic Mechanism Reasoning

Active132 months ago
Python

Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes Indian English, Hindi, Marathi, Kannada and Telugu, including the heavily code-mixed speech typical of real consultations (English drug names and clinical…

Active1.2K2 months ago
Python

A clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.

Active02 months ago
Python

Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active472 months ago
Python
Active42.8K2 months ago
Python
Active1.3K2 months ago
Python

Technical Report 🧬

Active5.2K2 months ago
Python

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2132 months ago
Python

MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…

Active1.7K2 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5782 months ago
Python

Longevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…

Active902 months ago
Python

!Format !Task !Params !Type !License

Active5.9K2 months ago
Python

HealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.

Active222 months ago
Python

ProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.

Active192 months ago
Python

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).

Active02 months ago
Python

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.

Active02 months ago
Python

A multilingual PII extractor for teams that need structured JSON from clinical and administrative text.

Active1382 months ago
Python

Contrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active222 months ago
Python
Active02 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5792 months ago
Python

464 fine-tuned DNABERT models for regulatory variant effect prediction

Active242 months ago
Python
Active02 months ago
Python

Trinity-Mini-AI-Scientist

Active132 months ago
Python

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active696.9K2 months ago
Python

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active1.3K2 months ago
Python

In retrieval systems, embedding models determine the quality of your search.

Active6.8K2 months ago
Python

MoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.

Active254.2K2 months ago
Python