Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
672 of 6,565 resources
Showing 401–450
This is a merge of pre-trained language models created using mergekit.
UniParser/MolDet
by UniParserFrom paper: "MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild" (ICCV2025 Accept)
mace-foundations/mace-mh-1
by mace-foundationsMACE-MH-1 is a foundation machine-learning interatomic potential (MLIP) that bridges molecular, surface, and materials chemistry through cross-domain learning:
ZJU-AI4H/Hulu-Med-4B
by ZJU-AI4HHulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Large Language and Vision Assistant for bioMedicine (i.e., “LLaVA-Med”) is a large language and vision model trained using a curriculum learning method for adapting LLaVA to the biomedical domain. It is an open-source release intended for research use only to facilitate reproducibility of the…
!1
Hussein-Abdallah/Qwen3-MedQA-FineTuned
by Hussein-AbdallahMedQA Github porject: https://github.com/hussien/MedQA Read the detailed expermintal report here: https://github.com/hussien/MedQA/blob/main/report/MedQA.pdf
ChemFIE-BED is a sentence-transformers based on gbyuvd/chemselfies-base-bertmlm fine-tuned on around (for now) 2 million pairs of valid molecules' SELFIES (Krenn et al. 2020) taken from COCONUTDB (Sorokina et al. 2021) and ChemBL34 (Zdrazil et al. 2023).
vandijklab/C2S-Scale-Gemma-2-27B
by vandijklabGitHub homepage: Cell2Sentence GitHub
This model is a fine-tuned version of google/medgemma-4b-it adapted for binary mammogram classification on the OMAMA 256×256 dataset. The dataset consists of ~154k mammogram image slices (.npz) with metadata JSONs providing labels (NonCancer, Cancer).
mradermacher/Gemma-2-2B-MedicalQA-Assistant-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
tahoebio/Tahoe-x1
by tahoebioTahoe-x1 is a family of perturbation-trained single-cell foundation models with up to 3 billion parameters, developed by Tahoe Therapeutics. Pretrained on 266 million single-cell transcriptomic profiles including the Tahoe-100M perturbation compendium, Tahoe-x1 achieves state-of-the-art performance…
imageomics/biocap
by imageomicsBioCAP is a foundation model for biology organismal images. It is trained on TreeOfLife-10M with synthetic captions (TreeOfLife-10M-Captions) as supervision on the basis of a CLIP model (ViT-B/16) pre-trained by OpenAI. BioCAP achieves state-of-the-art performance on text-image retrieval tasks.
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
Specialized model for Species Entity Recognition - Species and organism names
elonlit/GeneJEPA
by elonlitGeneJEPA is a Joint-Embedding Predictive Architecture (JEPA) trained for self-supervised representation learning on scRNA-seq. It uses a Perceiver-style encoder to handle sparse, high-dimensional gene count vectors and a Fourier-feature tokenizer for numerical tokenization.
DermLIP is a vision-language model for dermatology, trained on the Derm1M dataset—the largest dermatological image-text corpus to date. This model variant (PanDerm-base-w-PubMed-256) utilizes domain-specific pretraining to deliver superior performance compared to other DermLIP variants..
ZJU-AI4H/SigLip-NaViT
by ZJU-AI4HHulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
biomni/Biomni-R0-32B-Preview
by biomni# Biomni-R0-32B-Preview This repo contains the weights of Biomni-R0-32B-Preview, a research preview of the series of biomedical AI agents trained by the Biomni team.
InstaDeepAI/instanovoplus-v1.1.0
by InstaDeepAIInstaNovoPlus is a diffusion-based model for de novo peptide sequencing from mass spectrometry data. This model leverages multinomial diffusion for accurate, database-free peptide identification for large-scale proteomics experiments.
Typhoon-Si-Med-Thinking-4B is Southeast Asia’s first state-of-the-art, small, and efficient medical reasoning model, jointly developed by Typhoon (SCB 10X) and the Siriraj Informatics and Data Innovation Center (SiData+) at Siriraj Hospital, Mahidol University.
stanfordmimi/MedVAL-4B
by stanfordmimiMedVAL-4B (medical text validator) is a language model fine-tuned to assess AI-generated medical text outputs at near physician-level reliability.
This model is a lightweight model pre-trained on SELFIES (Self-Referencing Embedded Strings) representations of molecules. It is trained on 2.7M unique and valid molecules taken from COCONUTDB and ChemBL34, with 7.3M total generated masked examples.
nvidia/AMPLIFY_350M
by nvidia> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.
nvidia/AMPLIFY_120M
by nvidia> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.
lingshu-medical-mllm/Lingshu-32B
by lingshu-medical-mllmWebsite 🤖 7B Model 🤖 32B Model MedEvalKit Technical Report Lingshu MCP
lingshu-medical-mllm/Lingshu-7B
by lingshu-medical-mllmWebsite 🤖 7B Model 🤖 32B Model MedEvalKit Technical Report Lingshu MCP
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
evo-design/evo-2-7b-8k-microviridae
by evo-designEvo 2 is a state of the art DNA language model for long context modeling and design. Evo 2 models DNA sequences at single-nucleotide resolution at up to 1 million base pair context length using the StripedHyena 2 architecture, using Savanna.
hu-lab/PlantGFM
by hu-lab# Model Card for Model ID PlantGFM is a genetic foundation model pre-trained on the complete genome sequences of 12 model plants, encompassing 108 billion nucleotides. Using the Hyena framework with 220 million parameters and a context length of 64K bp, PlantGFM models sequences at…
kuleshov-group/PlantCAD2-Small-l24-d0768
by kuleshov-groupNovoMolGen is a family of molecular foundation models trained on 1.5 billion ZINC-22 molecules with Llama architectures and FlashAttention. It achieves state-of-the-art performance on both unconstrained and goal-directed molecule generation tasks.
Palmyra-Med, a powerful LLM designed for healthcare
Neeto-1.0-8b is an openly released biomedical large language model (LLM) created by BYOL Academy to assist learners and practitioners with medical exam study, literature understanding, and structured clinical reasoning.
ByteDance-Seed/bamboo_mixer
by ByteDance-SeedThis repository contains the official model of the paper A Unified Predictive and Generative Solution for Liquid Electrolyte Formulation.
sagawa/ReactionT5v2-forward
by sagawaThis is a ReactionT5 pre-trained to predict the products of reactions. You can use the demo here.
This is a ReactionT5 pre-trained to predict the reactants of reactions. You can use the demo here.
This repos contains the biomedicine MLLM developed from Qwen2.5-VL-3B-Instruct in our paper: On Domain-Adaptive Post-Training for Multimodal Large Language Models. The correspoding training dataset is in biomed-visual-instructions.
zhw-e8/LAMAR
by zhw-e8# LAMAR LAMAR is a Foundation Language Model for RNA Regulation, which achieves better or comparable performance compared to baseline models in various RNA regulation tasks, helping to decipher the rules of RNA regulation.
Specialized model for Chemical Entity Recognition - Identifies chemical compounds and substances in biomedical literature
Specialized model for Chemical Entity Recognition - Identifies chemical compounds and substances in biomedical literature
Specialized model for Chemical Entity Recognition - Chemical entities from the BC5CDR dataset
lokeshch19/ModernPubMedBERT
by lokeshch19A specialized medical embedding model fine-tuned from Clinical ModernBERT using InfoNCE contrastive learning on PubMed title-abstract pairs.
ameya98/JAMUN
by ameya98JAMUN is a novel approach for generating conformational ensembles of protein structures, presented in the paper JAMUN: Bridging Smoothed Molecular Dynamics and Score-Based Learning for Conformational Ensembles.