Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
881 of 7,064 resources
Showing 551–600
mradermacher/Biomni-R0-32B-Preview-i1-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
Support our open-source dataset and model releases!
This is a merge of pre-trained language models created using mergekit.
UniParser/MolDet
by UniParserFrom paper: "MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild" (ICCV2025 Accept)
mace-foundations/mace-mh-1
by mace-foundationsMACE-MH-1 is a foundation machine-learning interatomic potential (MLIP) that bridges molecular, surface, and materials chemistry through cross-domain learning:
ZJU-AI4H/Hulu-Med-4B
by ZJU-AI4HHulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Large Language and Vision Assistant for bioMedicine (i.e., “LLaVA-Med”) is a large language and vision model trained using a curriculum learning method for adapting LLaVA to the biomedical domain. It is an open-source release intended for research use only to facilitate reproducibility of the…
microsoft/llava-med-7b-delta
by microsoftThis model was added by Hugging Face staff.
# MedicalLlama3.2-11B-IT ## Model Description This is a fine-tuned version of meta-llama/Llama-3.2-11B-Vision-Instruct on a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset.
!1
InstaDeepAI/BulkRNABert
by InstaDeepAIBulkRNABert is a transformer-based, encoder-only language model pre-trained on bulk RNA-seq profiles from the TCGA dataset using self-supervised masked language modeling, following the original BERT framework. The model is trained to reconstruct randomly masked gene expression values from their…
Hussein-Abdallah/Qwen3-MedQA-FineTuned
by Hussein-AbdallahMedQA Github porject: https://github.com/hussien/MedQA Read the detailed expermintal report here: https://github.com/hussien/MedQA/blob/main/report/MedQA.pdf
ChemFIE-BED is a sentence-transformers based on gbyuvd/chemselfies-base-bertmlm fine-tuned on around (for now) 2 million pairs of valid molecules' SELFIES (Krenn et al. 2020) taken from COCONUTDB (Sorokina et al. 2021) and ChemBL34 (Zdrazil et al. 2023).
vandijklab/C2S-Scale-Gemma-2-27B
by vandijklabGitHub homepage: Cell2Sentence GitHub
AI-follower99/chemical-reaction-predictor
by AI-follower99基于模板的化学反应产物预测模型,使用图注意力网络(GAT)预测反应中心。
Complete layer-wise protein embeddings for 236,252 human proteins using ESMC models
This model is a fine-tuned version of google/medgemma-4b-it adapted for binary mammogram classification on the OMAMA 256×256 dataset. The dataset consists of ~154k mammogram image slices (.npz) with metadata JSONs providing labels (NonCancer, Cancer).
mradermacher/Gemma-2-2B-MedicalQA-Assistant-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
tahoebio/Tahoe-x1
by tahoebioTahoe-x1 is a family of perturbation-trained single-cell foundation models with up to 3 billion parameters, developed by Tahoe Therapeutics. Pretrained on 266 million single-cell transcriptomic profiles including the Tahoe-100M perturbation compendium, Tahoe-x1 achieves state-of-the-art performance…
imageomics/biocap
by imageomicsBioCAP is a foundation model for biology organismal images. It is trained on TreeOfLife-10M with synthetic captions (TreeOfLife-10M-Captions) as supervision on the basis of a CLIP model (ViT-B/16) pre-trained by OpenAI. BioCAP achieves state-of-the-art performance on text-image retrieval tasks.
weidawang/Chem-R-8B
by weidawangThe Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
Specialized model for Chemical Entity Recognition - Chemical entities from the BC5CDR dataset
Specialized model for Species Entity Recognition - Species and organism names
elonlit/GeneJEPA
by elonlitGeneJEPA is a Joint-Embedding Predictive Architecture (JEPA) trained for self-supervised representation learning on scRNA-seq. It uses a Perceiver-style encoder to handle sparse, high-dimensional gene count vectors and a Fourier-feature tokenizer for numerical tokenization.
DermLIP is a vision-language model for dermatology, trained on the Derm1M dataset—the largest dermatological image-text corpus to date. This model variant (PanDerm-base-w-PubMed-256) utilizes domain-specific pretraining to deliver superior performance compared to other DermLIP variants..
ZJU-AI4H/SigLip-NaViT
by ZJU-AI4HHulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
biomni/Biomni-R0-32B-Preview
by biomni# Biomni-R0-32B-Preview This repo contains the weights of Biomni-R0-32B-Preview, a research preview of the series of biomedical AI agents trained by the Biomni team.
InstaDeepAI/instanovoplus-v1.1.0
by InstaDeepAIInstaNovoPlus is a diffusion-based model for de novo peptide sequencing from mass spectrometry data. This model leverages multinomial diffusion for accurate, database-free peptide identification for large-scale proteomics experiments.
Typhoon-Si-Med-Thinking-4B is Southeast Asia’s first state-of-the-art, small, and efficient medical reasoning model, jointly developed by Typhoon (SCB 10X) and the Siriraj Informatics and Data Innovation Center (SiData+) at Siriraj Hospital, Mahidol University.
stanfordmimi/MedVAL-4B
by stanfordmimiMedVAL-4B (medical text validator) is a language model fine-tuned to assess AI-generated medical text outputs at near physician-level reliability.
This model is a lightweight model pre-trained on SELFIES (Self-Referencing Embedded Strings) representations of molecules. It is trained on 2.7M unique and valid molecules taken from COCONUTDB and ChemBL34, with 7.3M total generated masked examples.
nvidia/AMPLIFY_350M
by nvidia> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.
nvidia/AMPLIFY_120M
by nvidia> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.
Here is the pretrained version of PolyTAO, the first pretrained generative language model for polymer design.
lingshu-medical-mllm/Lingshu-32B
by lingshu-medical-mllmWebsite 🤖 7B Model 🤖 32B Model MedEvalKit Technical Report Lingshu MCP
lingshu-medical-mllm/Lingshu-7B
by lingshu-medical-mllmWebsite 🤖 7B Model 🤖 32B Model MedEvalKit Technical Report Lingshu MCP
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
evo-design/evo-2-7b-8k-microviridae
by evo-designEvo 2 is a state of the art DNA language model for long context modeling and design. Evo 2 models DNA sequences at single-nucleotide resolution at up to 1 million base pair context length using the StripedHyena 2 architecture, using Savanna.
stanfordmimi/RoentGen-v2
by stanfordmimihu-lab/PlantGFM
by hu-lab# Model Card for Model ID PlantGFM is a genetic foundation model pre-trained on the complete genome sequences of 12 model plants, encompassing 108 billion nucleotides. Using the Hyena framework with 220 million parameters and a context length of 64K bp, PlantGFM models sequences at…
kuleshov-group/PlantCAD2-Small-l24-d0768
by kuleshov-groupNovoMolGen is a family of molecular foundation models trained on 1.5 billion ZINC-22 molecules with Llama architectures and FlashAttention. It achieves state-of-the-art performance on both unconstrained and goal-directed molecule generation tasks.