Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

887 of 7,078 resources

Showing 551–600

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

Idle18110 months ago
Python

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology. This model version was continually pretrained on ~14 million cancer transcriptomes…

Idle1610 months ago
Python

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

Idle30710 months ago
Python

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

Idle1710 months ago
Python

## Model Overview PlantBiMoE is a DNA language model trained on 42 representative plant species genomes. More specifically, PlantBiMoE uses the BiMamba and SparseMoE architecture with a masked language modeling objective to leverage highly available genotype data from 42 different plant speices to…

Idle1010 months ago

For a convenient overview and download list, visit our model page for this model.

Idle16610 months ago
Python

Support our open-source dataset and model releases!

Idle6010 months ago
Python

This is a merge of pre-trained language models created using mergekit.

Idle4110 months ago

From paper: "MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild" (ICCV2025 Accept)

Idle010 months ago

MACE-MH-1 is a foundation machine-learning interatomic potential (MLIP) that bridges molecular, surface, and materials chemistry through cross-domain learning:

Idle010 months ago

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Idle4.5K10 months ago
Python

Large Language and Vision Assistant for bioMedicine (i.e., “LLaVA-Med”) is a large language and vision model trained using a curriculum learning method for adapting LLaVA to the biomedical domain. It is an open-source release intended for research use only to facilitate reproducibility of the…

Idle9.9K10 months ago
Python

This model was added by Hugging Face staff.

Idle29410 months ago
Python

# MedicalLlama3.2-11B-IT ## Model Description This is a fine-tuned version of meta-llama/Llama-3.2-11B-Vision-Instruct on a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset.

Idle2011 months ago

BulkRNABert is a transformer-based, encoder-only language model pre-trained on bulk RNA-seq profiles from the TCGA dataset using self-supervised masked language modeling, following the original BERT framework. The model is trained to reconstruct randomly masked gene expression values from their…

Idle26911 months ago
Python

MedQA Github porject: https://github.com/hussien/MedQA Read the detailed expermintal report here: https://github.com/hussien/MedQA/blob/main/report/MedQA.pdf

Idle14311 months ago

ChemFIE-BED is a sentence-transformers based on gbyuvd/chemselfies-base-bertmlm fine-tuned on around (for now) 2 million pairs of valid molecules' SELFIES (Krenn et al. 2020) taken from COCONUTDB (Sorokina et al. 2021) and ChemBL34 (Zdrazil et al. 2023).

Idle11311 months ago
Python

GitHub homepage: Cell2Sentence GitHub

Idle1.5K11 months ago
Python

基于模板的化学反应产物预测模型,使用图注意力网络(GAT)预测反应中心。

Idle5611 months ago

Complete layer-wise protein embeddings for 236,252 human proteins using ESMC models

Idle011 months ago
Idle798.5K11 months ago
Python

This model is a fine-tuned version of google/medgemma-4b-it adapted for binary mammogram classification on the OMAMA 256×256 dataset. The dataset consists of ~154k mammogram image slices (.npz) with metadata JSONs providing labels (NonCancer, Cancer).

Idle1511 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Idle46811 months ago
Python

!Main Image

Idle23211 months ago
Python

!Main Image

Idle48211 months ago

Tahoe-x1 is a family of perturbation-trained single-cell foundation models with up to 3 billion parameters, developed by Tahoe Therapeutics. Pretrained on 266 million single-cell transcriptomic profiles including the Tahoe-100M perturbation compendium, Tahoe-x1 achieves state-of-the-art performance…

Idle5911 months ago

BioCAP is a foundation model for biology organismal images. It is trained on TreeOfLife-10M with synthetic captions (TreeOfLife-10M-Captions) as supervision on the basis of a CLIP model (ViT-B/16) pre-trained by OpenAI. BioCAP achieves state-of-the-art performance on text-image retrieval tasks.

Idle27711 months ago
Idle2011 months ago

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle14.1K11 months ago
Python

Specialized model for Chemical Entity Recognition - Chemical entities from the BC5CDR dataset

Idle43.3K11 months ago

Specialized model for Species Entity Recognition - Species and organism names

Idle911 months ago

GeneJEPA is a Joint-Embedding Predictive Architecture (JEPA) trained for self-supervised representation learning on scRNA-seq. It uses a Perceiver-style encoder to handle sparse, high-dimensional gene count vectors and a Fourier-feature tokenizer for numerical tokenization.

Idle011 months ago

DermLIP is a vision-language model for dermatology, trained on the Derm1M dataset—the largest dermatological image-text corpus to date. This model variant (PanDerm-base-w-PubMed-256) utilizes domain-specific pretraining to deliver superior performance compared to other DermLIP variants..

Idle12912 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Idle812 months ago

# Biomni-R0-32B-Preview This repo contains the weights of Biomni-R0-32B-Preview, a research preview of the series of biomedical AI agents trained by the Biomni team.

Idle5.5K12 months ago

InstaNovoPlus is a diffusion-based model for de novo peptide sequencing from mass spectrometry data. This model leverages multinomial diffusion for accurate, database-free peptide identification for large-scale proteomics experiments.

Idle261 year ago

Typhoon-Si-Med-Thinking-4B is Southeast Asia’s first state-of-the-art, small, and efficient medical reasoning model, jointly developed by Typhoon (SCB 10X) and the Siriraj Informatics and Data Innovation Center (SiData+) at Siriraj Hospital, Mahidol University.

Idle511 year ago

MedVAL-4B (medical text validator) is a language model fine-tuned to assess AI-generated medical text outputs at near physician-level reliability.

Idle2031 year ago
Python

This model is a lightweight model pre-trained on SELFIES (Self-Referencing Embedded Strings) representations of molecules. It is trained on 2.7M unique and valid molecules taken from COCONUTDB and ChemBL34, with 7.3M total generated masked examples.

Idle601 year ago
Python

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

Idle341 year ago
Python

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

Idle5831 year ago
Python

Here is the pretrained version of PolyTAO, the first pretrained generative language model for polymer design.

Idle1051 year ago
Python

Evo 2 is a state-of-the-art DNA language model trained autoregressively on trillions of DNA tokens.

Idle1411 year ago

Website    🤖 7B Model    🤖 32B Model    MedEvalKit    Technical Report    Lingshu MCP

Idle7061 year ago
Python

Website    🤖 7B Model    🤖 32B Model    MedEvalKit    Technical Report    Lingshu MCP

Idle7.1K1 year ago
Python

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle2.9K1 year ago
Python