Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

54 of 7,064 resources

Showing 1–50

A taxonomy-informed sparse DNA foundation model for microbial genomics.

Active5614 days ago
Python

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Kimi K3, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with GPT-5.6 TerraPro, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Claude Opus 5, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

Contrastive LEarning with Soft Targets from TCRdist. Checkpoint SCEPTR6LACsoft800k20ep_bs1024.

Active121 week ago

Three genomic foundation models, packaged together for local inference on Apple silicon.

Active01 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active271 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active261 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active261 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active301 week ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active1.1K3 weeks ago
Python

PhenoSeq is a Gaussian diffusion model that generates scGPT RNA-seq embeddings conditioned on ViT-L microscopy imaging features. Given fluorescence microscopy images of a cell or well, it predicts a 512-dimensional scGPT embedding representing the transcriptomic state of individual cells — enabling…

Active601 month ago

scRep is a PyTorch model for extracting cell embeddings from single-cell RNA-seq AnnData (.h5ad) inputs. This repository is a standalone Hugging Face release bundle: it includes checkpoint weights, the exact paired gene vocabulary, the inference implementation, and runnable examples.

Active01 month ago
Active241 month ago
Python
Active7611 month ago
Python

In retrieval systems, embedding models determine the quality of your search.

Active6.8K2 months ago
Python

MoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.

Active255.3K2 months ago
Python

scUNVEIL is a pretrained foundation model for human single-cell RNA sequencing (scRNA-seq).

Active1482 months ago

# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language models…

Active773 months ago
Python

PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.

Active923 months ago
Python

Jolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:

Active4043 months ago
Python

JEPA-based spatial-transcriptomics foundation model (TERRA). Code & docs: https://github.com/Lotfollahi-lab/terra

Active03 months ago

This Hugging Face repository stores the official resources for RAMER (reaction-aware multimodal enzyme function representation model).

Active03 months ago

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active04 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active10K4 months ago
Python

GENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.

Active214 months ago
Python
Active954 months ago
Python
Active824 months ago
Python

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Active1475 months ago
Python

A patient-level disease classification model trained on single-cell RNA-seq data. Given a matrix of gene expression profiles (one row per cell), the model produces a disease-category prediction for the patient.

Active765 months ago
Python

A 50% magnitude-pruned version of facebook/esm2t1235MUR50D optimized for efficient drug discovery inference on Apple Silicon.

Active425 months ago
Python

DrugCLIP is a dual-encoder multimodal model (SchNet 3D Graph Neural Network + DistilBERT Text Encoder) mapped to a shared 128-dimensional latent space. It is designed to evaluate and retrieve novel 3D molecular structures by aligning them with natural language therapeutic intents and clinical…

Active05 months ago

Github | Cite

Idle306 months ago

Github | Cite

Idle326 months ago

RNAElectra is a nucleotide-resolution RNA language model trained using an ELECTRA-style objective for efficient and discriminative representation learning. The model produces contextualized embeddings for RNA sequences and is designed for downstream transcriptomic and regulatory modeling tasks.

Idle23K7 months ago
Python

BulkRNABert is a transformer-based, encoder-only language model pre-trained on bulk RNA-seq profiles from the TCGA dataset using self-supervised masked language modeling, following the original BERT framework. The model is trained to reconstruct randomly masked gene expression values from their…

Idle26911 months ago
Python

GeneJEPA is a Joint-Embedding Predictive Architecture (JEPA) trained for self-supervised representation learning on scRNA-seq. It uses a Perceiver-style encoder to handle sparse, high-dimensional gene count vectors and a Fourier-feature tokenizer for numerical tokenization.

Idle011 months ago

> A CMR-report contrastive model combining Vision Transformers and pretrained text encoders.

Idle141 year ago
Idle3.5K1 year ago
Python

This repository contains pre-trained models from RadImageNet, a large-scale radiologic image dataset designed to facilitate transfer learning for medical imaging applications.

Idle01 year ago

Welcome to IBM's series of large foundation models for sustainable materials. Our models span a variety of representations and modalities, including SMILES, SELFIES, 3D atom positions, 3D density grids, molecular graphs, and other formats.

Idle931 year ago
Python
Idle3.9K1 year ago
Python

> [!IMPORTANT] > 🎉 Check out the latest version of Phikon here: Phikon-v2 > > Phikon is a self-supervised learning model for histopathology trained with iBOT.

Idle19.8K1 year ago
Python

Model Size: 7B

Idle201 year ago

Model Size: 8B (English)

Idle561 year ago

한국어 모델을 이용한 SapBERT(Self-alignment pretraining for BERT)입니다. 한·영 의료 용어 사전인 KOSTOM을 사용해 한국어 용어와 영어 용어를 정렬했습니다. 참고: SapBERT, Original Code

Idle191 year ago

## Quick Start ```Python from transformers import AutoTokenizer, AutoModel

Stale382 years ago
Python

A Vision Transformer (ViT) image classification model. \ Trained on 15M histology patches from PAIP and TCGA. \ Used the MoCo v3 self supervised learning method.

Stale192 years ago
Python