Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language(1)
License
Source
Type
36 of 7,068 resources
pantoniadis/RAGenome
by pantoniadisPaper: RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts
athanzli/MicroGlot
by athanzliA taxonomy-informed sparse DNA foundation model for microbial genomics.
Three genomic foundation models, packaged together for local inference on Apple silicon.
Gaolaboratory/iona-base-400m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.
Gaolaboratory/iona-base-200m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.
Gaolaboratory/iona-base-100m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.
Gaolaboratory/iona-base-50m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
Manhph2211/ECG-Scan
by Manhph2211JerrryNie/ConceptCLIP
by JerrryNiezeroentropy/zembed-1-embedding
by zeroentropyIn retrieval systems, embedding models determine the quality of your search.
ibm-research/MoLFormer-XL-both-10pct
by ibm-researchMoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.
EthanGao123/CellHermes-v1.0
by EthanGao123# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language models…
PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.
raidium/Jolia
by raidiumJolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
AIRI-Institute/genatator-pipeline
by AIRI-InstituteGENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.
ScientaLab/eva-rna
by ScientaLabManhph2211/D-BETA
by Manhph2211MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.
ConvergeBio/virtual-cell-patient
by ConvergeBioA patient-level disease classification model trained on single-cell RNA-seq data. Given a matrix of gene expression profiles (one row per cell), the model produces a disease-category prediction for the patient.
stephenjun8192/esm2-35m-sparse50
by stephenjun8192A 50% magnitude-pruned version of facebook/esm2t1235MUR50D optimized for efficient drug discovery inference on Apple Silicon.
FreakingPotato/RNAElectra
by FreakingPotatoRNAElectra is a nucleotide-resolution RNA language model trained using an ELECTRA-style objective for efficient and discriminative representation learning. The model produces contextualized embeddings for RNA sequences and is designed for downstream transcriptomic and regulatory modeling tasks.
InstaDeepAI/BulkRNABert
by InstaDeepAIBulkRNABert is a transformer-based, encoder-only language model pre-trained on bulk RNA-seq profiles from the TCGA dataset using self-supervised masked language modeling, following the original BERT framework. The model is trained to reconstruct randomly masked gene expression values from their…
helical-ai/helix-mRNA
by helical-aiibm-research/materials.smi-ted
by ibm-researchWelcome to IBM's series of large foundation models for sustainable materials. Our models span a variety of representations and modalities, including SMILES, SELFIES, 3D atom positions, 3D density grids, molecular graphs, and other formats.
zhihan1996/DNABERT-S
by zhihan1996> [!IMPORTANT] > 🎉 Check out the latest version of Phikon here: Phikon-v2 > > Phikon is a self-supervised learning model for histopathology trained with iBOT.
qiuhuachuan/PsyChat
by qiuhuachuan## Quick Start ```Python from transformers import AutoTokenizer, AutoModel
A Vision Transformer (ViT) image classification model. \ Trained on 15M histology patches from PAIP and TCGA. \ Used the MoCo v3 self supervised learning method.
A Vision Transformer (ViT) image classification model. \ Trained by Owkin on 40 million pan-cancer histology tiles from TCGA-COAD.
A Vision Transformer (ViT) image classification model. \ Trained on 2M histology patches from TCGA-BRCA.
datasets: - UMLS