Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
881 of 7,064 resources
Showing 401–450
ZhejiangLab/OneGenome-Rice
by ZhejiangLabOGR is a foundational model for AI-driven precision breeding and functional genomics in rice. It is a generative genomic foundation model trained to process DNA sequences up to 1 million base pairs in length, with 1.25B total parameters and a Mixture-of-Experts (MoE) architecture.
ByteDance-Seed/byteff2
by ByteDance-SeedThis repository contains the model used for the paper Bridging Quantum Mechanics to Organic Liquid Properties via a Universal Force Field。
A domain-optimized reasoning model built on DeepSeek-R1-Distill-Qwen-32B, refined through a multi-stage pipeline of GPTQ quantization-aware training and QLoRA fine-tuning. Achieves 84% on MedQA — within 4 points of GPT-4o — in a ~20GB package that fits on a single L40/L40s GPU.
AIRI-Institute/moderngena-base
by AIRI-Institute# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…
hussenmi/scimilarity_expanded_model
by hussenmiAn extended version of SCimilarity, a metric-learning model for single-cell RNA-seq that maps cells to a unified 128-dimensional embedding space. The original model and method are described in:
stephenjun8192/esm2-35m-sparse50
by stephenjun8192A 50% magnitude-pruned version of facebook/esm2t1235MUR50D optimized for efficient drug discovery inference on Apple Silicon.
simmani91/GeneLinguaLM-v5
by simmani91GeneLinguaLM is a multimodal model that generates natural language descriptions of protein functions from amino acid sequences.
Duchifat-2.3-Instruct is a state-of-the-art, instruction-tuned Large Language Model developed by TopAI. As the flagship of the Duchifat series, this model represents a fundamental breakthrough in how Hebrew is processed, reasoned, and generated in the LLM era.
valencelabs/mars-fm
by valencelabsThis repository contains the PyTorch model weights for MarS-FM (Markov Space Flow Matching) trained on the MD-CATH dataset. This model was introduced in the ICLR 2026 paper: MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models.
mradermacher/AniMUL-v1-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
A generalist foundation model for healthcare capable of handling diverse medical data modalities.
westlake-repl/Evolla-10B
by westlake-replA frontier protein-language generative model — because proteins deserve better small talk.
homerquan/DrugClip
by homerquanDrugCLIP is a dual-encoder multimodal model (SchNet 3D Graph Neural Network + DistilBERT Text Encoder) mapped to a shared 128-dimensional latent space. It is designed to evaluate and retrieve novel 3D molecular structures by aligning them with natural language therapeutic intents and clinical…
IQuestLab/IQuest-UBio-MolFM-V1
by IQuestLabUBio-MolFM is a foundation model suite for molecular modeling, specifically designed for bio-systems. This model, UBio-MolFM-V1 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer architecture. Refer to the technique report for more details: UBio-MolFM (arXiv:2602.17709).
Fine-tuned version of google/gemma-4-E4B-it across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.
learning-unit/L1-16B-A3B
by learning-unitL1 (Learning Unit 1) is the first language model from Lunit and Lunit Consortium, purpose-built for the medical domain. Derived from Gravity-16B-A3B-Base, L1 is designed for clinical reasoning and decision support.
hugging-science/breast-cancer-detector-2
by hugging-science> Note: This checkpoint was donated to Huggingface-science to support open medical AI research
alegendaryfish/CodonTranslator
by alegendaryfishCodonTranslator is a protein-conditioned codon sequence generation model trained on the representative-only data_v3 release.
westlake-repl/Evolla-10B-hf
by westlake-replA frontier protein-language generative model — because proteins deserve better small talk.
epicmajorman/Gemma4-Biomedical-E4B-gguf
by epicmajormanA specialized biomedical AI assistant created by Major Grant, built on Google's Gemma 4 E4B foundation with OpenMed training data. GGUF format for efficient local inference.
## Model Description This is a lightweight, high-performance image classification model built to diagnose histopathological scans of lung and colon tissues. This model was specifically designed for rapid web deployment without sacrificing clinical accuracy.
AnakinHuang/brainscope-scgpt-disease
by AnakinHuangPrior-Labs/tabpfn_2_6
by Prior-Labs### Model Overview TabPFN-2.6 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.
AllenInstitute/DNA-sequence-model
by AllenInstituteThis repository contains DNA-sequence modeling resources associated with the basal ganglia (BG) cell atlas package. It serves as a centralized entry point for sequence-based regulatory analyses across multiple companion studies.
🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning
wanglab/bioreason-pro-sft
by wanglab🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning
Prior-Labs/tabpfn_2_5
by Prior-Labs### Model Overview TabPFN-2.5 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.
docling-project/MarkushGrapher-2
by docling-projectMarkushGrapher-2 is an end-to-end multimodal model for recognizing chemical structures from patent document images. It jointly encodes vision, text, and layout information to convert Markush structure images into machine-readable CXSMILES representations.
Verdugie/STEM-Oracle-27B
by Verdugie# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.
GENTEL-Lab/EVA
by GENTEL-LabEVA is a generative foundation model for universal RNA modeling and design, trained on OpenRNA v1 — a curated atlas of 114 million full-length RNA sequences spanning all domains of life.
docling-project/ChemicalOCR
by docling-projectChemicalOCR is a compact vision-language model fine-tuned specifically for optical character recognition (OCR) in chemical structure images. It extracts text and bounding boxes from molecular drawings, enabling the recognition of atom labels, abbreviations, and descriptive text within chemical…
alis-sila/mitotic-transformer
by alis-silaA biologically & cosmologically inspired causal language model based on the "Cosmology of the Living Cell" (Mother Theory)
This is a finetuned EVO2 model for chromosome classification, trained for 20 epochs.
Fine-tuned BGE-M3 on Chinese medical question-answer retrieval using hard negative mining and triple-path InfoNCE loss (dense + sparse + ColBERT).
UniParser/MolDetv2-YOLO26
by UniParserThis repository provides the YOLO26-based version of MolDetv2 model.
Xaira-Therapeutics/X-Cell
by Xaira-TherapeuticsA diffusion language model for genome-scale perturbation prediction across diverse cellular contexts.
ibm-research/trajcast.models-arxiv2025
by ibm-researchThis repository comprises a collection of TrajCast models, a framework for forecasting molecular dynamics (MD) trajectories using autoregressive equivariant message-passing networks. Provided with a starting configuration comprising information about atom types, atomic positions, and velocities,…
ClinicDx1/ClinicDx
by ClinicDx1ClinicDx V1 is a fine-tuned multimodal clinical decision support (CDS) model based on google/medgemma-4b-it. It is trained to generate structured, evidence-grounded clinical assessments from patient presentations, integrating a retrieval-augmented knowledge base (KB) pipeline and an audio input…
A Chemprop v2 multi-component MPNN model that predicts 7 spectroscopic properties of organic chromophores from molecular structure (SMILES) and solvent.
changlab/miniMTI-CRC
by changlabjheuschkel/SynCodonLM-V2
by jheuschkel- This repository contains code to utilize the model, and reproduce results of the paper Advancing Codon Language Modeling with Synonymous Codon Constrained Masking. - Unlike other Codon Language Models, SynCodonLM was trained with logit-level control, masking logits for non-synonymous codons.
A PyTorch port of AlphaGenome, the DNA sequence model from Google DeepMind that predicts hundreds of genomic tracks at single base-pair resolution from sequences up to 1M bp.