Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
674 of 6,569 resources
Showing 251–300
InstaDeepAI/instanovo-phospho-v1.0.0
by InstaDeepAIInstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.
InstaDeepAI/instanovo-v1.0.0
by InstaDeepAI# InstaNovo: De novo Peptide Sequencing Model ## Model Description
InstaDeepAI/instanovo-v1.1.0
by InstaDeepAI# InstaNovo: De novo Peptide Sequencing Model ## Model Description
jackxinning/Leanly_AI
by jackxinningharshitsiwach/qwen3.5-0.8b-peptide-steroid
by harshitsiwachThis model is a fine-tuned version of Qwen 3.5 0.8B on a specialized dataset covering biochemistry, peptides, and steroids. It is optimized for providing detailed information on compound mechanisms, dosage (including gender-specific considerations), cycle planning, and physiological effects.
A native MLX port of OpenMed/privacy-filter-nemotron, affine-quantized to 8-bit for fast on-device PII detection on Apple Silicon. For the unquantized BF16 reference, see OpenMed/privacy-filter-nemotron-mlx.
LexBwmn/ACE-V1
by LexBwmn# ACE-V1.1: Brain Tumor Detection !Python!Format > [!CAUTION] > MEDICAL RESEARCH USE ONLY. ACE-V1.1 is NOT a cleared medical device. It must not be used for primary diagnosis or clinical decision-making. All outputs must be verified by a qualified professional.
SongKun909/Qwen2.5-7B-Battery-Expert-LoRA
by SongKun909## Introduction (简介) This model is a domain-specific expert fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA (Low-Rank Adaptation). It is specifically designed for Fine-grained Information Extraction (IE) of technical indicator quintuples from highly complex lithium-ion battery patents.
mradermacher/Qwopus3.5-27B-v3.5-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
microsoft/skala-1.0
by microsoftIn pursuit of the universal functional for density functional theory (DFT), the OneDFT team from Microsoft Research AI for Science has developed the Skala-1.0 exchange-correlation functional, as introduced in Accurate and scalable exchange-correlation with deep learning (arXiv v5), Luise et al.
ZhejiangLab/OneGenome-Rice
by ZhejiangLabOGR is a foundational model for AI-driven precision breeding and functional genomics in rice. It is a generative genomic foundation model trained to process DNA sequences up to 1 million base pairs in length, with 1.25B total parameters and a Mixture-of-Experts (MoE) architecture.
ByteDance-Seed/byteff2
by ByteDance-SeedThis repository contains the model used for the paper Bridging Quantum Mechanics to Organic Liquid Properties via a Universal Force Field。
A domain-optimized reasoning model built on DeepSeek-R1-Distill-Qwen-32B, refined through a multi-stage pipeline of GPTQ quantization-aware training and QLoRA fine-tuning. Achieves 84% on MedQA — within 4 points of GPT-4o — in a ~20GB package that fits on a single L40/L40s GPU.
AIRI-Institute/moderngena-base
by AIRI-Institute# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…
hussenmi/scimilarity_expanded_model
by hussenmiAn extended version of SCimilarity, a metric-learning model for single-cell RNA-seq that maps cells to a unified 128-dimensional embedding space. The original model and method are described in:
Duchifat-2.3-Instruct is a state-of-the-art, instruction-tuned Large Language Model developed by TopAI. As the flagship of the Duchifat series, this model represents a fundamental breakthrough in how Hebrew is processed, reasoned, and generated in the LLM era.
valencelabs/mars-fm
by valencelabsThis repository contains the PyTorch model weights for MarS-FM (Markov Space Flow Matching) trained on the MD-CATH dataset. This model was introduced in the ICLR 2026 paper: MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models.
westlake-repl/Evolla-10B
by westlake-replA frontier protein-language generative model — because proteins deserve better small talk.
IQuestLab/IQuest-UBio-MolFM-V1
by IQuestLabUBio-MolFM is a foundation model suite for molecular modeling, specifically designed for bio-systems. This model, UBio-MolFM-V1 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer architecture. Refer to the technique report for more details: UBio-MolFM (arXiv:2602.17709).
Fine-tuned version of google/gemma-4-E4B-it across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.
pregH/MolecularDiffusion
by pregHlearning-unit/L1-16B-A3B
by learning-unitL1 (Learning Unit 1) is the first language model from Lunit and Lunit Consortium, purpose-built for the medical domain. Derived from Gravity-16B-A3B-Base, L1 is designed for clinical reasoning and decision support.
hugging-science/breast-cancer-detector-2
by hugging-science> Note: This checkpoint was donated to Huggingface-science to support open medical AI research
alegendaryfish/CodonTranslator
by alegendaryfishCodonTranslator is a protein-conditioned codon sequence generation model trained on the representative-only data_v3 release.
westlake-repl/Evolla-10B-hf
by westlake-replA frontier protein-language generative model — because proteins deserve better small talk.
epicmajorman/Gemma4-Biomedical-E4B-gguf
by epicmajormanA specialized biomedical AI assistant created by Major Grant, built on Google's Gemma 4 E4B foundation with OpenMed training data. GGUF format for efficient local inference.
## Model Description This is a lightweight, high-performance image classification model built to diagnose histopathological scans of lung and colon tissues. This model was specifically designed for rapid web deployment without sacrificing clinical accuracy.
Prior-Labs/tabpfn_2_6
by Prior-Labs### Model Overview TabPFN-2.6 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.
AllenInstitute/DNA-sequence-model
by AllenInstituteThis repository contains DNA-sequence modeling resources associated with the basal ganglia (BG) cell atlas package. It serves as a centralized entry point for sequence-based regulatory analyses across multiple companion studies.
🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning
wanglab/bioreason-pro-sft
by wanglab🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning
Prior-Labs/tabpfn_2_5
by Prior-Labs### Model Overview TabPFN-2.5 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.
docling-project/MarkushGrapher-2
by docling-projectMarkushGrapher-2 is an end-to-end multimodal model for recognizing chemical structures from patent document images. It jointly encodes vision, text, and layout information to convert Markush structure images into machine-readable CXSMILES representations.
Verdugie/STEM-Oracle-27B
by Verdugie# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.
GENTEL-Lab/EVA
by GENTEL-LabEVA is a generative foundation model for universal RNA modeling and design, trained on OpenRNA v1 — a curated atlas of 114 million full-length RNA sequences spanning all domains of life.
docling-project/ChemicalOCR
by docling-projectChemicalOCR is a compact vision-language model fine-tuned specifically for optical character recognition (OCR) in chemical structure images. It extracts text and bounding boxes from molecular drawings, enabling the recognition of atom labels, abbreviations, and descriptive text within chemical…
Fine-tuned BGE-M3 on Chinese medical question-answer retrieval using hard negative mining and triple-path InfoNCE loss (dense + sparse + ColBERT).
UniParser/MolDetv2
by UniParserCompared to MolDet, our new MolDetv2 model leverages more manually annotated training data, with further optimizations specifically for reducing molecular false detections and improving bounding box regression, achieving stronger performance with a smaller model.
UniParser/MolDetv2-YOLO26
by UniParserThis repository provides the YOLO26-based version of MolDetv2 model.
Xaira-Therapeutics/X-Cell
by Xaira-TherapeuticsA diffusion language model for genome-scale perturbation prediction across diverse cellular contexts.
ibm-research/trajcast.models-arxiv2025
by ibm-researchThis repository comprises a collection of TrajCast models, a framework for forecasting molecular dynamics (MD) trajectories using autoregressive equivariant message-passing networks. Provided with a starting configuration comprising information about atom types, atomic positions, and velocities,…
ClinicDx1/ClinicDx
by ClinicDx1ClinicDx V1 is a fine-tuned multimodal clinical decision support (CDS) model based on google/medgemma-4b-it. It is trained to generate structured, evidence-grounded clinical assessments from patient presentations, integrating a retrieval-augmented knowledge base (KB) pipeline and an audio input…