Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

674 of 6,569 resources

Showing 251300

InstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.

Active113 months ago

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

Active143 months ago

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

Active143 months ago
Active6.2K3 months ago

PickyBinders/tea

by PickyBinders

!Model Architecture

Active142.7K3 months ago

This model is a fine-tuned version of Qwen 3.5 0.8B on a specialized dataset covering biochemistry, peptides, and steroids. It is optimized for providing detailed information on compound mechanisms, dosage (including gender-specific considerations), cycle planning, and physiological effects.

Active6483 months ago

FitCareer_AI Introduction

Active03 months ago

A native MLX port of OpenMed/privacy-filter-nemotron, affine-quantized to 8-bit for fast on-device PII detection on Apple Silicon. For the unquantized BF16 reference, see OpenMed/privacy-filter-nemotron-mlx.

Active2.2K3 months ago

# ACE-V1.1: Brain Tumor Detection !Python!Format > [!CAUTION] > MEDICAL RESEARCH USE ONLY. ACE-V1.1 is NOT a cleared medical device. It must not be used for primary diagnosis or clinical decision-making. All outputs must be verified by a qualified professional.

Active03 months ago

## Introduction (简介) This model is a domain-specific expert fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA (Low-Rank Adaptation). It is specifically designed for Fine-grained Information Extraction (IE) of technical indicator quintuples from highly complex lithium-ion battery patents.

Active103 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active634 months ago
Python

In pursuit of the universal functional for density functional theory (DFT), the OneDFT team from Microsoft Research AI for Science has developed the Skala-1.0 exchange-correlation functional, as introduced in Accurate and scalable exchange-correlation with deep learning (arXiv v5), Luise et al.

Active6.5K4 months ago

OGR is a foundational model for AI-driven precision breeding and functional genomics in rice. It is a generative genomic foundation model trained to process DNA sequences up to 1 million base pairs in length, with 1.25B total parameters and a Mixture-of-Experts (MoE) architecture.

Active134 months ago

This repository contains the model used for the paper Bridging Quantum Mechanics to Organic Liquid Properties via a Universal Force Field。

Active04 months ago

A domain-optimized reasoning model built on DeepSeek-R1-Distill-Qwen-32B, refined through a multi-stage pipeline of GPTQ quantization-aware training and QLoRA fine-tuning. Achieves 84% on MedQA — within 4 points of GPT-4o — in a ~20GB package that fits on a single L40/L40s GPU.

Active1134 months ago

# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…

Active4284 months ago

An extended version of SCimilarity, a metric-learning model for single-cell RNA-seq that maps cells to a unified 128-dimensional embedding space. The original model and method are described in:

Active04 months ago

Duchifat-2.3-Instruct is a state-of-the-art, instruction-tuned Large Language Model developed by TopAI. As the flagship of the Duchifat series, this model represents a fundamental breakthrough in how Hebrew is processed, reasoned, and generated in the LLM era.

Active1544 months ago
Python

This repository contains the PyTorch model weights for MarS-FM (Markov Space Flow Matching) trained on the MD-CATH dataset. This model was introduced in the ICLR 2026 paper: MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models.

Active04 months ago

!image

Active6014 months ago
Python

A frontier protein-language generative model — because proteins deserve better small talk.

Active164 months ago
Active245.4K4 months ago
Python

UBio-MolFM is a foundation model suite for molecular modeling, specifically designed for bio-systems. This model, UBio-MolFM-V1 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer architecture. Refer to the technique report for more details: UBio-MolFM (arXiv:2602.17709).

Active54 months ago

Fine-tuned version of google/gemma-4-E4B-it across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.

Active1K4 months ago
Python
Active04 months ago

L1 (Learning Unit 1) is the first language model from Lunit and Lunit Consortium, purpose-built for the medical domain. Derived from Gravity-16B-A3B-Base, L1 is designed for clinical reasoning and decision support.

Active844 months ago
Python

> Note: This checkpoint was donated to Huggingface-science to support open medical AI research

Active224 months ago

CodonTranslator is a protein-conditioned codon sequence generation model trained on the representative-only data_v3 release.

Active04 months ago

A frontier protein-language generative model — because proteins deserve better small talk.

Active102.3K4 months ago

!Banner.

Active5604 months ago

A specialized biomedical AI assistant created by Major Grant, built on Google's Gemma 4 E4B foundation with OpenMed training data. GGUF format for efficient local inference.

Active2084 months ago

## Model Description This is a lightweight, high-performance image classification model built to diagnose histopathological scans of lung and colon tissues. This model was specifically designed for rapid web deployment without sacrificing clinical accuracy.

Active44 months ago
Python

### Model Overview TabPFN-2.6 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.

Active7564 months ago

This repository contains DNA-sequence modeling resources associated with the basal ganglia (BG) cell atlas package. It serves as a centralized entry point for sequence-based regulatory analyses across multiple companion studies.

Active04 months ago

🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning

Active764 months ago

🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning

Active504 months ago

### Model Overview TabPFN-2.5 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.

Active8.5K5 months ago

MarkushGrapher-2 is an end-to-end multimodal model for recognizing chemical structures from patent document images. It jointly encodes vision, text, and layout information to convert Markush structure images into machine-readable CXSMILES representations.

Active1735 months ago
Python

# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.

Active1425 months ago
Python

EVA is a generative foundation model for universal RNA modeling and design, trained on OpenRNA v1 — a curated atlas of 114 million full-length RNA sequences spanning all domains of life.

Active05 months ago

ChemicalOCR is a compact vision-language model fine-tuned specifically for optical character recognition (OCR) in chemical structure images. It extracts text and bounding boxes from molecular drawings, enabling the recognition of atom labels, abbreviations, and descriptive text within chemical…

Active4715 months ago
Python

Github | Cite

Active95 months ago

Github | Cite

Active125 months ago

Fine-tuned BGE-M3 on Chinese medical question-answer retrieval using hard negative mining and triple-path InfoNCE loss (dense + sparse + ColBERT).

Active05 months ago

Compared to MolDet, our new MolDetv2 model leverages more manually annotated training data, with further optimizations specifically for reducing molecular false detections and improving bounding box regression, achieving stronger performance with a smaller model.

Active05 months ago

This repository provides the YOLO26-based version of MolDetv2 model.

Active05 months ago

A diffusion language model for genome-scale perturbation prediction across diverse cellular contexts.

Active05 months ago

This repository comprises a collection of TrajCast models, a framework for forecasting molecular dynamics (MD) trajectories using autoregressive equivariant message-passing networks. Provided with a starting configuration comprising information about atom types, atomic positions, and velocities,…

Active1055 months ago

ClinicDx V1 is a fine-tuned multimodal clinical decision support (CDS) model based on google/medgemma-4b-it. It is trained to generate structured, evidence-grounded clinical assessments from patient presentations, integrating a retrieval-augmented knowledge base (KB) pipeline and an audio input…

Active455 months ago