Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source(1)
Type(1)
677 of 6,573 resources
Showing 151–200
weblab-LLM-M/AscleLM-1-10B
by weblab-LLM-MBabajaan/KAU-BioMedLLM
by BabajaanKAU-BioMedLLM is a research prototype for source-grounded biomedical variant interpretation. The current public release contains the LoRA adapter and documentation for a guarded report-generation system built around a curated biomedical evidence panel, citation enforcement, and abstention when…
arcinstitute/Stack-Large
by arcinstituteStack is a large-scale encoder-decoder foundation model for single-cell biology. It introduces a novel tabular attention architecture that enables both intra- and inter-cellular information flow, setting cell-by-gene matrix chunks as the basic input data unit.
divyeshkamalanaban/nucleotide-transformer-2.5b-multi-species-NF4-Q4
by divyeshkamalanabanThis model is an NF4 (Normal Float 4-bit) quantized version of the base model InstaDeepAI/nucleotide-transformer-2.5b-multi-species. The checkpoint was quantized using the BitsAndBytes library with double quantization enabled and BF16 computation.
yuhtong/DNABERT-S-binferno
by yuhtong> [!WARNING] > This is a model trained on publicly available data. While we've done our best to curate the data, the model performance can still improve. Proceed with caution.
## Important Notice If you are using GENERator for sequence generation, please ensure that the length of each input sequence is a multiple of 6. This can be achieved by either: 1. Padding the sequence on the left with 'A' (left padding); 2. Truncating the sequence from the left (left truncation).
tahoebio/Rhaister
by tahoebioBack to basics: Observed statistics are sufficient to predict drug responses
tokyotech-llm/Medical-GPT-OSS-Swallow-120B
by tokyotech-llmMedical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
tokyotech-llm/Medical-Qwen3-Swallow-30B-A3B
by tokyotech-llmMedical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
moritztng/boltz-2
by moritztngMirror of the Boltz-2 structure- and affinity-prediction weights, packaged for use with tt-bio on Tenstorrent hardware. The files are byte-for-byte identical to the upstream Boltz-2 release; this repo simply hosts them on the Hugging Face Hub so tt-bio can fetch them with huggingface_hub like every…
This repository contains an MLX packaging of OpenMed/OpenMed-PII-ClinicalE5-Small-33M-v1 for Apple Silicon inference with OpenMed.
qwen35-9b-medical is an Ollama/GGUF medical assistant profile based on Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF, distributed locally through the Ollama model kwangsuklee/Qwen3.5-9B.Q4KM-Claude-4.6-Opus-Reasoning-Distilled-v2.
RomeroLab-Duke/prism-antibody
by RomeroLab-DukePRISM is an antibody language model that jointly predicts amino acid identity and germline/non-germline (GL/NGL) position classification, enabling developability-aware antibody sequence modeling.
the-matter-lab/clari
by the-matter-labThis repository contains data and checkpoints for the paper: Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching (arXiv).
paradoxdan/nano-scGPT
by paradoxdan# nano-scGPT The simplest, fastest repository for scGPT inference, (soon) finetuning and trianing, with minimal dependencies. It reimplements the original scGPT from scratch. nanoscgpt/model.py is pure PyTorch in ~270 lines of code, and nanoscgpt/scGPT_tokenizer.py turns raw scRNA data into model…
mradermacher/Phi-4-Instruct-Bioaligned-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
# Instruction For more information, visit our GitHub repository: https://github.com/medfound/medfound
This repository contains GGUF files for gemma4-12b-bioinfo, a fine-tuned Gemma 4 12B model for bioinformatics and computational biology.
gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.
modelid = "DuanYi/R3LMHepG2"
Bioaligned/Phi-4-Instruct-Bioaligned
by BioalignedA merged (ready-to-use) version of microsoft/phi-4 fine-tuned for biological R&D reasoning via QLoRA and evaluated on the Bioalignment Benchmark.
*GenerRNA is a generative pre-trained language model for de novo RNA sequence design. It is a Transformer (decoder-only, GPT-style) model that learns the "language" of RNA from millions of natural sequences and can generate novel, realistic RNA sequences without any structural input, functional…
biohub/esm3-sm-open-v1
by biohubesm3-sm-open-v1 is trained on 2.78 billion natural proteins. With synthetic data augmentation, this led to 3.15 billion protein sequences, 236 million protein structures, and 539 million proteins with function annotations, totaling 771 billion tokens.
Edoardo-BS/HuBERT-ECG-SFT-CardioLearning-large
by Edoardo-BSOriginal code at (https://github.com/Edoar-do/HuBERT-ECG)
Original code at https://github.com/Edoar-do/HuBERT-ECG
Original code at https://github.com/Edoar-do/HuBERT-ECG
Original code at https://github.com/Edoar-do/HuBERT-ECG
genzeonplatform/cliniguard-vitals-ner
by genzeonplatformCliniGuard Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.
genzeonplatform/cliniguard-ner
by genzeonplatformCliniGuard NER is a clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in clinical text.
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
biohub/esmc-600m-2024-12
by biohubThis set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.
biohub/ESMC-6B
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
biohub/ESMC-600M
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
biohub/ESMC-300M
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
A Chinese medical reasoning model fine-tuned from Qwen3.5-4B using a two-stage training pipeline: Supervised Fine-Tuning (SFT) for format alignment, followed by Group Sequence Policy Optimization (GSPO) with an LLM-as-Judge reward function.
PhysicsWallahAI/Aryabhata-2.0
by PhysicsWallahAIAryabhata 2 is a reasoning-focused language model developed by PhysicsWallah for competitive STEM examinations (JEE, NEET). It is obtained by post-training GPT-OSS-20B via reinforcement learning on a curated curriculum of Physics, Chemistry, Mathematics, and General Reasoning questions — achieving…
FINAL-Bench/Darwin-218B-Delphi
by FINAL-Bench> VIDRAFT FINAL-Bench — chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.
AIRI-Institute/genatator-pipeline
by AIRI-InstituteGENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.
HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.
## Model Description ProtGPT3-112M is a single-sequence autoregressive protein language model for protein sequence generation. It is the smallest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.