Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

887 of 7,078 resources

Showing 301–350

Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Active153 months ago
Python

Medical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Active2403 months ago
Python

!Pretrainloss

Active03 months ago

Mirror of the Boltz-2 structure- and affinity-prediction weights, packaged for use with tt-bio on Tenstorrent hardware. The files are byte-for-byte identical to the upstream Boltz-2 release; this repo simply hosts them on the Hugging Face Hub so tt-bio can fetch them with huggingface_hub like every…

Active03 months ago

This repository contains an MLX packaging of OpenMed/OpenMed-PII-ClinicalE5-Small-33M-v1 for Apple Silicon inference with OpenMed.

Active3054 months ago

qwen35-9b-medical is an Ollama/GGUF medical assistant profile based on Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF, distributed locally through the Ollama model kwangsuklee/Qwen3.5-9B.Q4KM-Claude-4.6-Opus-Reasoning-Distilled-v2.

Active764 months ago
C

As part of the ENCODE 4 Project, we trained BPNet models on 2,339 ENCODE transcription factor ChIP-seq experiments spanning 788 targets across 175 biosamples. Here, we provide all models for open-source use.

Active74 months ago

CoralBay: A Self-Supervised CT Foundation Model

Active724 months ago

PRISM is an antibody language model that jointly predicts amino acid identity and germline/non-germline (GL/NGL) position classification, enabling developability-aware antibody sequence modeling.

Active04 months ago

the-matter-lab/clari

by the-matter-lab

This repository contains data and checkpoints for the paper: Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching (arXiv).

Active04 months ago

JAX/Equinox parameters for Promera, a dual-purpose biomolecular generative model for structure prediction and binder design. These weights were converted module-by-module from the official PyTorch checkpoint bjing-mit/promera (promera_2606.ckpt) and validated numerically against it (per-module…

Active04 months ago

# nano-scGPT The simplest, fastest repository for scGPT inference, (soon) finetuning and trianing, with minimal dependencies. It reimplements the original scGPT from scratch. nanoscgpt/model.py is pure PyTorch in ~270 lines of code, and nanoscgpt/scGPT_tokenizer.py turns raw scRNA data into model…

Active464 months ago

For a convenient overview and download list, visit our model page for this model.

Active4014 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active7064 months ago
Python

# Instruction For more information, visit our GitHub repository: https://github.com/medfound/medfound

Active9004 months ago

This repository contains GGUF files for gemma4-12b-bioinfo, a fine-tuned Gemma 4 12B model for bioinformatics and computational biology.

Active884 months ago
C

gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.

Active484 months ago
Python

modelid = "DuanYi/R3LMHepG2"

Active104 months ago
Python

A merged (ready-to-use) version of microsoft/phi-4 fine-tuned for biological R&D reasoning via QLoRA and evaluated on the Bioalignment Benchmark.

Active254 months ago

*GenerRNA is a generative pre-trained language model for de novo RNA sequence design. It is a Transformer (decoder-only, GPT-style) model that learns the "language" of RNA from millions of natural sequences and can generate novel, realistic RNA sequences without any structural input, functional…

Active04 months ago
Active124 months ago
Python

esm3-sm-open-v1 is trained on 2.78 billion natural proteins. With synthetic data augmentation, this led to 3.15 billion protein sequences, 236 million protein structures, and 539 million proteins with function annotations, totaling 771 billion tokens.

Active2.6K4 months ago
Python

Original code at (https://github.com/Edoar-do/HuBERT-ECG)

Active654 months ago
Python

Original code at (https://github.com/Edoar-do/HuBERT-ECG)

Active444 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active1124 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3.1K4 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3484 months ago
Python

CliniGuard Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.

Active74 months ago
Python

CliniGuard NER is a clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in clinical text.

Active34 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active04 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active10K4 months ago
Python

This set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.

Active1.1K4 months ago
Python

![Language: English]()

Active1.3K4 months ago
Python

A Chinese medical reasoning model fine-tuned from Qwen3.5-4B using a two-stage training pipeline: Supervised Fine-Tuning (SFT) for format alignment, followed by Group Sequence Policy Optimization (GSPO) with an LLM-as-Judge reward function.

Active6.2K4 months ago

Aryabhata 2 is a reasoning-focused language model developed by PhysicsWallah for competitive STEM examinations (JEE, NEET). It is obtained by post-training GPT-OSS-20B via reinforcement learning on a curated curriculum of Physics, Chemistry, Mathematics, and General Reasoning questions — achieving…

Active1974 months ago

> VIDRAFT FINAL-Bench — chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.

Active634 months ago
Python

GENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.

Active214 months ago
Python

HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.

Active314 months ago
Python

## Model Description ProtGPT3-112M is a single-sequence autoregressive protein language model for protein sequence generation. It is the smallest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active2.6K4 months ago
Python

ProtGPT3-10B is a single-sequence autoregressive protein language model for protein sequence generation. It is the largest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active994 months ago
Python

This repository hosts release artifacts for ReCLIP:

Active04 months ago

!Protein-ligand interaction header

Active64 months ago
Python

GGUF quantizations of HuggingFaceBio/Carbon-3B — a generative DNA foundation model — for efficient inference with llama.cpp.

Active6204 months ago

A HuggingFace-compatible repackaging of PlasmidGPT (Shao, 2024) — a GPT-2-style decoder pretrained on 153k engineered plasmid sequences from Addgene. Loadable with standard AutoModelForCausalLM and AutoTokenizer. Used as the base for PlasmidGPT-SFT and PlasmidGPT-GRPO.

Active1024 months ago
Python

1. 概述 2. 数据处理流程 3. 模型架构 4. 损失函数设计 5. 训练流程 6. 配置参数

Active04 months ago

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active9674 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active1834 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active4724 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5K4 months ago
Python