Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

881 of 7,064 resources

Showing 301–350

qwen35-9b-medical is an Ollama/GGUF medical assistant profile based on Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF, distributed locally through the Ollama model kwangsuklee/Qwen3.5-9B.Q4KM-Claude-4.6-Opus-Reasoning-Distilled-v2.

Active763 months ago
C

As part of the ENCODE 4 Project, we trained BPNet models on 2,339 ENCODE transcription factor ChIP-seq experiments spanning 788 targets across 175 biosamples. Here, we provide all models for open-source use.

Active73 months ago

CoralBay: A Self-Supervised CT Foundation Model

Active723 months ago

PRISM is an antibody language model that jointly predicts amino acid identity and germline/non-germline (GL/NGL) position classification, enabling developability-aware antibody sequence modeling.

Active03 months ago

the-matter-lab/clari

by the-matter-lab

This repository contains data and checkpoints for the paper: Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching (arXiv).

Active03 months ago

JAX/Equinox parameters for Promera, a dual-purpose biomolecular generative model for structure prediction and binder design. These weights were converted module-by-module from the official PyTorch checkpoint bjing-mit/promera (promera_2606.ckpt) and validated numerically against it (per-module…

Active03 months ago

# nano-scGPT The simplest, fastest repository for scGPT inference, (soon) finetuning and trianing, with minimal dependencies. It reimplements the original scGPT from scratch. nanoscgpt/model.py is pure PyTorch in ~270 lines of code, and nanoscgpt/scGPT_tokenizer.py turns raw scRNA data into model…

Active464 months ago

For a convenient overview and download list, visit our model page for this model.

Active4014 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active7064 months ago
Python

# Instruction For more information, visit our GitHub repository: https://github.com/medfound/medfound

Active9004 months ago

This repository contains GGUF files for gemma4-12b-bioinfo, a fine-tuned Gemma 4 12B model for bioinformatics and computational biology.

Active884 months ago
C

gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.

Active484 months ago
Python

modelid = "DuanYi/R3LMHepG2"

Active104 months ago
Python

A merged (ready-to-use) version of microsoft/phi-4 fine-tuned for biological R&D reasoning via QLoRA and evaluated on the Bioalignment Benchmark.

Active254 months ago

*GenerRNA is a generative pre-trained language model for de novo RNA sequence design. It is a Transformer (decoder-only, GPT-style) model that learns the "language" of RNA from millions of natural sequences and can generate novel, realistic RNA sequences without any structural input, functional…

Active04 months ago
Active124 months ago
Python

esm3-sm-open-v1 is trained on 2.78 billion natural proteins. With synthetic data augmentation, this led to 3.15 billion protein sequences, 236 million protein structures, and 539 million proteins with function annotations, totaling 771 billion tokens.

Active2.6K4 months ago
Python

Original code at (https://github.com/Edoar-do/HuBERT-ECG)

Active654 months ago
Python

Original code at (https://github.com/Edoar-do/HuBERT-ECG)

Active444 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active1124 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3.1K4 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3484 months ago
Python

CliniGuard Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.

Active74 months ago
Python

CliniGuard NER is a clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in clinical text.

Active34 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active04 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active10K4 months ago
Python

This set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.

Active1.1K4 months ago
Python

![Language: English]()

Active1.3K4 months ago
Python

A Chinese medical reasoning model fine-tuned from Qwen3.5-4B using a two-stage training pipeline: Supervised Fine-Tuning (SFT) for format alignment, followed by Group Sequence Policy Optimization (GSPO) with an LLM-as-Judge reward function.

Active6.2K4 months ago

Aryabhata 2 is a reasoning-focused language model developed by PhysicsWallah for competitive STEM examinations (JEE, NEET). It is obtained by post-training GPT-OSS-20B via reinforcement learning on a curated curriculum of Physics, Chemistry, Mathematics, and General Reasoning questions — achieving…

Active1974 months ago

> VIDRAFT FINAL-Bench — chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.

Active634 months ago
Python

GENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.

Active214 months ago
Python

HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.

Active314 months ago
Python

## Model Description ProtGPT3-112M is a single-sequence autoregressive protein language model for protein sequence generation. It is the smallest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active2.6K4 months ago
Python

ProtGPT3-10B is a single-sequence autoregressive protein language model for protein sequence generation. It is the largest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active994 months ago
Python

This repository hosts release artifacts for ReCLIP:

Active04 months ago

!Protein-ligand interaction header

Active64 months ago
Python

GGUF quantizations of HuggingFaceBio/Carbon-3B — a generative DNA foundation model — for efficient inference with llama.cpp.

Active6204 months ago

A HuggingFace-compatible repackaging of PlasmidGPT (Shao, 2024) — a GPT-2-style decoder pretrained on 153k engineered plasmid sequences from Addgene. Loadable with standard AutoModelForCausalLM and AutoTokenizer. Used as the base for PlasmidGPT-SFT and PlasmidGPT-GRPO.

Active1024 months ago
Python

1. 概述 2. 数据处理流程 3. 模型架构 4. 损失函数设计 5. 训练流程 6. 配置参数

Active04 months ago

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active9674 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active1834 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active4724 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5K4 months ago
Python

A 53K-parameter weight-shared transformer that learns SMILES grammar by applying one small block 8 times. It reaches 95.3% validity on ZINC-250K — outperforming an unshared GPT 10× larger (87.6%).

Active04 months ago

This set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.

Active6.2K4 months ago
Python
Active19.5K4 months ago
Python
Active954 months ago
Python

Sparse Autoencoder (SAE) trained on residue-level embeddings from ESM-2 (650M, layer 33) for interpretability research on protein language models.

Active184 months ago