Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

373 of 6,565 resources

Showing 101150

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active1122 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3.1K2 months ago
Python

Original code at https://github.com/Edoar-do/HuBERT-ECG

Active3482 months ago
Python

CliniGuard Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.

Active72 months ago
Python

CliniGuard NER is a clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in clinical text.

Active32 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active02 months ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active10K2 months ago
Python

This set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.

Active2.3K2 months ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active1.6M2 months ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active11.2K2 months ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active160K2 months ago
Python

![Language: English]()

Active1.3K2 months ago
Python

> VIDRAFT FINAL-Bench — chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.

Active632 months ago
Python

GENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.

Active212 months ago
Python

HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.

Active312 months ago
Python

## Model Description ProtGPT3-112M is a single-sequence autoregressive protein language model for protein sequence generation. It is the smallest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active2.6K2 months ago
Python

ProtGPT3-10B is a single-sequence autoregressive protein language model for protein sequence generation. It is the largest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.

Active992 months ago
Python

!Protein-ligand interaction header

Active62 months ago
Python

A HuggingFace-compatible repackaging of PlasmidGPT (Shao, 2024) — a GPT-2-style decoder pretrained on 153k engineered plasmid sequences from Addgene. Loadable with standard AutoModelForCausalLM and AutoTokenizer. Used as the base for PlasmidGPT-SFT and PlasmidGPT-GRPO.

Active1022 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active9672 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active1832 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Active4722 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5K2 months ago
Python

This set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.

Active6.2K2 months ago
Python

# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.

Active4.3K2 months ago
Python
Active166.9K2 months ago
Python
Active982 months ago
Python

Modelo fine-tunado com LoRA (MLX / Apple Silicon) para assistência clínica em saúde da mulher.

Active1253 months ago
Python
Active823 months ago
Python

This is the pretrained VetBERT model from the github repo: https://github.com/havocy28/VetBERT

Active313 months ago
Python

Target-Conditioned Molecular Ideation Model for Drug Discovery Research

Active03 months ago
Python

Genos-m is a foundation model for human-associated microbial genomes. It is trained to model microbial DNA sequences at single-nucleotide resolution and supports ultra-long genomic contexts up to one million tokens.

Active313 months ago
Python

Qwen3-8B-syco_med-gated-attention-FT is a plug-and-play gated attention weight released for AI safety research.

Active03 months ago
Python

LoRA fine-tune of Qwen3.5-9B on synthetic clinical triage Q&A pairs generated from PubMed Central open-access papers. The model is specialized for emergency-medicine decision-making: triaging patients, applying clinical decision rules, and generating protocol-grounded triage recommendations.

Active03 months ago
Python

Gemma 4 E2B fine-tuned on 225K drug–target pairs for novel small-molecule generation.

Active253 months ago
Python

- 2025-05-15: We identified a bug in the Bacformer Large code on HuggingFace which resulted in a significant drop in the quality of the output embeddings. This is now fixed, but if you downloaded or cached the model before this date, re-download and use the latest model revision before running…

Active8K3 months ago
Python

- 2025-05-15: We identified a bug in the Bacformer Large code on HuggingFace which resulted in a significant drop in the quality of the output embeddings. This is now fixed, but if you downloaded or cached the model before this date, re-download and use the latest model revision before running…

Active5473 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active7033 months ago
Python

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Active1473 months ago
Python

A patient-level disease classification model trained on single-cell RNA-seq data. Given a matrix of gene expression profiles (one row per cell), the model produces a disease-category prediction for the patient.

Active763 months ago
Python

MedPsy-4B is a state-of-the-art, text-only medical and healthcare language model purpose-built for edge deployment. Built on top of Qwen3-4B-Thinking-2507 and post-trained with a multi-stage pipeline (supervised fine-tuning + reinforcement learning) on curated medical data, it surpasses models…

Active1.5K3 months ago
Python

MedPsy-1.7B is a state-of-the-art, text-only medical and healthcare language model purpose-built for edge and smartphone deployment. Built on top of Qwen3-1.7B (operated in thinking mode, i.e. with enable_thinking=True) and post-trained with a multi-stage pipeline (supervised fine-tuning +…

Active6083 months ago
Python

## Introduction (简介) This model is a domain-specific expert fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA (Low-Rank Adaptation). It is specifically designed for Fine-grained Information Extraction (IE) of technical indicator quintuples from highly complex lithium-ion battery patents.

Active103 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active634 months ago
Python

Duchifat-2.3-Instruct is a state-of-the-art, instruction-tuned Large Language Model developed by TopAI. As the flagship of the Duchifat series, this model represents a fundamental breakthrough in how Hebrew is processed, reasoned, and generated in the LLM era.

Active1544 months ago
Python

!image

Active6014 months ago
Python
Active255.2K4 months ago
Python