Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source(1)
Type
477 of 7,078 resources
Showing 151–200
fairydance/molexar-10m-base
by fairydanceMolexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding and…
fairydance/molexar-10m-omni
by fairydanceMolexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,…
iti-visual-analytics/GRamma-12B
by iti-visual-analyticsGRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.
Full weight-level fine-tuning of InstaDeepAI/nucleotide-transformer-v2-50m-multi-species for binary DNA sequence classification on two GenomicBenchmarks tasks. All parameters are updated rather than using LoRA or a frozen backbone, with a leakage-free train/validation/test protocol and multi-seed…
QLoRA adapter for Llama-3.1-8B-Instruct, fine-tuned on PubMedQA for yes / no / maybe biomedical question answering (run5).
BioMatrix is a multimodal biological foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
nikitaredy/medictron-7B
by nikitaredyA domain-adapted clinical LLM fine-tuned on synthetic Indian medical Q&A records using QLoRA (4-bit quantization) with Unsloth 2x speedup. Built to power the conversational AI layer.
ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein sequence generation.
weblab-LLM-M/AscleLM-1-10B
by weblab-LLM-MBabajaan/KAU-BioMedLLM
by BabajaanKAU-BioMedLLM is a research prototype for source-grounded biomedical variant interpretation. The current public release contains the LoRA adapter and documentation for a guarded report-generation system built around a curated biomedical evidence panel, citation enforcement, and abstention when…
HauserGroup/ApeTokenizer-SMILES
by HauserGroupApeTokenizer-SMILES is an Atom Pair Encoding (APE) tokenizer for SMILES strings, trained on ~2M unique canonical SMILES from ChEMBL 36. It is the SMILES counterpart to ApeTokenizer-SELFIES, released alongside ModernMolBERT.
divyeshkamalanaban/nucleotide-transformer-2.5b-multi-species-NF4-Q4
by divyeshkamalanabanThis model is an NF4 (Normal Float 4-bit) quantized version of the base model InstaDeepAI/nucleotide-transformer-2.5b-multi-species. The checkpoint was quantized using the BitsAndBytes library with double quantization enabled and BF16 computation.
## Important Notice If you are using GENERator for sequence generation, please ensure that the length of each input sequence is a multiple of 6. This can be achieved by either: 1. Padding the sequence on the left with 'A' (left padding); 2. Truncating the sequence from the left (left truncation).
tokyotech-llm/Medical-GPT-OSS-Swallow-120B
by tokyotech-llmMedical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
tokyotech-llm/Medical-Qwen3-Swallow-30B-A3B
by tokyotech-llmMedical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
mradermacher/Phi-4-Instruct-Bioaligned-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
mradermacher/gemma4-12b-bioinfo-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.
modelid = "DuanYi/R3LMHepG2"
Wildstash/DentalGPT
by Wildstashbiohub/esm3-sm-open-v1
by biohubesm3-sm-open-v1 is trained on 2.78 billion natural proteins. With synthetic data augmentation, this led to 3.15 billion protein sequences, 236 million protein structures, and 539 million proteins with function annotations, totaling 771 billion tokens.
Edoardo-BS/HuBERT-ECG-SFT-CardioLearning-large
by Edoardo-BSOriginal code at (https://github.com/Edoar-do/HuBERT-ECG)
Edoardo-Coppola/HuBERT-ECG-SFT-CardioLearning-large
by Edoardo-CoppolaOriginal code at (https://github.com/Edoar-do/HuBERT-ECG)
Original code at https://github.com/Edoar-do/HuBERT-ECG
Original code at https://github.com/Edoar-do/HuBERT-ECG
Original code at https://github.com/Edoar-do/HuBERT-ECG
genzeonplatform/cliniguard-vitals-ner
by genzeonplatformCliniGuard Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.
genzeonplatform/cliniguard-ner
by genzeonplatformCliniGuard NER is a clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in clinical text.
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
biohub/esmc-600m-2024-12
by biohubThis set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.
FINAL-Bench/Darwin-218B-Delphi
by FINAL-Bench> VIDRAFT FINAL-Bench — chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.
AIRI-Institute/genatator-pipeline
by AIRI-InstituteGENATATOR-PIPELINE is a Hugging Face pipeline for ab initio gene annotation from genomic DNA. It accepts a FASTA file, finds candidate transcript intervals, assigns transcript type, predicts exon and CDS structure, and writes a GFF3 annotation file.
HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.
## Model Description ProtGPT3-112M is a single-sequence autoregressive protein language model for protein sequence generation. It is the smallest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.
ProtGPT3-10B is a single-sequence autoregressive protein language model for protein sequence generation. It is the largest model in the ProtGPT3 family, an open-source suite of promptable and aligned protein language models ranging from 112M to 10B parameters.
poolside-laguna-hackathon/protein-ligand-design
by poolside-laguna-hackathon!Protein-ligand interaction header
UCL-CSSB/PlasmidGPT
by UCL-CSSBA HuggingFace-compatible repackaging of PlasmidGPT (Shao, 2024) — a GPT-2-style decoder pretrained on 153k engineered plasmid sequences from Addgene. Loadable with standard AutoModelForCausalLM and AutoTokenizer. Used as the base for PlasmidGPT-SFT and PlasmidGPT-GRPO.
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
For a convenient overview and download list, visit our model page for this model.
biohub/esmc-300m-2024-12
by biohubThis set of model weights was released with the GitHub-compatible esm package format. The models here are kept for backwards compatibility, but we recommend you use the HuggingFace-compatible model weights at biohub/ESMC-6B (or biohub/ESMC-300M / biohub/ESMC-600M) instead.
ScientaLab/eva-rna
by ScientaLabwellsondahostaraguaia/consultas-medica-saude-mulher
by wellsondahostaraguaiaModelo fine-tunado com LoRA (MLX / Apple Silicon) para assistência clínica em saúde da mulher.
Manhph2211/D-BETA
by Manhph2211havocy28/VetBERT
by havocy28This is the pretrained VetBERT model from the github repo: https://github.com/havocy28/VetBERT
Hamdan003/inventmol-r1
by Hamdan003Target-Conditioned Molecular Ideation Model for Drug Discovery Research