Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

684 of 6,584 resources

Showing 101150

💻 Github | 📄 Report (Coming soon...) | 🚀 Demo

Active2541 month ago
Python

!image

Active2.5K1 month ago
Python

A 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smooths…

Active1.4K1 month ago
Python

Fine-tuned ESM-2 650M with LoRA for predicting protein subcellular localization (10 classes).

Active181 month ago
Python

PhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iNaturalist records of deciduous woody plants using a two-stage semi-supervised…

Active391 month ago
Python

PhenoVision is a Vision Transformer (ViT-Large) model fine-tuned to detect flowers and fruits in plant photographs. It was trained on 1.5 million human-annotated iNaturalist images and has been used to generate over 30 million new phenology records across 119,000+ plant species, vastly expanding…

Active451 month ago
Python

> NEXUS domain specialist for medical Q&A and clinical reasoning — lightweight & uncensored.

Active16.4K1 month ago

> NEXUS domain specialist for medical Q&A and clinical reasoning — lightweight & uncensored.

Active1.6K1 month ago
Python
Active01 month ago
Python
Active7081 month ago
Python

Distilling Boltz: Flow Maps for Fast All-Atom Cofolding

Active01 month ago

!Screenshot 2026-07-05 at 2.33.47 AM

Active1.1K1 month ago
Python

### Model Overview TabPFN-3 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/TabPFN. More details can be found in the Model Report.

Active12.5K1 month ago

This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactive…

Active01 month ago
Python

Fine-tuned gSelformer-MV-Dyn model for predicting dynamic surface tension (DST) curves from SMILES, temperature, and concentration.

Active151 month ago

A native MLX port of OpenMed/privacy-filter-multilingual-v2 for Apple Silicon PII detection and de-identification with OpenMed. This is the unquantized BF16 reference artifact. For the 8-bit sibling, see OpenMed/privacy-filter-multilingual-v2-mlx-8bit.

Active751 month ago

A native MLX port of OpenMed/privacy-filter-multilingual-v2, affine-quantized to 8-bit for faster and smaller Apple Silicon PII detection with OpenMed. For the unquantized BF16 reference, see OpenMed/privacy-filter-multilingual-v2-mlx.

Active361 month ago
Active01 month ago

!ESMC ProtHash Banner

Active461 month ago

!Python

Active01 month ago

🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT

Active201 month ago

🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT

Active4601 month ago
Python
Active6221 month ago

Heretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.

Active2351 month ago
Python

HantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.

Active01 month ago
Python

For a convenient overview and download list, visit our model page for this model.

Active4261 month ago
Python

Xinghe1-9B (杏核) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.

Active461 month ago
Python

Pretrained LDARNet (~2M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active01 month ago

Pretrained LDARNet (~110M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active01 month ago

# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language models…

Active991 month ago
Python

This repository contains Chemistry fine-tuned Qwen3-4B checkpoints from the local SciKnowEval-style generalization setup.

Active861 month ago
Python

PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.

Active461 month ago
Python

33 lightweight, standalone models that predict discrete Alzheimer's-disease–related phenotypes (e.g. medication use, APOE genotype, vascular pathology, sex) from SomaScan plasma proteomics. The full list is in phenotypes.tsv.

Active02 months ago

69 lightweight, standalone models that predict continuous Alzheimer's-disease–related phenotypes (cognition, neuropathology burden, motor/functional measures, demographics, a genetic risk score, and longitudinal change) from SomaScan plasma proteomics.

Active02 months ago

This repository contains a drop-in, Hugging Face–compatible checkpoint converted from https://huggingface.co/microsoft/llava-med-v1.5-mistral-7b. You can load it with the exact same code you use for the original model—no extra conversion steps required.

Active4.3K2 months ago

👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here.

Active3702 months ago
Python

Sexo-FR is a French-language conversational language model that provides reliable, caring, and evidence-based sexual health information (information en santé sexuelle). It is part of a French public-health initiative whose goal is to make trustworthy sexual-health information more accessible to the…

Active302 months ago
Python

👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here.

Active6162 months ago
Python

A lightweight plasma-protein aging clock that predicts chronological age from 85 unique inflammatory proteins measured by SomaScan (125 aptamers / SomaScan features). The model is a TabM† student distilled from a TabPFN v2 teacher, so it runs at inference without any TabPFN dependency (small,…

Active02 months ago

A lightweight plasma-protein aging clock that predicts chronological age from 92 Olink Inflammation-panel proteins. The model is a TabM† student distilled from a TabPFN v2 teacher, so it runs at inference without any TabPFN dependency (small, DUA-friendly artifacts).

Active02 months ago
Active1.7K2 months ago
Python

Jolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:

Active4042 months ago
Python

Molexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding and…

Active152 months ago
Python

Molexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,…

Active112 months ago
Python

JEPA-based spatial-transcriptomics foundation model (TERRA). Code & docs: https://github.com/Lotfollahi-lab/terra

Active02 months ago

GRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.

Active192 months ago
Python

Full weight-level fine-tuning of InstaDeepAI/nucleotide-transformer-v2-50m-multi-species for binary DNA sequence classification on two GenomicBenchmarks tasks. All parameters are updated rather than using LoRA or a frozen backbone, with a leakage-free train/validation/test protocol and multi-seed…

Active02 months ago
Python

> Source code, training scripts, and inference utilities for this model: > github.com/NVIDIA-BioNeMo/KERMT > (v2.0 branch / v2.0.0 release tag)

Active02 months ago