Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

674 of 6,569 resources

Showing 101–150

Active01 month ago
Python
Active7081 month ago
Python
Active431 month ago
Python

Distilling Boltz: Flow Maps for Fast All-Atom Cofolding

Active01 month ago

!Screenshot 2026-07-05 at 2.33.47 AM

Active1.1K1 month ago
Python

### Model Overview TabPFN-3 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/TabPFN. More details can be found in the Model Report.

Active12K1 month ago

This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactive…

Active01 month ago
Python

Fine-tuned gSelformer-MV-Dyn model for predicting dynamic surface tension (DST) curves from SMILES, temperature, and concentration.

Active141 month ago

A native MLX port of OpenMed/privacy-filter-multilingual-v2 for Apple Silicon PII detection and de-identification with OpenMed. This is the unquantized BF16 reference artifact. For the 8-bit sibling, see OpenMed/privacy-filter-multilingual-v2-mlx-8bit.

Active751 month ago

A native MLX port of OpenMed/privacy-filter-multilingual-v2, affine-quantized to 8-bit for faster and smaller Apple Silicon PII detection with OpenMed. For the unquantized BF16 reference, see OpenMed/privacy-filter-multilingual-v2-mlx.

Active361 month ago
Active01 month ago

!ESMC ProtHash Banner

Active461 month ago

!Python

Active01 month ago

πŸ€— Blog | πŸ“„ Paper | πŸ’» Code | 🌐 FineMed | 🩺 DoctoBERT

Active201 month ago

πŸ€— Blog | πŸ“„ Paper | πŸ’» Code | 🌐 FineMed | 🩺 DoctoBERT

Active4601 month ago
Python
Active6221 month ago

Heretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.

Active2351 month ago
Python

HantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.

Active01 month ago
Python

For a convenient overview and download list, visit our model page for this model.

Active4261 month ago
Python

Xinghe1-9B (杏核) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.

Active461 month ago
Python

Pretrained LDARNet (~2M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active01 month ago

Pretrained LDARNet (~110M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active01 month ago

# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language models…

Active841 month ago
Python

This repository contains Chemistry fine-tuned Qwen3-4B checkpoints from the local SciKnowEval-style generalization setup.

Active861 month ago
Python

PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.

Active461 month ago
Python

33 lightweight, standalone models that predict discrete Alzheimer's-disease–related phenotypes (e.g. medication use, APOE genotype, vascular pathology, sex) from SomaScan plasma proteomics. The full list is in phenotypes.tsv.

Active01 month ago

69 lightweight, standalone models that predict continuous Alzheimer's-disease–related phenotypes (cognition, neuropathology burden, motor/functional measures, demographics, a genetic risk score, and longitudinal change) from SomaScan plasma proteomics.

Active01 month ago

This repository contains a drop-in, Hugging Face–compatible checkpoint converted from https://huggingface.co/microsoft/llava-med-v1.5-mistral-7b. You can load it with the exact same code you use for the original modelβ€”no extra conversion steps required.

Active4.3K1 month ago

πŸ‘‹ Join our LiGHT community. πŸ“– Check out the MeditronFO blog and MeditronFO preprint. πŸ”œ If you are a clinician join the MOOVE initiative here.

Active3701 month ago
Python

Sexo-FR is a French-language conversational language model that provides reliable, caring, and evidence-based sexual health information (information en santΓ© sexuelle). It is part of a French public-health initiative whose goal is to make trustworthy sexual-health information more accessible to the…

Active301 month ago
Python

πŸ‘‹ Join our LiGHT community. πŸ“– Check out the MeditronFO blog and MeditronFO preprint. πŸ”œ If you are a clinician join the MOOVE initiative here.

Active6161 month ago
Python

A lightweight plasma-protein aging clock that predicts chronological age from 85 unique inflammatory proteins measured by SomaScan (125 aptamers / SomaScan features). The model is a TabM† student distilled from a TabPFN v2 teacher, so it runs at inference without any TabPFN dependency (small,…

Active02 months ago

A lightweight plasma-protein aging clock that predicts chronological age from 92 Olink Inflammation-panel proteins. The model is a TabM† student distilled from a TabPFN v2 teacher, so it runs at inference without any TabPFN dependency (small, DUA-friendly artifacts).

Active02 months ago
Active1.7K2 months ago
Python

Jolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:

Active4042 months ago
Python

Molexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding and…

Active152 months ago
Python

Molexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,…

Active112 months ago
Python

JEPA-based spatial-transcriptomics foundation model (TERRA). Code & docs: https://github.com/Lotfollahi-lab/terra

Active02 months ago

GRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.

Active192 months ago
Python

Full weight-level fine-tuning of InstaDeepAI/nucleotide-transformer-v2-50m-multi-species for binary DNA sequence classification on two GenomicBenchmarks tasks. All parameters are updated rather than using LoRA or a frozen backbone, with a leakage-free train/validation/test protocol and multi-seed…

Active02 months ago
Python

> Source code, training scripts, and inference utilities for this model: > github.com/NVIDIA-BioNeMo/KERMT > (v2.0 branch / v2.0.0 release tag)

Active02 months ago

QLoRA adapter for Llama-3.1-8B-Instruct, fine-tuned on PubMedQA for yes / no / maybe biomedical question answering (run5).

Active212 months ago
Python

BioMatrix is a multimodal biological foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.

Active1062 months ago
Python
Active02 months ago

A domain-adapted clinical LLM fine-tuned on synthetic Indian medical Q&A records using QLoRA (4-bit quantization) with Unsloth 2x speedup. Built to power the conversational AI layer.

Active852 months ago
Python

This Hugging Face repository stores the official resources for RAMER (reaction-aware multimodal enzyme function representation model).

Active02 months ago

ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein sequence generation.

Active1.6K2 months ago
Python
Active82 months ago
Python

KAU-BioMedLLM is a research prototype for source-grounded biomedical variant interpretation. The current public release contains the LoRA adapter and documentation for a guarded report-generation system built around a curated biomedical evidence panel, citation enforcement, and abstention when…

Active02 months ago
Python