Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

884 of 7,068 resources

Showing 351–400

Sparse Autoencoder (SAE) trained on residue-level embeddings from ESM-2 (650M, layer 33) for interpretability research on protein language models.

Active184 months ago

DISCO (DIffusion for Sequence-structure CO-design) is a multimodal generative model that simultaneously co-designs protein sequences and 3D structures, conditioned on and co-folded with arbitrary biomolecules — including small-molecule ligands, DNA, and RNA.

Active04 months ago

A PyTorch MLP that predicts odor descriptors from a molecule's SMILES string using Morgan (ECFP4) fingerprints. Given any molecule, the model outputs a smell profile across 50 odor categories.

Active04 months ago

COMO (Closed-loop Optical Molecule recOgnition) is a deep learning framework for Optical Chemical Structure Recognition (OCSR). It recognizes chemical structure diagrams from images and predicts SMILES strings with atom-level 2D coordinates and bond matrices.

Active04 months ago

Modelo fine-tunado com LoRA (MLX / Apple Silicon) para assistência clínica em saúde da mulher.

Active1254 months ago
Python

vitreg4so150mp14ls_dino-v2-bio is a Bio-DINO image encoder for natural photographs of living organisms. It uses a SoViT-150M/14 Vision Transformer with 4 register tokens and 133.6M backbone parameters, trained with a DINOv2-style self-supervised objective on approximately 31 million curated images…

Active3.9K4 months ago

vitreg1s14lsdino-v2-dist-bio is a compact Bio-DINO image encoder distilled from the larger Bio-DINO SoViT-150M/14 model. It keeps the same natural-photography biodiversity scope as the teacher model, but uses a much smaller ViT-S/14-style student with 21.7M backbone parameters and 384-dimensional…

Active5224 months ago
Active824 months ago
Python

This is the pretrained VetBERT model from the github repo: https://github.com/havocy28/VetBERT

Active944 months ago
Python

This repository contains the full Bio-DINO DINOv2 training weights for a SoViT-150M/14 Vision Transformer trained on natural photographs of living organisms. It is the companion release to the Birder backbone checkpoints at .

Active1324 months ago

To render LaTex in your README, wrap the code in \\( and \\). Example: \\(\frac{1}{2}\\)

Active93.5K4 months ago

This model is a domain-specific large language model fine-tuned from QwQ-32B, specialized in the field of perovskite solar cells, particularly focusing on precursor additives. It is designed to assist researchers, engineers, and material scientists by providing knowledge, insights, and suggestions…

Active34 months ago

Target-Conditioned Molecular Ideation Model for Drug Discovery Research

Active04 months ago
Python

Genos-m is a foundation model for human-associated microbial genomes. It is trained to model microbial DNA sequences at single-nucleotide resolution and supports ultra-long genomic contexts up to one million tokens.

Active314 months ago
Python

Data repository for SynPlanner — an open-source tool for retrosynthetic planning.

Active04 months ago
Active6024 months ago

Qwen3-8B-syco_med-gated-attention-FT is a plug-and-play gated attention weight released for AI safety research.

Active04 months ago
Python

LoRA fine-tune of Qwen3.5-9B on synthetic clinical triage Q&A pairs generated from PubMed Central open-access papers. The model is specialized for emergency-medicine decision-making: triaging patients, applying clinical decision rules, and generating protocol-grounded triage recommendations.

Active04 months ago
Python

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Active1904 months ago

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Active674 months ago

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Active3064 months ago

Gemma 4 E2B fine-tuned on 225K drug–target pairs for novel small-molecule generation.

Active254 months ago
Python

- 2025-05-15: We identified a bug in the Bacformer Large code on HuggingFace which resulted in a significant drop in the quality of the output embeddings. This is now fixed, but if you downloaded or cached the model before this date, re-download and use the latest model revision before running…

Active8K4 months ago
Python

- 2025-05-15: We identified a bug in the Bacformer Large code on HuggingFace which resulted in a significant drop in the quality of the output embeddings. This is now fixed, but if you downloaded or cached the model before this date, re-download and use the latest model revision before running…

Active5474 months ago
Python

This repository contains the model weights for Zatom-1, a general-purpose model that unifies generative modeling and representation learning of 3D molecules and materials. Introduced in Zatom-1: Towards a Multimodal Foundation Model for 3D Molecules and Materials, Zatom-1 is a Transformer trained…

Active04 months ago

This repository hosts JAX exports of MatterSim v1.0.0 for use with kUPS, a JAX-native molecular-simulation toolkit. Each artefact is a self-contained .zip containing the serialized JAX computation graph, the original model parameters, and the minimal metadata needed to run inference.

Active05 months ago
Active85 months ago

Meow-Omni 1 is the world’s first Multimodal Large Language Model (MLLM) specifically engineered for Computational Ethology. It natively co-embeds four distinct modalities—Text, Video, Audio, and Biological Time-Series—to decode the latent intentions of non-verbal species.

Active1195 months ago

For a convenient overview and download list, visit our model page for this model.

Active7035 months ago
Python

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Active1475 months ago
Python

This U‑Net model predicts tilt, z scanner drift, and other large‑scale imaging artifacts present in Atomic Force Microscopy (AFM) height maps. It outputs a background image, the same size and scale as the raw AFM image, which can be subtracted (via the accompanying afMLevel code) to produce a…

Active05 months ago

This U‑Net model masks features in Atomic Force Microscopy (AFM) height maps. It outputs a probability mask image, the same size as the raw AFM image; the accompanying python package, afMLevel code then applies a threshold (typically 0.5) to produce a binary mask.

Active05 months ago

A patient-level disease classification model trained on single-cell RNA-seq data. Given a matrix of gene expression profiles (one row per cell), the model produces a disease-category prediction for the patient.

Active765 months ago
Python
Active05 months ago

MedPsy-4B is a state-of-the-art, text-only medical and healthcare language model purpose-built for edge deployment. Built on top of Qwen3-4B-Thinking-2507 and post-trained with a multi-stage pipeline (supervised fine-tuning + reinforcement learning) on curated medical data, it surpasses models…

Active1K5 months ago
Python

MedPsy-1.7B is a state-of-the-art, text-only medical and healthcare language model purpose-built for edge and smartphone deployment. Built on top of Qwen3-1.7B (operated in thinking mode, i.e. with enable_thinking=True) and post-trained with a multi-stage pipeline (supervised fine-tuning +…

Active6085 months ago
Python

Predict olfactory descriptors from SMILES strings. A GINE-based graph neural network trained on the Leffingwell PMP dataset, reproducing the Principal Odor Map (POM) approach.

Active45 months ago

> [!WARNING] > This is a baseline model trained on publicly available data. While we've done our best to curate the data, the model performance is quite poor. Proceed with caution.

Active245 months ago

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

Active185 months ago

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

Active325 months ago
Active6.2K5 months ago

PickyBinders/tea

by PickyBinders

!Model Architecture

Active5645 months ago

This model is a fine-tuned version of Qwen 3.5 0.8B on a specialized dataset covering biochemistry, peptides, and steroids. It is optimized for providing detailed information on compound mechanisms, dosage (including gender-specific considerations), cycle planning, and physiological effects.

Active6485 months ago

FitCareer_AI Introduction

Active05 months ago

A native MLX port of OpenMed/privacy-filter-nemotron, affine-quantized to 8-bit for fast on-device PII detection on Apple Silicon. For the unquantized BF16 reference, see OpenMed/privacy-filter-nemotron-mlx.

Active2.2K5 months ago

# ACE-V1.1: Brain Tumor Detection !Python!Format > [!CAUTION] > MEDICAL RESEARCH USE ONLY. ACE-V1.1 is NOT a cleared medical device. It must not be used for primary diagnosis or clinical decision-making. All outputs must be verified by a qualified professional.

Active05 months ago

## Introduction (简介) This model is a domain-specific expert fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA (Low-Rank Adaptation). It is specifically designed for Fine-grained Information Extraction (IE) of technical indicator quintuples from highly complex lithium-ion battery patents.

Active105 months ago
Python