Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

881 of 7,064 resources

Showing 401–450

OGR is a foundational model for AI-driven precision breeding and functional genomics in rice. It is a generative genomic foundation model trained to process DNA sequences up to 1 million base pairs in length, with 1.25B total parameters and a Mixture-of-Experts (MoE) architecture.

Active135 months ago

This repository contains the model used for the paper Bridging Quantum Mechanics to Organic Liquid Properties via a Universal Force Field。

Active05 months ago

A domain-optimized reasoning model built on DeepSeek-R1-Distill-Qwen-32B, refined through a multi-stage pipeline of GPTQ quantization-aware training and QLoRA fine-tuning. Achieves 84% on MedQA — within 4 points of GPT-4o — in a ~20GB package that fits on a single L40/L40s GPU.

Active1135 months ago

# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…

Active4285 months ago

An extended version of SCimilarity, a metric-learning model for single-cell RNA-seq that maps cells to a unified 128-dimensional embedding space. The original model and method are described in:

Active05 months ago

A 50% magnitude-pruned version of facebook/esm2t1235MUR50D optimized for efficient drug discovery inference on Apple Silicon.

Active425 months ago
Python

📝 arXiv • 📖 IEEE TNNLS • 🤗 Model • 🧩 Codes

Active1295 months ago

GeneLinguaLM is a multimodal model that generates natural language descriptions of protein functions from amino acid sequences.

Active05 months ago

Duchifat-2.3-Instruct is a state-of-the-art, instruction-tuned Large Language Model developed by TopAI. As the flagship of the Duchifat series, this model represents a fundamental breakthrough in how Hebrew is processed, reasoned, and generated in the LLM era.

Active1545 months ago
Python

This repository contains the PyTorch model weights for MarS-FM (Markov Space Flow Matching) trained on the MD-CATH dataset. This model was introduced in the ICLR 2026 paper: MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models.

Active05 months ago

!image

Active6015 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active3125 months ago
Python

A generalist foundation model for healthcare capable of handling diverse medical data modalities.

Active715 months ago
Python

A frontier protein-language generative model — because proteins deserve better small talk.

Active165 months ago
Active231.8K5 months ago
Python

DrugCLIP is a dual-encoder multimodal model (SchNet 3D Graph Neural Network + DistilBERT Text Encoder) mapped to a shared 128-dimensional latent space. It is designed to evaluate and retrieve novel 3D molecular structures by aligning them with natural language therapeutic intents and clinical…

Active05 months ago

UBio-MolFM is a foundation model suite for molecular modeling, specifically designed for bio-systems. This model, UBio-MolFM-V1 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer architecture. Refer to the technique report for more details: UBio-MolFM (arXiv:2602.17709).

Active55 months ago

Fine-tuned version of google/gemma-4-E4B-it across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.

Active1K5 months ago
Python

L1 (Learning Unit 1) is the first language model from Lunit and Lunit Consortium, purpose-built for the medical domain. Derived from Gravity-16B-A3B-Base, L1 is designed for clinical reasoning and decision support.

Idle846 months ago
Python

> Note: This checkpoint was donated to Huggingface-science to support open medical AI research

Idle226 months ago

CodonTranslator is a protein-conditioned codon sequence generation model trained on the representative-only data_v3 release.

Idle06 months ago

A frontier protein-language generative model — because proteins deserve better small talk.

Idle102.3K6 months ago

!Banner.

Idle7456 months ago

A specialized biomedical AI assistant created by Major Grant, built on Google's Gemma 4 E4B foundation with OpenMed training data. GGUF format for efficient local inference.

Idle2086 months ago

## Model Description This is a lightweight, high-performance image classification model built to diagnose histopathological scans of lung and colon tissues. This model was specifically designed for rapid web deployment without sacrificing clinical accuracy.

Idle46 months ago
Python
Idle06 months ago

### Model Overview TabPFN-2.6 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.

Idle2.4K6 months ago

This repository contains DNA-sequence modeling resources associated with the basal ganglia (BG) cell atlas package. It serves as a centralized entry point for sequence-based regulatory analyses across multiple companion studies.

Idle06 months ago

🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning

Idle766 months ago

🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning

Idle506 months ago

### Model Overview TabPFN-2.5 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.

Idle19.2K6 months ago

MarkushGrapher-2 is an end-to-end multimodal model for recognizing chemical structures from patent document images. It jointly encodes vision, text, and layout information to convert Markush structure images into machine-readable CXSMILES representations.

Idle1566 months ago
Python

# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.

Idle1426 months ago
Python

EVA is a generative foundation model for universal RNA modeling and design, trained on OpenRNA v1 — a curated atlas of 114 million full-length RNA sequences spanning all domains of life.

Idle06 months ago

ChemicalOCR is a compact vision-language model fine-tuned specifically for optical character recognition (OCR) in chemical structure images. It extracts text and bounding boxes from molecular drawings, enabling the recognition of atom labels, abbreviations, and descriptive text within chemical…

Idle4716 months ago
Python

A biologically & cosmologically inspired causal language model based on the "Cosmology of the Living Cell" (Mother Theory)

Idle86 months ago

Github | Cite

Idle306 months ago

Github | Cite

Idle326 months ago

This is a finetuned EVO2 model for chromosome classification, trained for 20 epochs.

Idle06 months ago

Fine-tuned BGE-M3 on Chinese medical question-answer retrieval using hard negative mining and triple-path InfoNCE loss (dense + sparse + ColBERT).

Idle06 months ago

This repository provides the YOLO26-based version of MolDetv2 model.

Idle06 months ago

A diffusion language model for genome-scale perturbation prediction across diverse cellular contexts.

Idle06 months ago

This repository comprises a collection of TrajCast models, a framework for forecasting molecular dynamics (MD) trajectories using autoregressive equivariant message-passing networks. Provided with a starting configuration comprising information about atom types, atomic positions, and velocities,…

Idle1056 months ago

ClinicDx V1 is a fine-tuned multimodal clinical decision support (CDS) model based on google/medgemma-4b-it. It is trained to generate structured, evidence-grounded clinical assessments from patient presentations, integrating a retrieval-augmented knowledge base (KB) pipeline and an audio input…

Idle456 months ago

A Chemprop v2 multi-component MPNN model that predicts 7 spectroscopic properties of organic chromophores from molecular structure (SMILES) and solvent.

Idle06 months ago
Idle36 months ago

- This repository contains code to utilize the model, and reproduce results of the paper Advancing Codon Language Modeling with Synonymous Codon Constrained Masking. - Unlike other Codon Language Models, SynCodonLM was trained with logit-level control, masking logits for non-synonymous codons.

Idle1017 months ago

A PyTorch port of AlphaGenome, the DNA sequence model from Google DeepMind that predicts hundreds of genomic tracks at single base-pair resolution from sequences up to 1M bp.

Idle817 months ago