Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

374 of 6,569 resources

Showing 201250

Support our open-source dataset and model releases!

Idle608 months ago
Python

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Idle12.3K9 months ago
Python

Large Language and Vision Assistant for bioMedicine (i.e., “LLaVA-Med”) is a large language and vision model trained using a curriculum learning method for adapting LLaVA to the biomedical domain. It is an open-source release intended for research use only to facilitate reproducibility of the…

Idle11.4K9 months ago
Python

ChemFIE-BED is a sentence-transformers based on gbyuvd/chemselfies-base-bertmlm fine-tuned on around (for now) 2 million pairs of valid molecules' SELFIES (Krenn et al. 2020) taken from COCONUTDB (Sorokina et al. 2021) and ChemBL34 (Zdrazil et al. 2023).

Idle1069 months ago
Python

GitHub homepage: Cell2Sentence GitHub

Idle1.3K9 months ago
Python
Idle798.7K9 months ago
Python

This model is a fine-tuned version of google/medgemma-4b-it adapted for binary mammogram classification on the OMAMA 256×256 dataset. The dataset consists of ~154k mammogram image slices (.npz) with metadata JSONs providing labels (NonCancer, Cancer).

Idle1510 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Idle46810 months ago
Python

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle14.1K10 months ago
Python

DermLIP is a vision-language model for dermatology, trained on the Derm1M dataset—the largest dermatological image-text corpus to date. This model variant (PanDerm-base-w-PubMed-256) utilizes domain-specific pretraining to deliver superior performance compared to other DermLIP variants..

Idle12910 months ago
Python

MedVAL-4B (medical text validator) is a language model fine-tuned to assess AI-generated medical text outputs at near physician-level reliability.

Idle20310 months ago
Python

This model is a lightweight model pre-trained on SELFIES (Self-Referencing Embedded Strings) representations of molecules. It is trained on 2.7M unique and valid molecules taken from COCONUTDB and ChemBL34, with 7.3M total generated masked examples.

Idle1411 months ago
Python

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

Idle3411 months ago
Python

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

Idle58311 months ago
Python

Website    🤖 7B Model    🤖 32B Model    MedEvalKit    Technical Report    Lingshu MCP

Idle1.5K11 months ago
Python

Website    🤖 7B Model    🤖 32B Model    MedEvalKit    Technical Report    Lingshu MCP

Idle4.1K11 months ago
Python

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle2.9K11 months ago
Python

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle17.9K11 months ago
Python

Palmyra-Med, a powerful LLM designed for healthcare

Idle5311 months ago
Python

中文版说明

Idle14711 months ago
Python

Neeto-1.0-8b is an openly released biomedical large language model (LLM) created by BYOL Academy to assist learners and practitioners with medical exam study, literature understanding, and structured clinical reasoning.

Idle7.7K11 months ago
Python
Idle5111 months ago
Python

This is a ReactionT5 pre-trained to predict the products of reactions. You can use the demo here.

Idle2K1 year ago
Python

This is a ReactionT5 pre-trained to predict the reactants of reactions. You can use the demo here.

Idle1.7K1 year ago
Python

This repos contains the biomedicine MLLM developed from Qwen2.5-VL-3B-Instruct in our paper: On Domain-Adaptive Post-Training for Multimodal Large Language Models. The correspoding training dataset is in biomed-visual-instructions.

Idle1211 year ago
Python

Specialized model for Chemical Entity Recognition - Identifies chemical compounds and substances in biomedical literature

Idle711 year ago
Python

Specialized model for Chemical Entity Recognition - Identifies chemical compounds and substances in biomedical literature

Idle104.1K1 year ago
Python

Specialized model for Chemical Entity Recognition - Chemical entities from the BC5CDR dataset

Idle258K1 year ago
Python

A specialized medical embedding model fine-tuned from Clinical ModernBERT using InfoNCE contrastive learning on PubMed title-abstract pairs.

Idle3.3K1 year ago
Python

Highly focused on medical Training datasets ; + Upgraded inplace

Idle1201 year ago
Python

This model classifies facial skin images into 6 common dermatological conditions using a fine-tuned EfficientNetV2B0 architecture.

Idle811 year ago
Python

darkknight25/deepseek-16b-medical-GPT is a fine-tuned version of deepseek-ai/deepseek-l6b-moe-chat, optimized for medical question answering, reasoning, and clinical summarization using QLoRA and open-access healthcare datasets.

Idle01 year ago
Python

Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.

Idle17.7K1 year ago
Python

This is a merge of pre-trained language models created using mergekit, combining the specialty and general reasoning skills of Esper 3 8b and Shining Valiant 3 8b.

Idle151 year ago
Python

For a convenient overview and download list, visit our model page for this model.

Idle3.6K1 year ago
Python

For a convenient overview and download list, visit our model page for this model.

Idle4651 year ago
Python

For a convenient overview and download list, visit our model page for this model.

Idle4281 year ago
Python

Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.

Idle6K1 year ago
Python
Idle3.7K1 year ago
Python

!image # Compumacy-Experimental_MF ## A Specialized Language Model for Clinical Psychology & Psychiatry

Idle311 year ago
Python

Welcome to IBM's series of large foundation models for sustainable materials. Our models span a variety of representations and modalities, including SMILES, SELFIES, 3D atom positions, 3D density grids, molecular graphs, and other formats.

Idle2121 year ago
Python
Idle1.9K1 year ago
Python
Idle2241 year ago
Python
Idle2341 year ago
Python