Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type(1)
374 of 6,569 resources
Showing 151–200
Fine-tuned version of google/gemma-4-E4B-it across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.
learning-unit/L1-16B-A3B
by learning-unitL1 (Learning Unit 1) is the first language model from Lunit and Lunit Consortium, purpose-built for the medical domain. Derived from Gravity-16B-A3B-Base, L1 is designed for clinical reasoning and decision support.
## Model Description This is a lightweight, high-performance image classification model built to diagnose histopathological scans of lung and colon tissues. This model was specifically designed for rapid web deployment without sacrificing clinical accuracy.
docling-project/MarkushGrapher-2
by docling-projectMarkushGrapher-2 is an end-to-end multimodal model for recognizing chemical structures from patent document images. It jointly encodes vision, text, and layout information to convert Markush structure images into machine-readable CXSMILES representations.
Verdugie/STEM-Oracle-27B
by Verdugie# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.
docling-project/ChemicalOCR
by docling-projectChemicalOCR is a compact vision-language model fine-tuned specifically for optical character recognition (OCR) in chemical structure images. It extracts text and bounding boxes from molecular drawings, enabling the recognition of atom labels, abbreviations, and descriptive text within chemical…
# GigaHeart ## A Cardiac-specific CT Foundation Model for Heart Transplantation
![Language: Multilingual]()
FreakingPotato/RNAElectra
by FreakingPotatoRNAElectra is a nucleotide-resolution RNA language model trained using an ELECTRA-style objective for efficient and discriminative representation learning. The model produces contextualized embeddings for RNA sequences and is designed for downstream transcriptomic and regulatory modeling tasks.
Matrix-Corp/Vortex-13b-V1
by Matrix-CorpVortex Scientific is a from-scratch AI model family designed for deep scientific reasoning. Built from the ground up with a novel hybrid state-space + attention architecture, optimized for consumer laptop hardware (Apple Silicon MacBooks and Nvidia 4060 laptop GPUs).
zeroentropy/zerank-1-small-reranker
by zeroentropyIn search enginers, rerankers are crucial for improving the accuracy of your retrieval system.
Hengchang-Liu/D3LM-from-nt
by Hengchang-LiuThis repository contains the model presented in D3LM: A Discrete DNA Diffusion Language Model for Bidirectional DNA Understanding and Generation.
mradermacher/Prototype-Virus-1B-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
UmbrellaInc/Prototype-Virus-1B
by UmbrellaInc!image/png
thelamapi/next-ocr
by thelamapi![Language: Multilingual]()
Sahal Shaji Mullappilly\, Mohammed Irfan K\, Omair Mohamed, Mohamed Zidan, Fahad Khan, Salman Khan, Rao Muhammad Anwer, and Hisham Cholakkal
A compact protein language model distilled from ProtGPT2 using complementary-regularizer distillation---a method that combines uncertainty-aware position weighting with calibration-aware label smoothing to achieve 31% better perplexity than standard knowledge distillation at 3.8x compression.
littleworth/protgpt2-distilled-small
by littleworthA compact protein language model distilled from ProtGPT2 using complementary-regularizer distillation---a method that combines uncertainty-aware position weighting with calibration-aware label smoothing to achieve 54% better perplexity than standard knowledge distillation at 9.4x compression.
littleworth/protgpt2-distilled-tiny
by littleworthA compact protein language model distilled from ProtGPT2 using complementary-regularizer distillation---a method that combines uncertainty-aware position weighting with calibration-aware label smoothing to achieve 87% better perplexity than standard knowledge distillation at 20x compression.
InstaDeepAI/NTv3_650M_post
by InstaDeepAIInstaDeepAI/NTv3_650M_post_131kb
by InstaDeepAIInstaDeepAI/NTv3_100M_post
by InstaDeepAIInstaDeepAI/NTv3_650M_pre
by InstaDeepAIInstaDeepAI/NTv3_8M_pre
by InstaDeepAIClinical-Reasoning-Hub/Diagnostic-Medicine-R1
by Clinical-Reasoning-HubGeneral-Medical-AI/UniMedVL
by General-Medical-AI🌟 Github | 📥 Model Download | 📚 Dataset | 📄 Paper Link | 🌐 Project Page
From Inquiry to Decision: Building Trustworthy Medical AI
Raziel1234/OSTLM
by Raziel1234A Neural Machine Translation (NMT) model based on a custom Transformer (Encoder-Decoder) architecture, trained from scratch. This model is designed to translate English sentences into Hebrew using multilingual encoding and specialized layer configurations.
winninghealth/WiNGPT2-Llama-3-8B-Chat
by winninghealthWiNGPT 是一个基于GPT的医疗垂直领域大模型,旨在将专业的医学知识、医疗信息、数据融会贯通,为医疗行业提供智能化的医疗问答、诊断支持和医学知识等信息服务,提高诊疗效率和医疗服务质量。
RationAI/LSP-DETR
by RationAIMatěj Pekár, Vít Musil, Rudolf Nenutil, Petr Holub, Tomáš Brázdil
AVAILABLE NOW THE LATEST ITERATION OF THE ALOE FAMILY! ALOE BETA 8B AND ALOE BETA 70B VERSIONS. These include: Better overall performance More thorough alignment and safety * License compatible with more uses
Aloe: A Family of Fine-tuned Open Healthcare LLMs
Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.
PII Detection Model | 44M Parameters | Open Source
PII Detection Model | 434M Parameters | Open Source
This is a MobileViT (Small) model fine-tuned on the Processed Diabetic Retinopathy dataset.
> [!NOTE] > Inspired by the thought of: what if you could speak to an offline medical assistant that doesn't decline to answer some of your questions?
microsoft/MediPhi-Guidelines
by microsoftThe MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e.
microsoft/MediPhi-Clinical
by microsoftThe MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e.
microsoft/MediPhi-PubMed
by microsoftThe MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e.
microsoft/MediPhi
by microsoftThe MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e.
microsoft/MediPhi-Instruct
by microsoftThe MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e.
## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.
## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology. This model version was continually pretrained on ~14 million cancer transcriptomes…
## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.
## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.
mradermacher/Biomni-R0-32B-Preview-i1-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.