Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source(1)
Type
477 of 7,078 resources
Showing 51–100
shikunpunk/ask-dao
by shikunpunk> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…
shikunpunk/ask-dao-v0.2-2ep
by shikunpunk知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题
Aurigene-AI/ChemFM-1B
by Aurigene-AISee the upstream model card for full details, training data and citation.
learning-unit/L1-30B-A5B
by learning-unitL1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.
UniParser/MolParser-Mobile
by UniParser💻 Github | 📄 Report | 🚀 Demo
mradermacher/Nidum-Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
NisargRhino/SmileBERTa
by NisargRhinoSmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.
Manhph2211/ECG-Scan
by Manhph2211Taykhoom/ModernGENA-large
by TaykhoomMinimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.
wearewaiv/mascaret
by wearewaivCENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.
JerrryNie/ConceptCLIP
by JerrryNieloveCloud/OmniTCR
by loveCloudOmniTCR is a component-aware autoregressive foundation model for learning relationships among peptide epitopes, major histocompatibility complex (MHC) molecules, T-cell receptor alpha chains (TRA) and T-cell receptor beta chains (TRB).
tzcfly/PertMind
by tzcflyPertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.
NavitraTechnologies01/navikinase-1.0
by NavitraTechnologies01A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.
Eleuthia is a binary protein variant classifier that predicts whether a mutated protein sequence is likely Pathogenic or Benign. It is fine-tuned from ESM-2 and intended for research support in variant prioritization workflows.
A 350M encoder that finds nine types of personally identifiable information across 17 languages and returns exact character spans for review and redaction.
prathmeshadsod/BondShift-Llama-3.3-70B-Instruct
by prathmeshadsodBondShift: Organic Mechanism Reasoning
Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes Indian English, Hindi, Marathi, Kannada and Telugu, including the heavily code-mixed speech typical of real consultations (English drug names and clinical…
Neurazum/VLbai-2.6AD
by NeurazumA clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.
GrimSqueaker/ProtSent-V2-ESMC-300M
by GrimSqueakerContrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
prov-gigapath/prov-gigapath-flash
by prov-gigapathprov-gigapath/prov-gigapath
by prov-gigapathAignostics/RudolfV-2-S
by AignosticsAignostics/RudolfV-2-B
by AignosticsAignostics/RudolfV-2
by AignosticsHuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report 🧬
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…
mradermacher/Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…
!Format !Task !Params !Type !License
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
GrimSqueaker/ProtSent-V2.5-35M
by GrimSqueakerProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.
Meddies/meddies-pii
by MeddiesA multilingual PII extractor for teams that need structured JSON from clinical and administrative text.
GrimSqueaker/ProtSent-V2-150M
by GrimSqueakerContrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
ZeroOneAI/ZEO-Med-2
by ZeroOneAImradermacher/BrainMed-8B-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
duttaprat/DeepVRegulome
by duttaprat464 fine-tuned DNABERT models for regulatory variant effect prediction
PatSnap/Hiro-OCSR
by PatSnapTrinity-Mini-AI-Scientist
zeroentropy/zerank-2-reranker
by zeroentropyIn search engines, rerankers are crucial for improving the accuracy of your retrieval system.
zeroentropy/zerank-1-reranker
by zeroentropyIn search engines, rerankers are crucial for improving the accuracy of your retrieval system.
zeroentropy/zembed-1-embedding
by zeroentropyIn retrieval systems, embedding models determine the quality of your search.
ibm-research/MoLFormer-XL-both-10pct
by ibm-researchMoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.