Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source(1)
Type
887 of 7,078 resources
Showing 701–750
google/hear
by googlemedicalai/ClinicalBERT
by medicalaiThis model card describes the ClinicalBERT model, which was trained on a large multicenter dataset with a large corpus of 1.2B words of diverse diseases we constructed. We then utilized a large-scale corpus of EHRs from over 3 million patient records to fine tune the base language model.
This is the full precision (f16) GGUF version of a model trained for medical chatbot and dental implant assistant tasks. It combines general doctor–patient dialogue understanding with domain-specific Q&A derived from Straumann® dental implant system manuals.
This project fine-tunes the meta-llama/Llama-4-Scout-17B-16E-Instruct model using a medical reasoning dataset (FreedomIntelligence/medical-o1-reasoning-SFT) with 4-bit quantization for memory-efficient training.
FreedomIntelligence/Apollo2-2B
by FreedomIntelligenceCovering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
FreedomIntelligence/Apollo2-9B
by FreedomIntelligenceCovering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
devmanpreet/Medical-GPT2-Classifier
by devmanpreetMedGPT is a GPT-2 model fine-tuned on the pubmed-200k-rct dataset. It classifies individual sentences from biomedical abstracts into one of five standard sections:
prithivMLmods/Food-101-93M
by prithivMLmods!zxdfdsxf.png
Protein solubility is a critical factor in both pharmaceutical research and production processes, as it can significantly impact the quality and function of a protein. This is an example for finetuning ibm/biomed.omics.bl.sm-ted-458m for protein solubility prediction (binary classification) based…
prithivMLmods/Indian-Western-Food-34
by prithivMLmods!fffffff.png
theislab/Nicheformer
by theislabNicheformer is a transformer-based model designed for understanding and predicting cellular niches and their interactions. The model uses masked language modeling to learn representations of cellular contexts and their relationships.
This deep learning model is designed for ECG image classification, fine-tuned using ResNet-50. It can classify ECG images into different categories to assist in heart disease detection.
If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
Mbilal755/Radiology_Bart
by Mbilal755This model summarizes radiology findings into accurate, informative impressions to improve radiologist-clinician communication.
WaltonFuture/Diabetica-7B
by WaltonFutureDiabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management
tahoebio/Tahoe-100M-SCVI-v1
by tahoebioAn SCVI model and minified AnnData of the Tahoe-100M dataset from Vevo Tx.
View label scheme (20 labels for 1 components)
PurvaTijare/PPTStab
by PurvaTijarePPTStab: Prediction and Designing of thermostable proteins with a desired melting temperature
nasa-impact/nasa-ibm-st.38m
by nasa-impactINDUS-Retriever-small (previously nasa-smd-ibm-st.38m) is a Bi-encoder sentence transformer model, that is fine-tuned from distilled version of nasa-smd-ibm-v0.1 encoder model. it is a smaller version of nasa-smd-ibm-st with better performance, using fewer parameters (shown below).
mradermacher/Dans-PersonalityEngine-V1.2.0-24b-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
QizhiPei/biot5-base
by QizhiPei## Example Usage ```python from transformers import AutoTokenizer, T5ForConditionalGeneration
한국어 모델을 이용한 SapBERT(Self-alignment pretraining for BERT)입니다. 한·영 의료 용어 사전인 KOSTOM을 사용해 한국어 용어와 영어 용어를 정렬했습니다. 참고: SapBERT, Original Code
DOEJGI/GenomeOcean-4B
by DOEJGIThis is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).
ai4colonoscopy/ColonGPT
by ai4colonoscopyThe Gradio Web UI allows you to use our examples or upload your images for inference.
StanfordShahLab/llama-base-4096-clmbr
by StanfordShahLabStanfordShahLab/gpt-base-512-clmbr
by StanfordShahLabThis model is a high-performance Named Entity Recognition (NER) model designed specifically for medical text. It identifies entities such as diseases, symptoms, procedures, medications, and healthcare providers with high precision and recall, making it ideal for clinical and healthcare applications.
songlab/gpn-brassicales
by songlab# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.
The Clinical Assertion and Negation Classification BERT is introduced in the paper Assertion Detection in Clinical Notes: Medical Language Models to the Rescue? . The model helps structure information in clinical patient letters by classifying medical conditions mentioned in the letter into…
BiomedCLIP is a biomedical vision-language foundation model that is pretrained on PMC-15M, a dataset of 15 million figure-caption pairs extracted from biomedical research articles in PubMed Central, using contrastive learning.
ArielLubonja/biobert-embeddings
by ArielLubonjaModel from this repo. Model used to be in Dropbox/GDrive, leading to issues with download
FremyCompany/BioLORD-2023
by FremyCompany# FremyCompany/BioLORD-2023 This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
FremyCompany/BioLORD-2023-M
by FremyCompany# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
johahi/borzoi-replicate-0
by johahiUsing llama.cpp release b4404 for quantization.
DiljitSingh14/smol-medical
by DiljitSingh14Henrychur/MMedS-Llama-3-8B
by Henrychur# MMedS-Llama3 💻Github Repo 🖨️arXiv Paper
# PathummaLLM-text-1.0.0-7B: Thai & China & English Large Language Model Instruct PathummaLLM-text-1.0.0-7B is a Thai 🇹🇭 & China 🇨🇳 & English 🇬🇧 large language model with 7 billion parameters, and it is Instruction finetune based on OpenThaiLLM-Prebuilt.
Accurate prediction of drug-target binding affinity is essential in the early stages of drug discovery. This is an example of finetuning ibm/biomed.omics.bl.sm-ted-400 the task. Prediction of binding affinities using pKd, the negative logarithm of the dissociation constant, which reflects the…