Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
881 of 7,064 resources
Showing 701–750
FreedomIntelligence/Apollo2-2B
by FreedomIntelligenceCovering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
FreedomIntelligence/Apollo2-9B
by FreedomIntelligenceCovering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
devmanpreet/Medical-GPT2-Classifier
by devmanpreetMedGPT is a GPT-2 model fine-tuned on the pubmed-200k-rct dataset. It classifies individual sentences from biomedical abstracts into one of five standard sections:
prithivMLmods/Food-101-93M
by prithivMLmods!zxdfdsxf.png
Protein solubility is a critical factor in both pharmaceutical research and production processes, as it can significantly impact the quality and function of a protein. This is an example for finetuning ibm/biomed.omics.bl.sm-ted-458m for protein solubility prediction (binary classification) based…
prithivMLmods/Indian-Western-Food-34
by prithivMLmods!fffffff.png
theislab/Nicheformer
by theislabNicheformer is a transformer-based model designed for understanding and predicting cellular niches and their interactions. The model uses masked language modeling to learn representations of cellular contexts and their relationships.
This deep learning model is designed for ECG image classification, fine-tuned using ResNet-50. It can classify ECG images into different categories to assist in heart disease detection.
If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
Mbilal755/Radiology_Bart
by Mbilal755This model summarizes radiology findings into accurate, informative impressions to improve radiologist-clinician communication.
WaltonFuture/Diabetica-7B
by WaltonFutureDiabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management
tahoebio/Tahoe-100M-SCVI-v1
by tahoebioAn SCVI model and minified AnnData of the Tahoe-100M dataset from Vevo Tx.
View label scheme (20 labels for 1 components)
PurvaTijare/PPTStab
by PurvaTijarePPTStab: Prediction and Designing of thermostable proteins with a desired melting temperature
nasa-impact/nasa-ibm-st.38m
by nasa-impactINDUS-Retriever-small (previously nasa-smd-ibm-st.38m) is a Bi-encoder sentence transformer model, that is fine-tuned from distilled version of nasa-smd-ibm-v0.1 encoder model. it is a smaller version of nasa-smd-ibm-st with better performance, using fewer parameters (shown below).
mradermacher/Dans-PersonalityEngine-V1.2.0-24b-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
QizhiPei/biot5-base
by QizhiPei## Example Usage ```python from transformers import AutoTokenizer, T5ForConditionalGeneration
한국어 모델을 이용한 SapBERT(Self-alignment pretraining for BERT)입니다. 한·영 의료 용어 사전인 KOSTOM을 사용해 한국어 용어와 영어 용어를 정렬했습니다. 참고: SapBERT, Original Code
DOEJGI/GenomeOcean-4B
by DOEJGIThis is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).
ai4colonoscopy/ColonGPT
by ai4colonoscopyThe Gradio Web UI allows you to use our examples or upload your images for inference.
StanfordShahLab/llama-base-4096-clmbr
by StanfordShahLabStanfordShahLab/gpt-base-512-clmbr
by StanfordShahLabThis model is a high-performance Named Entity Recognition (NER) model designed specifically for medical text. It identifies entities such as diseases, symptoms, procedures, medications, and healthcare providers with high precision and recall, making it ideal for clinical and healthcare applications.
songlab/gpn-brassicales
by songlab# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.
The Clinical Assertion and Negation Classification BERT is introduced in the paper Assertion Detection in Clinical Notes: Medical Language Models to the Rescue? . The model helps structure information in clinical patient letters by classifying medical conditions mentioned in the letter into…
BiomedCLIP is a biomedical vision-language foundation model that is pretrained on PMC-15M, a dataset of 15 million figure-caption pairs extracted from biomedical research articles in PubMed Central, using contrastive learning.
ArielLubonja/biobert-embeddings
by ArielLubonjaModel from this repo. Model used to be in Dropbox/GDrive, leading to issues with download
FremyCompany/BioLORD-2023
by FremyCompany# FremyCompany/BioLORD-2023 This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
FremyCompany/BioLORD-2023-M
by FremyCompany# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
johahi/borzoi-replicate-0
by johahiUsing llama.cpp release b4404 for quantization.
DiljitSingh14/smol-medical
by DiljitSingh14Henrychur/MMedS-Llama-3-8B
by Henrychur# MMedS-Llama3 💻Github Repo 🖨️arXiv Paper
# PathummaLLM-text-1.0.0-7B: Thai & China & English Large Language Model Instruct PathummaLLM-text-1.0.0-7B is a Thai 🇹🇭 & China 🇨🇳 & English 🇬🇧 large language model with 7 billion parameters, and it is Instruction finetune based on OpenThaiLLM-Prebuilt.
Accurate prediction of drug-target binding affinity is essential in the early stages of drug discovery. This is an example of finetuning ibm/biomed.omics.bl.sm-ted-400 the task. Prediction of binding affinities using pKd, the negative logarithm of the dissociation constant, which reflects the…
T-cell receptor (TCR) binding to immunogenic peptides (epitopes) presented by major histocompatibility complex (MHC) molecules is a critical mechanism in the adaptive immune system, essential for antigen recognition and triggering immune responses.
Drugs must satisfy stringent criteria for both efficacy and safety. This model predicts the likelihood of FDA approval for small-molecule drugs, represented using SMILES (Simplified Molecular Input Line Entry System) strings.
Drugs must satisfy stringent criteria for both efficacy and safety. This model predicts the likelihood of failure in clinical toxicity trials for small-molecule drugs, represented using SMILES (Simplified Molecular Input Line Entry System) strings.
Drugs targeting the central nervous system must meet stringent criteria for both efficacy and safety, including their ability to penetrate the blood-brain barrier (BBB). This model predicts the likelihood of small-molecule drugs crossing the BBB, a critical factor in CNS drug development.
Accurate prediction of drug-target binding affinity is essential in the early stages of drug discovery. Traditionally, binding affinities are measured through high-throughput screening experiments, which, while accurate, are resource-intensive and limited in their scalability to evaluate large sets…
ibm-research/biomed.omics.bl.sm.ma-ted-458m
by ibm-researchThe ibm/biomed.omics.bl.sm.ma-ted-458m model is a biomedical foundation model trained on over 2 billion biological samples across multiple modalities, including proteins, small molecules, and single-cell gene data. Designed for robust performance, it achieves state-of-the-art results over a variety…