Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

881 of 7,064 resources

Showing 701–750

Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.

Idle131 year ago

Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.

Idle651 year ago

MedGPT is a GPT-2 model fine-tuned on the pubmed-200k-rct dataset. It classifies individual sentences from biomedical abstracts into one of five standard sections:

Idle01 year ago

!zxdfdsxf.png

Idle2771 year ago
Python

Protein solubility is a critical factor in both pharmaceutical research and production processes, as it can significantly impact the quality and function of a protein. This is an example for finetuning ibm/biomed.omics.bl.sm-ted-458m for protein solubility prediction (binary classification) based…

Idle871 year ago

!11.png

Idle841 year ago
Python

!fffffff.png

Idle271 year ago
Python

Nicheformer is a transformer-based model designed for understanding and predicting cellular niches and their interactions. The model uses masked language modeling to learn representations of cellular contexts and their relationships.

Idle5851 year ago

This deep learning model is designed for ECG image classification, fine-tuned using ResNet-50. It can classify ECG images into different categories to assist in heart disease detection.

Idle41 year ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Idle2.7K1 year ago
Python

This model summarizes radiology findings into accurate, informative impressions to improve radiologist-clinician communication.

Idle421 year ago
Python

Diabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management

Idle711 year ago
Python

An SCVI model and minified AnnData of the Tahoe-100M dataset from Vevo Tx.

Idle01 year ago

View label scheme (20 labels for 1 components)

Idle391 year ago
Python

PPTStab: Prediction and Designing of thermostable proteins with a desired melting temperature

Idle01 year ago
Python

INDUS-Retriever-small (previously nasa-smd-ibm-st.38m) is a Bi-encoder sentence transformer model, that is fine-tuned from distilled version of nasa-smd-ibm-v0.1 encoder model. it is a smaller version of nasa-smd-ibm-st with better performance, using fewer parameters (shown below).

Idle111 year ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Idle6851 year ago
Python

## Example Usage ```python from transformers import AutoTokenizer, T5ForConditionalGeneration

Idle1911 year ago
Python

한국어 모델을 이용한 SapBERT(Self-alignment pretraining for BERT)입니다. 한·영 의료 용어 사전인 KOSTOM을 사용해 한국어 용어와 영어 용어를 정렬했습니다. 참고: SapBERT, Original Code

Idle191 year ago

This is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).

Idle8581 year ago

The Gradio Web UI allows you to use our examples or upload your images for inference.

Idle11 year ago
Python
Idle71 year ago
Idle601 year ago

This model is a high-performance Named Entity Recognition (NER) model designed specifically for medical text. It identifies entities such as diseases, symptoms, procedures, medications, and healthcare providers with high precision and recall, making it ideal for clinical and healthcare applications.

Idle281 year ago
Python

# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.

Idle5361 year ago
Python

!image/png

Idle6041 year ago
Python

The Clinical Assertion and Negation Classification BERT is introduced in the paper Assertion Detection in Clinical Notes: Medical Language Models to the Rescue? . The model helps structure information in clinical patient letters by classifying medical conditions mentioned in the letter into…

Idle1.5K1 year ago
Python

BiomedCLIP is a biomedical vision-language foundation model that is pretrained on PMC-15M, a dataset of 15 million figure-caption pairs extracted from biomedical research articles in PubMed Central, using contrastive learning.

Idle213.2K1 year ago

Model from this repo. Model used to be in Dropbox/GDrive, leading to issues with download

Idle141 year ago
Python

# FremyCompany/BioLORD-2023 This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.

Idle431.8K1 year ago
Python

# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.

Idle37.1K1 year ago
Python

HuatuoGPT-o1-72B

Idle3621 year ago

HuatuoGPT-o1-7B

Idle6821 year ago
Idle58.3K1 year ago

Using llama.cpp release b4404 for quantization.

Idle8601 year ago
Idle321 year ago

# MMedS-Llama3 💻Github Repo 🖨️arXiv Paper

Idle9481 year ago
Python

HuatuoGPT-o1-8B

Idle1.5K1 year ago

# PathummaLLM-text-1.0.0-7B: Thai & China & English Large Language Model Instruct PathummaLLM-text-1.0.0-7B is a Thai 🇹🇭 & China 🇨🇳 & English 🇬🇧 large language model with 7 billion parameters, and it is Instruction finetune based on OpenThaiLLM-Prebuilt.

Idle9331 year ago

Accurate prediction of drug-target binding affinity is essential in the early stages of drug discovery. This is an example of finetuning ibm/biomed.omics.bl.sm-ted-400 the task. Prediction of binding affinities using pKd, the negative logarithm of the dissociation constant, which reflects the…

Idle9.9K1 year ago

T-cell receptor (TCR) binding to immunogenic peptides (epitopes) presented by major histocompatibility complex (MHC) molecules is a critical mechanism in the adaptive immune system, essential for antigen recognition and triggering immune responses.

Idle411 year ago

Drugs must satisfy stringent criteria for both efficacy and safety. This model predicts the likelihood of FDA approval for small-molecule drugs, represented using SMILES (Simplified Molecular Input Line Entry System) strings.

Idle261 year ago

Drugs must satisfy stringent criteria for both efficacy and safety. This model predicts the likelihood of failure in clinical toxicity trials for small-molecule drugs, represented using SMILES (Simplified Molecular Input Line Entry System) strings.

Idle291 year ago

Drugs targeting the central nervous system must meet stringent criteria for both efficacy and safety, including their ability to penetrate the blood-brain barrier (BBB). This model predicts the likelihood of small-molecule drugs crossing the BBB, a critical factor in CNS drug development.

Idle311 year ago

Accurate prediction of drug-target binding affinity is essential in the early stages of drug discovery. Traditionally, binding affinities are measured through high-throughput screening experiments, which, while accurate, are resource-intensive and limited in their scalability to evaluate large sets…

Idle151 year ago

The ibm/biomed.omics.bl.sm.ma-ted-458m model is a biomedical foundation model trained on over 2 billion biological samples across multiple modalities, including proteins, small molecules, and single-cell gene data. Designed for robust performance, it achieves state-of-the-art results over a variety…

Idle3591 year ago