Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source(1)
Type(1)
677 of 6,573 resources
Showing 601–650
RaphaelMourad/Mistral-DNA-v1-138M-bacteria
by RaphaelMouradThe Mistral-DNA-v1-138M-bacteria Large Language Model (LLM) is a pretrained generative DNA text model with 17.31M parameters x 8 experts = 138.5M parameters. It is derived from Mistral-7B-v0.1 model, which was simplified for DNA: the number of layers and the hidden size were reduced.
tmberooney/medllama-merged
by tmberooneyModel Card for "medllama" ---------------------------
sagawa/ReactionT5v1-forward
by sagawaThis is a ReactionT5 pre-trained to predict the products of reactions.
minwoosun/uce-100m
by minwoosunUniversal Cell Embeddings (UCE) is a foundation model designed for single-cell RNA sequencing data analysis. UCE generates a universal representation of cells that captures the molecular diversity across different cell types, tissues, and species.
InstaDeepAI/agro-nucleotide-transformer-1b
by InstaDeepAI## Model Overview AgroNT is a DNA language model trained on primarily edible plant genomes. More specifically, AgroNT uses the transformer architecture with self-attention and a masked language modeling objective to leverage highly available genotype data from 48 different plant speices to learn…
mradermacher/SEMIKONG-70B-v2-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
# Medical-Llama3-v2 Fine-Tuned Llama3 for Medical Q&A This repository provides a fine-tuned version of the powerful Llama3 8B model, specifically designed to answer medical questions in an informative way. It leverages the rich knowledge contained in the AI Medical Chatbot dataset…
This is an official model checkpoint for Asclepius-Mistral-7B-v0.3 (arxiv). This model is an enhanced version of Asclepius-7B, by replacing the base model with Mistral-7B-v0.3 and increasing the max sequence length to 8192.
This is an official model checkpoint for Asclepius-Llama3-8B (arxiv). This model is an enhanced version of Asclepius-7B, by replacing the base model with Llama-3 and increasing the max sequence length to 8192.
mradermacher/Medichat-V2-Llama3-8B-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
Henrychur/MMed-Llama-3-8B
by Henrychur# MMedLM 💻Github Repo 🖨️arXiv Paper
nasa-impact/nasa-smd-ibm-st
by nasa-impactThis model is deprecated. please use the updated sentence transformer model here: https://huggingface.co/nasa-impact/nasa-smd-ibm-st-v2. Alternatively, you can also use distilled version of the model here: https://huggingface.co/nasa-impact/nasa-ibm-st.38m
johnsnowlabs/JSL-MedLlama-3-8B-v2.0
by johnsnowlabs# JSL-MedLlama-3-8B-v2.0
Reference: R. Luu and M.J. Buehler, "BioinspiredLLM: Conversational Large Language Model for the Mechanics of Biological and Bio-Inspired Materials," Adv. Science, 2023, DOI: https://doi.org/10.1002/advs.202306724
clinicalnlplab/finetuned-Llama-2-13b-hf-PubmedQA
by clinicalnlplabMedical mT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain
# ChemLLM-7B-Chat-1.5-DPO: LLM for Chemistry and Molecule Science ChemLLM-7B-Chat-1.5-DPO, The First Open-source Large Language Model for Chemistry and Molecule Science, Build based on InternLM-2 with ❤
!image/png
This model is a fine-tuned version of DeBERTa on the PubMED Dataset.
Using llama.cpp release b2440 for quantization.
Goekdeniz-Guelmez/Hyperion-2.0-Mistral-7B-GGUF
by Goekdeniz-Guelmez!image/png
Using llama.cpp commit fa97464 for quantization.
!image/png
# Mr-Grammatology-clinical-problems-Mistral-7B-0.5 !image/png
This is a merge of pre-trained language models created using mergekit.
BioMistral/BioMistral-7B-SLERP
by BioMistralThis is a merge of pre-trained language models created using mergekit.
BioMistral/BioMistral-7B-DARE
by BioMistralThis is a merge of pre-trained language models created using mergekit.
This modelcard aims to be a base template for new models. It has been generated using this raw template.
knowledgator/SMILES2IUPAC-canonical-base
by knowledgatorSMILES2IUPAC-canonical-base was designed to accurately translate SMILES chemical names to IUPAC standards.
songlab/tokenizer-dna-mlm
by songlabrootstrap-org/Alzheimer-Classifier-Demo
by rootstrap-org### Model Description A machine learning model for waste classification
muzammil-eds/tinyllama-2.5T-Clinical-v2
by muzammil-eds# TinyLlama-1.1B
The T5 Large for Medical Text Summarization is a specialized variant of the T5 transformer model, fine-tuned for the task of summarizing medical text. This model is designed to generate concise and coherent summaries of medical documents, research papers, clinical notes, and other…
starmpcc/Asclepius-13B
by starmpccThis is official model checkpoint for Asclepius-13B (arxiv). This model is the first publicly shareable clinical LLM, trained with synthetic data.
TachyHealth/Thealth_Mixtral-8x7B
by TachyHealthepfl-llm/meditron-70b
by epfl-llmDetails coming soon
# Meditron 70B - GGUF - Model creator: EPFL LLM Team - Original model: Meditron 70B
A Vision Transformer (ViT) image classification model. \ Trained on 15M histology patches from PAIP and TCGA. \ Used the MoCo v3 self supervised learning method.
AmelieSchreiber/esm_interact
by AmelieSchreiberThis model was finetuned on concatenated pairs of interacting proteins in much the same way as PepMLM. It is meant to generate interaction partners for proteins using the masked language modeling capabilities of ESM-2. The model is not well tested, so use with caution.
Rostlab/ProstT5
by RostlabProstT5 is a protein language model (pLM) which can translate between protein sequence and structure. !ProstT5 pre-training and inference
## Model Description The "Bird Species Classifier" is a state-of-the-art image classification model designed to identify various bird species from images. It uses the EfficientNet architecture and has been fine-tuned to achieve high accuracy in recognizing a wide range of bird species.
QuiltNet-B-32 is a CLIP ViT-B/32 vision-language foundation model trained on the Quilt-1M dataset curated from representative histopathology videos. It can perform various vision-language processing (VLP) tasks such as cross-modal retrieval, image classification, and visual question answering.
Galahad3x/QAModelForPatho
by Galahad3xQuestion Answering Model for the PathoTHREAT Project
Pre-trained weights and exported models for our spine segmentation project. The source code, designed to reproduce our test results and facilitate training and running inference on your own data, is available on GitHub: https://github.com/MMIV-ML/fastMONAI/tree/master/research
MentaLLaMA-chat-7B is part of the MentaLLaMA project, the first open-source large language model (LLM) series for interpretable mental health analysis with instruction-following capability. This model is finetuned based on the Meta LLaMA2-chat-7B foundation model and the full IMHI instruction…
This model may be overfit to some extent (see below). Try running this notebook on the datasets linked to in the notebook. See if you can figure out why the metrics differ so much on the datasets. Is it due to something like sequence similarity in the train/test split?