Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,068 resources
Showing 1,051–1,100
BioMistral/BioMistral-7B-SLERP
by BioMistralThis is a merge of pre-trained language models created using mergekit.
BioMistral/BioMistral-7B-DARE
by BioMistralThis is a merge of pre-trained language models created using mergekit.
k-mer counting, filtering, and graph traversal.
knowledgator/SMILES2IUPAC-canonical-base
by knowledgatorSMILES2IUPAC-canonical-base was designed to accurately translate SMILES chemical names to IUPAC standards.
A package for benchmarking of models for _de novo_ molecular design.
Molecular descriptor calculator based on [RDKit](http://www.rdkit.org/).
Protein structure prediction from ESM models
muzammil-eds/tinyllama-2.5T-Clinical-v2
by muzammil-eds# TinyLlama-1.1B
The T5 Large for Medical Text Summarization is a specialized variant of the T5 transformer model, fine-tuned for the task of summarizing medical text. This model is designed to generate concise and coherent summaries of medical documents, research papers, clinical notes, and other…
Huawei's 3D high-resolution global weather forecast model at 0.25° resolution, first AI method to comprehensively outperform traditional NWP across all variables and lead times, integrated into ECMWF operational forecasts (Nature 2023)
starmpcc/Asclepius-13B
by starmpccThis is official model checkpoint for Asclepius-13B (arxiv). This model is the first publicly shareable clinical LLM, trained with synthetic data.
3D Equivariant Diffusion for Target-Aware Molecule Generation (ICLR2023)
TachyHealth/Thealth_Mixtral-8x7B
by TachyHealthSingle-cell BERT for gene expression
PINN research collection
Google ViT model is finetuned on lung and colon histopathology image classification dataset. The dataset is available on Kaggle.
epfl-llm/meditron-70b
by epfl-llmDetails coming soon
# Meditron 70B - GGUF - Model creator: EPFL LLM Team - Original model: Meditron 70B
A Vision Transformer (ViT) image classification model. \ Trained on 15M histology patches from PAIP and TCGA. \ Used the MoCo v3 self supervised learning method.
ErnestBeckham/MulticancerViT
by ErnestBeckhamThis is Vision Transformer model trained for cancer classification. To make single model to predict any cancer, I trained this ViT model. following are the cancer types that model can predict: Brain cancer Breast Cancer (histopathology) Lung & Colon Cancer (histopathology) Cervical Caner Kidney…
OpenChem is a deep learning toolkit for Computational Chemistry with PyTorch backend.
AmelieSchreiber/esm_interact
by AmelieSchreiberThis model was finetuned on concatenated pairs of interacting proteins in much the same way as PepMLM. It is meant to generate interaction partners for proteins using the masked language modeling capabilities of ESM-2. The model is not well tested, so use with caution.
Secure text-to-visualization through standardized chart specifications
Rostlab/ProstT5
by RostlabProstT5 is a protein language model (pLM) which can translate between protein sequence and structure. !ProstT5 pre-training and inference
A Vision Transformer (ViT) image classification model. \ Trained by Owkin on 40 million pan-cancer histology tiles from TCGA-COAD.
A Vision Transformer (ViT) image classification model. \ Trained by Owkin on 40M pan-cancer histology tiles from TCGA. \ Fine-tuned on LC25000's lung subset.
## Model Description The "Bird Species Classifier" is a state-of-the-art image classification model designed to identify various bird species from images. It uses the EfficientNet architecture and has been fine-tuned to achieve high accuracy in recognizing a wide range of bird species.
DGL-LifeSci is a [DGL](https://www.dgl.ai/)-based package for various applications in life science with graph neural network.
A Vision Transformer (ViT) image classification model. \ Trained on 2M histology patches from TCGA-BRCA.
Tonic/mistralmed
by TonicThis is a medicine-focussed mistral fine tuned using keivalya/MedQuad-MedicalQnADataset
Galahad3x/QAModelForPatho
by Galahad3xQuestion Answering Model for the PathoTHREAT Project
Write-once-read-many table for large datasets.
First foundation model for weather and climate by Microsoft, Vision Transformer-based architecture trained on heterogeneous datasets (ICML 2023)
MentaLLaMA-chat-7B is part of the MentaLLaMA project, the first open-source large language model (LLM) series for interpretable mental health analysis with instruction-following capability. This model is finetuned based on the Meta LLaMA2-chat-7B foundation model and the full IMHI instruction…
This model may be overfit to some extent (see below). Try running this notebook on the datasets linked to in the notebook. See if you can figure out why the metrics differ so much on the datasets. Is it due to something like sequence similarity in the train/test split?
First vision-and-language foundation model for pathology AI, fine-tuned from CLIP on 249K image-caption pairs, enabling open-ended visual-semantic search and zero-shot diagnosis across histopathology (Pathology Foundation, 376+ stars)
ayoubkirouane/Med_English2Spanish
by ayoubkirouane+ Model Name: Med_English2Spanish + Model Type: Transformer-based Neural Machine Translation (NMT) Model + Task: English to Spanish Medical Translation
This model is a fine-tuned model based on the Llama 2_7b architecture. It has been specifically trained on a dataset comprising USMLE (United States Medical Licensing Examination) questions and answers, as well as conversations between doctors and patients.
Screen a bacterial assembly (contigs/CDS or proteins) for nucleotide or protein sequences. Pipeline that screens for presence of genes of interest (GOI) in bacterial assemblies. Generates multiple CSVs and plots that describe which genes are present and how variable their sequence is. Can use DNA or protein query sequences (GOIs) and DNA contigs/fastas or protein fastas as database (db) to search in.
DeepDILI is a tool for Deep Learning-Powered Drug-Induced Liver Injury Prediction Using Model-Level Representation
An open, extensible Python framework for GPU-accelerated alchemical free energy calculations.
mtag is a Python-based command line tool for jointly analyzing multiple sets of GWAS summary statistics as described by Turley et. al. (2018). It can also be used as a tool to meta-analyze GWAS results.
项目地址:https://github.com/iioSnail/chinesemedicalner
This is a Japanese RoBERTa base model pre-trained on academic articles in medical sciences collected by Japan Science and Technology Agency (JST).
I present a demo showcasing retinal vessel segmentation using the U-Net model, which is a well-known and widely used model in medical image segmentation. The model was trained on the DRIVE dataset, and the training process was conducted on Google Colab.
datasets: - UMLS
In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains.