Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,027 of 6,573 resources
Showing 801–850
prithivMLmods/Indian-Western-Food-34
by prithivMLmods!fffffff.png
Vision-language pathology foundation model using contrastive learning on histopathology image-text pairs, enabling zero-shot classification, slide-level retrieval, and multimodal reasoning across diverse cancer types (Mahmood Lab, 494+ stars)
This deep learning model is designed for ECG image classification, fine-tuned using ResNet-50. It can classify ECG images into different categories to assist in heart disease detection.
Python wrapper for [bedtools](https://github.com/arq5x/bedtools).
WaltonFuture/Diabetica-7B
by WaltonFutureDiabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management
PurvaTijare/PPTStab
by PurvaTijarePPTStab: Prediction and Designing of thermostable proteins with a desired melting temperature
nasa-impact/nasa-ibm-st.38m
by nasa-impactINDUS-Retriever-small (previously nasa-smd-ibm-st.38m) is a Bi-encoder sentence transformer model, that is fine-tuned from distilled version of nasa-smd-ibm-v0.1 encoder model. it is a smaller version of nasa-smd-ibm-st with better performance, using fewer parameters (shown below).
mradermacher/Dans-PersonalityEngine-V1.2.0-24b-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
Neural optical understanding for academic documents, transforms scientific PDFs to Markdown with mathematical formula support
Generate comprehensive reviews from arXiv papers and convert to blog posts
Microsoft's AI-powered ab initio biomolecular dynamics simulation achieving quantum-mechanical accuracy for proteins with 10,000+ atoms, orders of magnitude faster than DFT using protein fragmentation and ML force fields (Nature 2024)
Equivariant graph attention Transformer (ICLR2023)
Extension of ProteinMPNN for protein sequence design in the context of small-molecule ligands, metal ions, and nucleic acids, enabling binding site engineering and co-factor redesign (Baker Lab)
Physics-AI hybrid modeling for fine-grained weather forecasting (NeurIPS'24)
Geometric deep learning model predicting transcriptional outcomes of novel single- and multi-gene perturbations using gene–gene knowledge graphs, 40% higher precision than prior methods on combinatorial perturbation prediction (Stanford, Nature Biotechnology 2024)
songlab/gpn-brassicales
by songlab# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.
Open-source medical large language model for complex clinical reasoning, extending the o1 long-chain-of-thought paradigm to biomedical question answering and diagnostic inference (FreedomIntelligence, 1.3K+ stars)
The Clinical Assertion and Negation Classification BERT is introduced in the paper Assertion Detection in Clinical Notes: Medical Language Models to the Rescue? . The model helps structure information in clinical patient letters by classifying medical conditions mentioned in the letter into…
ArielLubonja/biobert-embeddings
by ArielLubonjaModel from this repo. Model used to be in Dropbox/GDrive, leading to issues with download
FremyCompany/BioLORD-2023
by FremyCompany# FremyCompany/BioLORD-2023 This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
FremyCompany/BioLORD-2023-M
by FremyCompany# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
A module for solving and visualizing the Schrödinger equation.
Comprehensive toolkit for high-quality PDF content extraction with layout detection, formula recognition, and OCR
Henrychur/MMedS-Llama-3-8B
by Henrychur# MMedS-Llama3 💻Github Repo 🖨️arXiv Paper
### Welcome to Nidum! At Nidum, we believe in pushing the boundaries of innovation by providing advanced and unrestricted AI models for every application. Dive into our world of possibilities and experience the freedom of Nidum-Llama-3.2-3B-Uncensored, tailored to meet diverse needs with…
mradermacher/Bio-Medical-Llama-3.1-8B-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
mradermacher/Breeze-Petro-7B-Instruct-v1-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
Single-cell transformer foundation model pretrained on 104M human transcriptomes via masked gene prediction, enabling transfer learning for cell type classification, gene network analysis, and in silico perturbation with limited labeled data (Nature 2023, V2 2024)
togethercomputer/evo-1-131k-base
by togethercomputerWe identified and fixed an issue related to a wrong permutation of some projections, which affects generation quality. To use the new model revision, please load as follows:
mradermacher/Hyperion-2.0-Mistral-7B-i1-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
peteparker456/medical_diagnosis_llama2
by peteparker456This model aims to be a base template for new models. It has been generated using this raw template.
Large-scale biomolecular instruction dataset for chemistry/biology LLMs (ICLR2024)
Large Language Models for automated open-domain scientific hypotheses discovery (ACL 2024, ICML Best Poster)
Descriptor computation(chemistry) and (optional) storage for machine learning.
nasa-impact/nasa-smd-ibm-v0.1
by nasa-impactIndus (previously known as nasa-smd-ibm-v0.1) is a RoBERTa-based, Encoder-only transformer model, domain-adapted for NASA Science Mission Directorate (SMD) applications. It's fine-tuned on scientific journals and articles relevant to NASA SMD, aiming to enhance natural language technologies like…
nasa-impact/nasa-smd-ibm-st-v2
by nasa-impactIndus-Retriever (nasa-smd-ibm-st-v2) is a Bi-encoder sentence transformer model, that is fine-tuned from nasa-smd-ibm-v0.1 encoder model. it is an updated version of nasa-smd-ibm-st with better performance (shown below). It's trained with 271 million examples along with a domain-specific dataset of…
digitalhealth-healthyliving/MediFlow
by digitalhealth-healthylivingMediFlow se trata de un modelo inicializado con xlnet-large-cased y adaptado con preguntas y especialidades para poder realizar Derivaciones Automatizadas en Servicios Hospitalarios. El dataset se puede encontrar de manera pública y se trata de MedDialog EN.
mradermacher/Palmyra-Med-70B-GGUF
by mradermacherIf you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.
> [!IMPORTANT] > Better using New version of ChemLLM! > AI4Chem/ChemLLM-7B-Chat-1.5-DPO or AI4Chem/ChemLLM-7B-Chat-1.5-SFT
yerevann/chemma-2b
by yerevannChemma-2B is a continually pretrained gemma-2b model for organic molecules. It is pretrained on 40B tokens covering 110M+ molecules from PubChem as well as their chemical properties (molecular weight, synthetic accessibility score, drug-likeness etc.) and similarities (Tanimoto distance between…
qiuhuachuan/PsyChat
by qiuhuachuan## Quick Start ```Python from transformers import AutoTokenizer, AutoModel
Universal chart comprehension and reasoning model
Utility that performs integrated analyses of 'gene' data (a set of genes or other genomic features) with 'peak' data (a set of regions, for example ChIP peaks) to identify the genes nearest to each peak, and vice versa.
Transform arXiv research papers into engaging presentations and YouTube-ready videos
Batteries included genomic analysis pipeline for variant and RNA-Seq analysis, structural variant calling, annotation, and prediction.
Convert PDF files into editable slides with three lines of code
Structure-aware prefix adaptation for integrating LLMs with knowledge graphs (ACM MM 2024)
Powerful and flexible machine learning platform for drug discovery, providing comprehensive tools for molecular property prediction, generative models, knowledge graph reasoning, and reaction prediction with PyTorch backend (1.5K+ stars)