Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

1,191 of 7,068 resources

Showing 951–1,000

## Example Usage ```python from transformers import AutoTokenizer, T5ForConditionalGeneration

Idle1911 year ago
Python

Generate comprehensive reviews from arXiv papers and convert to blog posts

Idle8481 year ago
Python
Apache-2.0

Microsoft's AI-powered ab initio biomolecular dynamics simulation achieving quantum-mechanical accuracy for proteins with 10,000+ atoms, orders of magnitude faster than DFT using protein fragmentation and ML force fields (Nature 2024)

Idle5821 year ago
Python
MIT

Equivariant graph attention Transformer (ICLR2023)

Idle2901 year ago
Python
MIT

Extension of ProteinMPNN for protein sequence design in the context of small-molecule ligands, metal ions, and nucleic acids, enabling binding site engineering and co-factor redesign (Baker Lab)

Idle6351 year ago
Python
MIT

The Gradio Web UI allows you to use our examples or upload your images for inference.

Idle11 year ago
Python

Physics-AI hybrid modeling for fine-grained weather forecasting (NeurIPS'24)

Idle1721 year ago
Python

Geometric deep learning model predicting transcriptional outcomes of novel single- and multi-gene perturbations using gene–gene knowledge graphs, 40% higher precision than prior methods on combinatorial perturbation prediction (Stanford, Nature Biotechnology 2024)

Idle4131 year ago
Python
MIT

This model is a high-performance Named Entity Recognition (NER) model designed specifically for medical text. It identifies entities such as diseases, symptoms, procedures, medications, and healthcare providers with high precision and recall, making it ideal for clinical and healthcare applications.

Idle281 year ago
Python

# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.

Idle5041 year ago
Python

Open-source medical large language model for complex clinical reasoning, extending the o1 long-chain-of-thought paradigm to biomedical question answering and diagnostic inference (FreedomIntelligence, 1.3K+ stars)

Idle1.4K1 year ago
Python

!image/png

Idle6041 year ago
Python

The Clinical Assertion and Negation Classification BERT is introduced in the paper Assertion Detection in Clinical Notes: Medical Language Models to the Rescue? . The model helps structure information in clinical patient letters by classifying medical conditions mentioned in the letter into…

Idle1.5K1 year ago
Python

Model from this repo. Model used to be in Dropbox/GDrive, leading to issues with download

Idle141 year ago
Python

# FremyCompany/BioLORD-2023 This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.

Idle431.8K1 year ago
Python

# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.

Idle37.1K1 year ago
Python

A module for solving and visualizing the Schrödinger equation.

Idle1.2K1 year ago
Python
BSD-3-Clause

Comprehensive toolkit for high-quality PDF content extraction with layout detection, formula recognition, and OCR

Idle10K1 year ago
Python
AGPL-3.0

# MMedS-Llama3 💻Github Repo 🖨️arXiv Paper

Idle9481 year ago
Python

### Welcome to Nidum! At Nidum, we believe in pushing the boundaries of innovation by providing advanced and unrestricted AI models for every application. Dive into our world of possibilities and experience the freedom of Nidum-Llama-3.2-3B-Uncensored, tailored to meet diverse needs with…

Idle4.3K1 year ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Idle1.2K1 year ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Idle1921 year ago
Python

Single-cell transformer foundation model pretrained on 104M human transcriptomes via masked gene prediction, enabling transfer learning for cell type classification, gene network analysis, and in silico perturbation with limited labeled data (Nature 2023, V2 2024)

Idle11 year ago
Python

We identified and fixed an issue related to a wrong permutation of some projections, which affects generation quality. To use the new model revision, please load as follows:

Idle2K1 year ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Idle5331 year ago
Python

This model aims to be a base template for new models. It has been generated using this raw template.

Idle1291 year ago
Python

Large-scale biomolecular instruction dataset for chemistry/biology LLMs (ICLR2024)

Idle2941 year ago
Python
MIT

Large Language Models for automated open-domain scientific hypotheses discovery (ACL 2024, ICML Best Poster)

Idle471 year ago
Python

The plant DNA large language models (LLMs) contain a series of foundation models based on different model architectures, which are pre-trained on various plant reference genomes. All the models have a comparable model size between 90 MB and 150 MB, BPE tokenizer is used for tokenization and 8000…

Idle3171 year ago
Python

Indus (previously known as nasa-smd-ibm-v0.1) is a RoBERTa-based, Encoder-only transformer model, domain-adapted for NASA Science Mission Directorate (SMD) applications. It's fine-tuned on scientific journals and articles relevant to NASA SMD, aiming to enhance natural language technologies like…

Idle552 years ago
Python

Indus-Retriever (nasa-smd-ibm-st-v2) is a Bi-encoder sentence transformer model, that is fine-tuned from nasa-smd-ibm-v0.1 encoder model. it is an updated version of nasa-smd-ibm-st with better performance (shown below). It's trained with 271 million examples along with a domain-specific dataset of…

Idle1.6K2 years ago
Python

MediFlow se trata de un modelo inicializado con xlnet-large-cased y adaptado con preguntas y especialidades para poder realizar Derivaciones Automatizadas en Servicios Hospitalarios. El dataset se puede encontrar de manera pública y se trata de MedDialog EN.

Stale62 years ago
Python

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

Stale3812 years ago
Python

> [!IMPORTANT] > Better using New version of ChemLLM! > AI4Chem/ChemLLM-7B-Chat-1.5-DPO or AI4Chem/ChemLLM-7B-Chat-1.5-SFT

Stale1.1K2 years ago
Python

Chemma-2B is a continually pretrained gemma-2b model for organic molecules. It is pretrained on 40B tokens covering 110M+ molecules from PubChem as well as their chemical properties (molecular weight, synthetic accessibility score, drug-likeness etc.) and similarities (Tanimoto distance between…

Stale472 years ago
Python

## Quick Start ```Python from transformers import AutoTokenizer, AutoModel

Stale382 years ago
Python

Universal chart comprehension and reasoning model

Stale1362 years ago
Python
NOASSERTION

Utility that performs integrated analyses of 'gene' data (a set of genes or other genomic features) with 'peak' data (a set of regions, for example ChIP peaks) to identify the genes nearest to each peak, and vice versa.

Stale52 years ago
Python
Artistic-2.0

Transform arXiv research papers into engaging presentations and YouTube-ready videos

Stale142 years ago
Python

Batteries included genomic analysis pipeline for variant and RNA-Seq analysis, structural variant calling, annotation, and prediction.

Stale1K2 years ago
Python
MIT

Materials informatics benchmark

Stale2182 years ago
Python
MIT

Convert PDF files into editable slides with three lines of code

Stale172 years ago
Python
GPL-3.0

Structure-aware prefix adaptation for integrating LLMs with knowledge graphs (ACM MM 2024)

Stale2132 years ago
Python
MIT

Powerful and flexible machine learning platform for drug discovery, providing comprehensive tools for molecular property prediction, generative models, knowledge graph reasoning, and reaction prediction with PyTorch backend (1.5K+ stars)

Stale1.6K2 years ago
Python
Apache-2.0

ChemFIE-SA is a BERT-like sequence classifier for predicting synthesis accessibility given a SELFIES string of a compound, fine-tuned from gbyuvd/chemselfies-base-bertmlm on DeepSA's expanded dataset from Wang et al. 2023.

Stale82 years ago
Python

This model is a BERT-like sequence classifier for 221 human protein drug targets, fine-tuned from gbyuvd/chemselfies-base-bertmlm on a dataset derived ChemBL34 (Zdrazil et al. 2023). It predicts potential drug targets using chemical structures represented as SELFIES (Self-Referencing Embedded…

Stale402 years ago
Python

Resources on ChIP-seq data which include papers, methods, links to software, and analysis.

Stale8542 years ago
Python
MIT

The Mistral-DNA-v1-138M-bacteria Large Language Model (LLM) is a pretrained generative DNA text model with 17.31M parameters x 8 experts = 138.5M parameters. It is derived from Mistral-7B-v0.1 model, which was simplified for DNA: the number of layers and the hidden size were reduced.

Stale162 years ago
Python

Model Card for "medllama" ---------------------------

Stale152 years ago
Python

UNIX-style FASTA manipulation tools.

Stale192 years ago
Python
MIT