This is your work, valued
Caltech Bioengineering PhD Student
TRILL. Sandbox for Deep-Learning based Computational Protein Design
★ 121protein_gibbs_sampler. Gibbs sampling for generating protein sequences
★ 1MOP-UP. A pipeline that constructs bipartite protein-sharing networks for bacteriophages to analyze host-range, taxonomy and horizontal gene transfer.
★ 1propka. PROPKA predicts the pKa values of ionizable groups in proteins and protein-ligand complexes based in the 3D structure.
★ 368posebusters. Plausibility checks for generated molecule poses.
★ 392PLACER. PLACER is graph neural network for local prediction of protein-ligand conformational ensembles.
★ 260BindFlow. A snakemake-based workflow for FEP and MM(PB/GB)SA calculations with GROMACS
★ 169MAGPIE. MAGPIE
★ 17IPSAE. Scoring function for interprotein interactions in AlphaFold2 and AlphaFold3
★ 254zoplicate. A plugin that does one thing only: Detect and manage duplicate items in Zotero.
★ 939FoldBeast. 3Di substitution models for BEAST 2
★ 5melonnpan. Model-based Genomically Informed High-dimensional Predictor of Microbial Community Metabolic Profiles
★ 44biobakery. bioBakery tools for meta'omic profiling
★ 322metawibele. MetaWIBELE: Workflow to Identify novel Bioactive Elements in microbiome
★ 23fugassem. FUGAsseM: Function predictor of Uncharacterized Gene products by Assessing high-dimensional community data in Microbiomes
★ 21gapseq. Informed prediction and analysis of bacterial metabolic pathways and genome-scale networks
★ 218DRAM. Distilled and Refined Annotation of Metabolism: A tool for the annotation and curation of function for microbial and viral genomes
★ 335humann. HUMAnN is the next generation of HUMAnN 1.0 (HMP Unified Metabolic Analysis Network).
★ 249zDB. zDB: comparative bacterial genomics made easy
★ 41pnabind. A python package and collection of scripts for computing protein surface meshes, chemical, electrostatic, geometric features, and building/training graph neural network models of protein-nucleic acid binding
★ 27boltzgen. BoltzGen: Toward Universal Binder Design
★ 1kCLAPE. contrastive learning and pre-trained encoder for protein-ligand binding sites prediction
★ 37cath-tools-genomescan. CATH: high-throughput protein structure/function annotations
★ 12PLM-interact. PLM-interact: extending protein language models to predict protein-protein interactions.
★ 72folddisco. Fast indexing and search of discontinuous motifs in protein structures
★ 206BattyBirdNET-Pi. A realtime acoustic bat & bird classification system for the Raspberry Pi 4/5 built on BattyBirdNET-Analyzer.
★ 114BattyBirdNET-Analyzer. BattyBirdNET analyzer for scientific audio data processing.
★ 73pi-hole. A black hole for Internet advertisements
★ 60kBirdNET-Pi. A realtime acoustic bird classification system for the Raspberry Pi 5, 4B 3B+ 0W2 and more. Built on the TFLite version of BirdNET.
★ 1.1kAll-Atom-Protein-Sequence-Generation. Python
★ 9PocketVina. GPU-accelerated protein-ligand docking with automated pocket detection, exploring through multi-pocket conditioning. Official Implementation of PocketVina
★ 74plaid. All-atom protein generation using latent diffusion, with compositional function and taxonomic prompts. http://bit.ly/plaid-proteins
★ 127hierarchical_diffusion_LM. Python
★ 10MADSci. Main repository for the Modular Autonomous Discovery for Science (MADSci) Framework
★ 71faker. Faker is a Python package that generates fake data for you.
★ 19kBAT.py. The Binding Affinity Tool (BAT.py) is a fully automated tool for absolute binding free energy (ABFE) and relative binding free energy (RBFE) calculations on protein-ligand systems, compatible with the AMBER and OpenMM simulation packages.
★ 226direct-file. Direct File
★ 4.6kevo2. Genome modeling and design across all domains of life
★ 4kBoltzDesign1. Python
★ 258masif_seed. Masif seed paper repository
★ 172CataPro. A generalized enzyme kinetics parameter prediction model.
★ 96SpeedPPI. Rapid protein-protein interaction network creation from multiple sequence alignments with Deep Learning
★ 90ProLIF. Interaction Fingerprints for protein-ligand complexes and more
★ 526bioemu. Inference code for scalable emulation of protein equilibrium ensembles with generative deep learning
★ 857merizo_search. Fast protein domain structure embedding+search tool
★ 31genomad. geNomad: Identification of mobile genetic elements
★ 326reseek. Protein structure alignment and search algorithm
★ 84BindCraft. User friendly and accurate binder design pipeline
★ 1.2kCatPred. Machine Learning models for in vitro enzyme kinetic parameter prediction
★ 94GraphEC. Python
★ 53ParmEd. Parameter/topology editor and molecular simulator
★ 452Insane. INSert membrANE - A simple, versatile tool for building coarse-grained simulation systems
★ 92GraphSol. Code of our JC paper: "Structure-aware protein solubility prediction from sequence through graph convolutional network and predicted contact map"
★ 77caltechdata_api. Python library for using the CaltechDATA API
★ 12Racer. C
★ 3.8kprevent-nf. PREVENT: PRotein Engineering by Variational frEe eNergy approximaTion
★ 13chai-lab. Chai-1, SOTA model for biomolecular structure prediction
★ 2kPSICHIC. PSICHIC (pronounced Psychic) - PhySIcoCHemICal graph neural network for learning protein-ligand interaction fingerprints from sequence data
★ 141Interformer. Jupyter Notebook
★ 127boltz. Official repository for the Boltz biomolecular interaction models
★ 4.1kColabFold. Making Protein folding accessible to all!
★ 2.9klocalcolabfold. ColabFold on your local PC
★ 871AI2BMD. AI-powered ab initio biomolecular dynamics simulation
★ 577dplm. The Family of Diffusion Protein Language Models (DPLM)
★ 341datamapplot. Creating beautiful plots of data maps
★ 1kAMPLIFY. Python
★ 135FLIP. A collection of tasks to probe the effectiveness of protein sequence representations in modeling aspects of protein design
★ 136tape. Tasks Assessing Protein Embeddings (TAPE), a set of five biologically relevant semi-supervised learning tasks spread across different domains of protein biology.
★ 739BioPhi. BioPhi is an open-source antibody design platform. It features methods for automated antibody humanization (Sapiens), humanness evaluation (OASis) and an interface for computer-assisted antibody sequence design.
★ 255Umol. Protein-ligand structure prediction
★ 241esm-AxP-GDL. Python
★ 20EvolvePro. PLM based active learning model for protein engineering
★ 86CARE. CARE: a Benchmark Suite for the Classification and Retrieval of Enzymes
★ 48AnnotationVocabulary. Jupyter Notebook
★ 7cheap-proteins. Joint embedding of protein sequence and structure with discrete and continuous compressions of protein folding model latent spaces. http://bit.ly/cheap-proteins
★ 155raygun. Template-based protein design with Raygun
★ 95CaLM. Protein language model trained on coding DNA
★ 54RiNALMo. RiboNucleic Acid (RNA) Language Model
★ 169RNA-FM. Nature Methods: RNA foundation model (together with RhoFold)
★ 386PSALM. Protein sequence domain annotation with a language model
★ 33protpardelle. Diffusion-based all-atom protein generative model.
★ 237Metabuli. Metabuli: specific and sensitive metagenomic classification via joint analysis of DNA and amino acid.
★ 180ProGen2-finetuning. Finetuning ProGen2 protein language model for generation of protein sequences from selected protein families.
★ 93esm. Jupyter Notebook
★ 2.9kECPICK. Python
★ 12spades. SPAdes Genome Assembler
★ 955phylophlan. Precise phylogenetic analysis of microbial isolates and genomes from metagenomes
★ 169Jellyfish. A fast multi-threaded k-mer counter
★ 547hostile. Precise host read removal
★ 128openff-forcefields. Force fields produced by the Open Force Field Initiative
★ 188openfe. The Open Free Energy toolkit
★ 312openff-toolkit. The Open Forcefield Toolkit provides implementations of the SMIRNOFF format, parameterization engine, and other tools. Documentation available at http://open-forcefield-toolkit.readthedocs.io
★ 404gLM. Genomic language model predicts protein co-regulation and function
★ 91transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kalphaflow. AlphaFold Meets Flow Matching for Generating Protein Ensembles
★ 535EvoProtGrad. Directed evolution of proteins in sequence space with gradients
★ 96pymbar. Python implementation of the multistate Bennett acceptance ratio (MBAR)
★ 313D-SCRIPT. A structure-aware interpretable deep learning model for sequence-based prediction of protein-protein interactions
★ 118foldseek. Foldseek enables fast and sensitive comparisons of large structure sets.
★ 1.3kprotein_scoring. Generating and scoring novel enzyme sequences with a variety of models and metrics
★ 72evo. Biological foundation modeling from molecular to genome scale
★ 1.5kPLM_SWE. Implementation code for the paper "Aggregating Residue-Level Protein Language Model Embeddings with Optimal Transport"
★ 25e2eatp. A fast and high-accuracy protein-ATP binding residue prediction via protein language model
★ 3m-ionic. Protein-metal ion interaction modeling
★ 6hmmer3di. hmmer3 and easel patched to support the 3di alphabet.
★ 9ECRECer. Dual-core Multi-agent Learning Framework For EC Number Prediction
★ 21genome_portal_api. Python package to access and download data from the ATCC Genome Portal (https://genomes.atcc.org/)
★ 12megaDNA. Jupyter Notebook
★ 76AToM-OpenMM. OpenMM-based framework for absolute and relative binding free energy calculations with the Alchemical Transfer Method
★ 172yank. An open, extensible Python framework for GPU-accelerated alchemical free energy calculations.
★ 203LigandMPNN. Python
★ 609rf_diffusion_all_atom. Public RFDiffusionAA repo
★ 483RoseTTAFold-All-Atom. Python
★ 814bebi103. Utilities for BE/Bi 103
★ 16NFL2K5-Resurrected.
★ 58chiaki-ng. Next-Generation of Chiaki (the open-source remote play client for PlayStation)
★ 2.6kComplete-Fire-Red-Upgrade. A complete upgrade for FireRed, including an upgraded Battle Engine.
★ 783genslm. GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics
★ 142polyply_1.0. Generate input parameters and coordinates for atomistic and coarse-grained simulations of polymers, ssDNA, and carbohydrates
★ 195hyena-dna. Official implementation for HyenaDNA, a long-range genomic foundation model built with Hyena
★ 799ByProt. Python
★ 218GeoDock. Flexible Protein-Protein Docking with a Multi-Track Iterative Transformer.
★ 101ADCP. AutoDock CrankPep for peptide and disordered protein docking
★ 63scikit-optimize. Sequential model-based optimization with a `scipy.optimize` interface
★ 2pLM-BLAST. Detection of remote homology by comparison of protein language model representations
★ 64bopp. Black-box optimization of peptides and proteins
★ 7SaProt. Saprot: Protein Language Model with Structural Alphabet (AA+3Di)
★ 613ColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kopenmmtools. A batteries-included toolkit for the GPU-accelerated OpenMM molecular simulation engine.
★ 335haddock3. Official repo of the modular BioExcel version of HADDOCK
★ 258making-it-rain. Cloud-based molecular simulations for everyone
★ 499simple-simulate-complex. Simple protein-ligand complex simulation with OpenMM
★ 95openmm. OpenMM is a toolkit for molecular simulation using high performance GPU code.
★ 1.9kmasif. MaSIF- Molecular surface interaction fingerprints. Geometric deep learning to decipher patterns in molecular surfaces.
★ 767meld. Modeling with limited data
★ 65PeSTo. Geometric deep learning method to predict protein binding interfaces from a protein structure.
★ 163nucleotide-transformer. Foundation Models for Genomics & Transcriptomics
★ 901genie. De Novo Protein Design by Equivariantly Diffusing Oriented Residue Clouds
★ 189evodiff. Generation of protein sequences and evolutionary alignments via discrete diffusion models
★ 676lightdock-rust. A Rust implementation of the LightDock macromolecular docking software
★ 31lightdock. Protein-protein, protein-peptide and protein-DNA docking framework based on the GSO algorithm
★ 403esmologs. Local homology search powered by ESM-2 language model, foldseek, hhsuite, and hmmer.
★ 18ThermoMPNN. GNN trained to predict changes in thermodynamic stability for protein point mutants
★ 259ProstT5. Bilingual Language Model for Protein Sequence and Structure
★ 318ThemePark. Fun ggplot themes for popular culture
★ 207MMseqs2. MMseqs2: ultra fast and sensitive search and clustering suite
★ 2.1kmurraylab_tools. Functions for building Echo picklists and directions for setting up source plates for the Echo. Mostly intended for setting up TX-TL experiments.
★ 2EpHod. Machine learning prediction of enzyme optimum pH
★ 56engineered-riboregulator-ML. Sequence-to-function deep learning frameworks for engineered riboregulators
★ 14CLEAN. CLEAN: a contrastive learning model for high-quality functional prediction of proteins
★ 323progen. Official release of the ProGen models
★ 705gnina. A deep learning framework for molecular docking
★ 956Ankh. Ankh: Optimized Protein Language Model
★ 249sleepwalk. Exploring dimension-reduced embeddings
★ 111pydock_tutorial. PyDock Tutorial
★ 37papers_for_protein_design_using_DL. List of papers about Proteins Design using Deep Learning
★ 2kprotein_generator. Joint sequence and structure generation with RoseTTAFold sequence space diffusion
★ 332openfold. Trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2
★ 3.4koddt. Open Drug Discovery Toolkit
★ 468excel-html. HTML
★ 1AutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kDiffDock-PP. Implementation of DiffDock-PP: Rigid Protein-Protein Docking with Diffusion Models in PyTorch (ICLR 2023 - MLDD Workshop)
★ 238langchain. The agent engineering platform.
★ 143kRFdiffusion. Code for running RFdiffusion
★ 3kTemStaPro. TemStaPro - a program for protein thermostability prediction using sequence representations from a protein language model.
★ 85DiffDock. Implementation of DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking
★ 1.6kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kLoRA. Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
★ 14kactiveSVC. Jupyter Notebook
★ 6repliclust. Natural Language-Based Synthetic Data Generation for Cluster Analysis.
★ 18diamond. Accelerated BLAST compatible local sequence aligner.
★ 1.3khighdimensional-decision-boundary-plot. Estimating and plotting the decision boundary (decision surface) of machine learning classifiers in higher dimensions (scikit-learn compatible)
★ 235species-aware-DNA-LM. Species-aware DNA language modeling
★ 18iqplot. iqplot: Bokeh plots with one quantitative variable
★ 10gget. 🧬 gget enables efficient querying of genomic reference databases
★ 1.2kMachine-learning-for-proteins. Listing of papers about machine learning for proteins.
★ 1.7kseqkit. A cross-platform and ultrafast toolkit for FASTA/Q file manipulation
★ 1.6kprotein_gibbs_sampler. Gibbs sampling for generating protein sequences
★ 58esm. Evolutionary Scale Modeling (esm): Pretrained language models for proteins
★ 4.2kColabDesign. Making Protein Design accessible to all via Google Colab!
★ 925protein-sequence-models. Python
★ 259bioCRNpyler. A modular compiler for biological chemical reaction networks
★ 54