I am an Assistant Professor at UMass Chan Medical School, where I use machine learning to study regulatory genomics.
pomegranate. Fast, flexible and easy to use probabilistic modelling in Python.
3.5kapricot. apricot implements submodular optimization for the purpose of selecting subsets of massive data sets to train machine learning models quickly. See the documentation page: https://apricot-select.readthedocs.io/en/latest/index.html
534tangermeme. Biological sequence analysis for the modern age.
307yahmm. Yet Another Hidden Markov Model repository.
250avocado. Avocado is a multi-scale deep tensor factorization model that learns a latent representation of the human epigenome and enables imputation of epigenomic experiments that have not yet been performed.
121ledidi. Ledidi turns any machine learning model into a biological sequence editor, allowing you to design sequences with desired properties.
109tfmodisco-lite. A lite implementation of tfmodisco, a motif discovery algorithm for genomics experiments.
92torchegranate. A temporary repository hosting a pomegranate re-write using PyTorch as the backend.
73bpnet-lite. This repository hosts a minimal version of a Python API for BPNet.
60cherimoya. Cherimoya is a lightweight sequence-to-function (S2F) model that is small, fast, and easy to use. Now usable by agents.
39notebooks. Various notebooks and tutorials on subjects of interest.
36dragonnfruit. A method for analyzing scATAC-seq experiments.
33memesuite-lite. A lightweight reimplementation of some of the algorithms in the MEME suite in Python.
33PyPore. Tools used to analyze data from nanopore-based experiments.
31bam2bw. A command-line tool for reading SAM/BAM files and converted them directly to bigwig files.
22rambutan. Prediction of the 3D structure of the genome through statistically significant Hi-C contacts.
21awesome-machine-learning. A curated list of awesome Machine Learning frameworks, libraries and software.
9constraint_graphs. Repository for the data and associated notebook for the manuscript "Finding the optimal Bayesian network given a constraint graph."
6antz. Bioinformatic tools for biologists, made pythonic!
5yabn. Yet Another Bayesian Network
2kiwano. Kiwano implements an approach for prioritizing epigenomic and transcriptomic characterization based on submodular selection and imputed experiments
2deep-learning-notes. Experiments with Deep Learning
2discern. Python
1pystruct. Simple structured learning framework for python
1hmmlearn. Hidden Markov Models in Python, with scikit-learn like API
1lifelines. Survival analysis in Python
1Abada. GUI for analyzing nanopore data, based on the Abada package.
1lychee. An experimental method.
1fertilizer. The Fertile Ground Hypothesis is that genomes are full of "almost-regulatory" regions that do not do anything themselves, but can be minimally edited to achieved subtle and precise activity. This package helps identify the fertile ground that is most useful for your design task.
1