This is your work, valued
Designing scalable machine learning algorithms for problems from Computational Genomics.
clustering_on_transcript_compatibility_counts. Clustering cells from single cell RNA seq assays
★ 46combinatorial_MAB. C++
★ 3Normalized_Trend_Filtering. Jupyter Notebook
★ 2ghostty. 👻 Ghostty is a fast, feature-rich, and cross-platform terminal emulator that uses platform-native UI and GPU acceleration.
★ 59kopenworker. Python
★ 11kzsh-completions. Additional completion definitions for Zsh.
★ 7.8kINSID3. [CVPR 2026 Oral] "INSID3: Training-Free In-Context Segmentation with DINOv3"
★ 705ohmyzsh. 🙃 A delightful community-driven (with 2,500+ contributors) framework for managing your zsh configuration. Includes 300+ optional plugins (rails, git, macOS, hub, docker, homebrew, node, php, python, etc), 140+ themes to spice up your morning, and an auto-update tool that makes it easy to keep up with the latest updates from the community.
★ 189kDiffusionBlocks. DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
★ 244cuda-oxide. cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
★ 3kArchon. The first open-source harness builder for AI coding. Make AI coding deterministic and repeatable.
★ 23kcolloquium. A markdown native slides tool for academics building with agents.
★ 235evalstats. Rigorous statistical analysis for LLM evaluations, from model and prompt comparisons to statistical inference resilient to LLM judge bias, including at small sample sizes. All defaults battle-tested in Monte Carlo simulations.
★ 112claude-code-nix. Always up-to-date Nix package for Claude Code with hourly updates from Anthropic's native releases
★ 460chafa. 📺🗿 Terminal graphics for the 21st century.
★ 5.1kzsh-autocomplete. 🤖 Real-time type-ahead completion for Zsh. Asynchronous find-as-you-type autocompletion.
★ 6.7kF3. [SIGMOD 2026] F3: The Open-Source Data File Format for the Future
★ 753Qwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kdeep-learning-with-python-notebooks. Jupyter notebooks for the code samples of the book "Deep Learning with Python"
★ 20kbitchat. bluetooth mesh chat, IRC vibes
★ 34kpmpp. Complete solutions to the Programming Massively Parallel Processors Edition 4
★ 823denoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 11kcontainer. A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for Apple silicon.
★ 48kequinox. Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/
★ 2.9kgonb. GoNB, a Go Notebook Kernel for Jupyter
★ 1kparquet-tools. A utility to deal with Parquet data
★ 205ladybird. Truly independent web browser
★ 65kdeconstructor. A beautiful and interactive web application that deconstructs words into their meaningful parts and explains their etymology.
★ 262dioxus. Fullstack app framework for web, desktop, and mobile.
★ 38kopenslide. C library for reading virtual slide images
★ 509openslide-python. Python bindings to OpenSlide
★ 432devenv. Fast, Declarative, Reproducible, and Composable Developer Environments using Nix
★ 7.2kkitty. If you live in the terminal, kitty is made for you! Cross-platform, fast, feature-rich, GPU based.
★ 34kstarship. ☄🌌️ The minimal, blazing-fast, and infinitely customizable prompt for any shell!
★ 59ksprs. sparse linear algebra library for rust
★ 626rten. ONNX neural network inference engine
★ 329store-interval-tree. A balanced unbounded interval-tree in Rust with associated values in the nodes
★ 13bazel-diff. Performs Bazel Target Diffing between two revisions in Git, allowing for Test Target Selection and Selective Building
★ 514llvm-project. The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.
★ 40kgrok-1. Grok open release
★ 52kcellranger. 10x Genomics Single Cell Analysis
★ 471scan-rs. Single-cell analysis methods in Rust
★ 34piecewise_linear_fit_py. fit piecewise linear data for a specified number of line segments
★ 351Normalized_Trend_Filtering. Jupyter Notebook
★ 2covid19india-cluster. :microscope: Covid19 India Cluster Graph
★ 985spectral_jaccard_similarity. Python
★ 15glmnet_python. Jupyter Notebook
★ 211spectral_jaccard_similarity. Python
★ 1tabula-py. Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame
★ 2.3ktn_test. post-clustering differential expression test
★ 37kgraph. A library for k-nearest neighbor search
★ 387nndescent. Implementation of Efficient K-NN Graph Construction Algorithm NN-Descent in C++
★ 30scikit-learn. scikit-learn: machine learning in Python
★ 67kbanditlib. Multi-armed bandit simulation library
★ 140JSAT. Java Statistical Analysis Tool, a Java library for Machine Learning
★ 794perceptions. Perceptions of Probability and Numbers
★ 867eakmeans. Implementation of fast exact k-means algorithms
★ 45trimed. The trimed algorithm for obtaining the medoid of a set
★ 3deepvariant. DeepVariant is an analysis pipeline that uses a deep neural network to call genetic variants from next-generation DNA sequencing data.
★ 3.8kpyro. Deep universal probabilistic programming with Python and PyTorch
★ 9kclustering_on_transcript_compatibility_counts. Clustering cells from single cell RNA seq assays
★ 46gonum. Gonum is a set of numeric libraries for the Go programming language. It contains libraries for matrices, statistics, optimization, and more
★ 8.4kreflow. A language and runtime for distributed, incremental data processing in the cloud
★ 975Meddit. Jupyter Notebook
★ 8fitting-random-labels. Example code for the paper "Understanding deep learning requires rethinking generalization"
★ 179jupyterlab. JupyterLab computational environment.
★ 15kseaborn. Statistical data visualization in Python
★ 14kdendrosplit. Jupyter Notebook
★ 23lastools. Tools for processing DALIGNER .las files
★ 3minimap2. A versatile pairwise aligner for genomic and spliced nucleotide sequences
★ 2.2kfailures_of_DL. Python
★ 90readme2tex. Renders TeXy Math for Github Readme - No longer needed with official MathTex support on GH
★ 911tufte-latex. A Tufte-inspired LaTeX class for producing handouts, papers, and books
★ 1.9khingeassembler.github.io. HINGE website
★ 3PixelDTGAN. A torch implementation of "Pixel-Level Domain Transfer"
★ 262HINGE-analyses. Analysis accompanying "HINGE: Long-Read Assembly Achieves Optimal Repeat Resolution" http://genome.cshlp.org/content/early/2017/03/20/gr.216465.116.full.pdf
★ 4DAMAPPER. Long read to reference genome mapping tool
★ 13GFA-spec. Graphical Fragment Assembly (GFA) Format Specification
★ 221pacbio-14-nctc-assemblies. 14 NCTC samples for testing PacBio assemblers
★ 4FALCON_unzip. foo
★ 2DAVIEWER. C
★ 6tensorflow. An Open Source Machine Learning Framework for Everyone
★ 197kDEXTRACTOR. Bax File Decoder and Data Compressor
★ 33TensorFlow-Examples. TensorFlow Tutorial and Examples for Beginners (support TF v1 & v2)
★ 44ksuffix-trees. Python implementation of Suffix Trees and Generalized Suffix Trees.
★ 126gephi. Gephi - The Open Graph Viz Platform
★ 6.6kDALIGNER. Find all significant local alignments between reads
★ 140DevNet. The DevNet project on github stores the PacBio DevNet website.
★ 115DAMASKER. Module to determine where repeats are and make soft-masks of said
★ 10vimrc. The ultimate Vim configuration (vimrc)
★ 32kdata-science-sequencing.github.io. Web site for the class EE 372 : Data Science for High-Throughput Sequencing at Stanford
★ 67DASCRUBBER. Alignment-based Scrubbing pipeline
★ 21Bandage. a Bioinformatics Application for Navigating De novo Assembly Graphs Easily
★ 672HINGE. Software accompanying "HINGE: Long-Read Assembly Achieves Optimal Repeat Resolution"
★ 65hacker-scripts. Based on a true story
★ 50k