minimap2. A versatile pairwise aligner for genomic and spliced nucleotide sequences
2.2kbwa. Burrow-Wheeler Aligner for short-read alignment (see minimap2 for long-read alignment)
1.8kseqtk. Toolkit for processing sequences in FASTA/Q formats
1.6kbioawk. BWK awk modified for biological data
642minigraph. Sequence-to-graph mapper and graph generator
483miniprot. Align proteins to genomes with splicing and frameshift
410miniasm. Ultrafast de novo assembly for long noisy reads (though having no consensus step)
356wgsim. Reads simulator
286gfatools. Tools for manipulating sequence graphs in the GFA and rGFA formats
252pangene. Constructing a pangenome gene graph
209psmc. Implementation of the Pairwise Sequentially Markovian Coalescent (PSMC) model
193biofast. Benchmarking programming languages/implementations for common tasks in Bioinformatics
184readfq. Fast multi-line FASTA/Q reader in several programming languages
177kmer-cnt. Code examples of fast and simple k-mer counters for tutorial purposes
176yak. Yet another k-mer analyzer
174cgranges. A C/C++ library for fast interval overlap queries (with a "bedtools coverage" example)
172bedtk. A simple toolset for BED files (warning: CLI may change before bedtk becomes stable)
145ksw2. Global alignment and alignment extension
143ropebwt3. BWT construction and search
129hickit. TAD calling, phase imputation, 3D modeling and more for diploid single-cell Hi-C (Dip-C) and general Hi-C
119dipcall. Reference-based variant calling pipeline for a pair of phased haplotype assemblies
115fermikit. De novo assembly based variant calling pipeline for Illumina short reads
110srf. SRF: Satellite Repeat Finder
107minimap. This repo is DEPRECATED. Please use minimap2, the successor of minimap.
106longdust. Identify long STRs, VNTRs, satellite DNA and other low-complexity regions in a genome
100minipileup. Simple pileup-based variant caller
95bgt. Flexible genotype query among 30,000+ samples whole-genome
95unimap. A EXPERIMENTAL fork of minimap2 optimized for assembly-to-reference alignment
87dna-nn. Model and predict short DNA sequence features with neural networks
81bfc. High-performance error correction for Illumina resequencing data
75fermi. A WGS de novo assembler based on the FMD-index for large genomes
74ropebwt2. Incremental construction of FM-index for DNA sequences
72fermi-lite. Standalone C library for assembling Illumina short reads in small regions
72tabtk. Toolkit for processing TAB-delimited format
62ref-gen. Human reference genome analysis sets
62htsbox. My experimental tools on top of htslib. NOT OFFICIAL!!!
60minisplice. Scoring GT/AG sites for improving spliced alignment
58miniwfa. A reimplementation of the WaveFront Alignment algorithm at low memory
51minisv. Lightweight mosaic/somatic SV caller for long reads (WIP)
36pre-pe. Preprocessing paired-end reads produced with experiment-specific protocols
32gffio. C
32fermi2. C
25lianti. Tools to process LIANTI sequence data
23rtgeval. Wrapper for RTG's vcfeval; DEPRECATED!
21sgdp-fermi. FermiKit small variant calls for public SGDP samples
17