This is your work, valued
Founding engineer working on lightning-fast LLM inference.
vllm-kvcompress. KV cache compression for high-throughput LLM inference
159Syntactically-Constrained-Sampling. LLM sampling method for enforcing syntax adherence in generated output
25tacklebox. Improved handling of PyTorch module hooks
5weighed-levenshtein-substring. Fork of https://github.com/infoscout/weighted-levenshtein
2splatnet. Python
1