PhD student at Princeton University, Princeton AI Lab Fellow
deep-delta-learning. Official Project Page for Deep Delta Learning (https://arxiv.org/abs/2601.00417)
356HLA. Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)
102AutoMathText. [ACL 2025 Findings] Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts (https://huggingface.co/papers/2402.07625)
92Matrix-SSL. [ICML 2024] Matrix Information Theory for Self-supervised Learning (https://arxiv.org/abs/2305.17326)
34FALCON. Official Project Page for FALCON: Fast-Weight Attention for Continual Learning (https://yifanzhang-pro.github.io/FALCON)
26Kernel-InfoNCE. [ICLR 2024] Contrastive Learning Is Spectral Clustering On Similarity Graph (https://arxiv.org/abs/2303.15103)
21MEA. Official Project Page for MEA: Matrix Exponential Attention
20quantum-lattice. Official Project Page for "Exact Coset Sampling for Quantum Lattice Algorithms" (https://arxiv.org/abs/2509.12341)
18lanser-cli. [Lanser-CLI] Official Implementation of "Reinforcement Learning from Compiler and Language Server Feedback" (https://arxiv.org/abs/2510.22907)
18M-MAE. [ICML 2024] Matrix Variational Masked Autoencoder (M-MAE) for ICML paper "Information Flow in Self-Supervised Learning" (https://arxiv.org/abs/2309.17281)
15how-to-train-frontier-models.
12BlueMO. BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series
8StackMathQA. StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange
7Revisiting-SVRPG-LLM-RL. HTML
7ICS. All 8 Labs of Course: Introduction to Computer Systems with textbook CSAPP
6residual-stream-duality. Official Project Page for Residual Stream Duality in Modern Transformer Architectures (https://arxiv.org/abs/2603.16039)
6RelationMatch. RelationMatch: Matching In-batch Relationships for Semi-supervised Learning (https://arxiv.org/abs/2305.10397)
5fast-weight-attention. Official Project Page for Falcon: Fast Weight Attention for Continual Learning
5syntax-semantics. TemplateMath: Syntactic Data Generation for Mathematical Problems
4Matrix-LLM. Official Implementation of ICML 2024 paper 'Matrix Cross-Entropy for Large Language Models' (https://arxiv.org/abs/2305.17326)
4lagrange-superintelligence. Lagrange Points Are the Next Frontier of Superintelligence
3Rethinking-SWA. Official Project Page for "Rethinking SWA": Why Short Sliding Window Attention Will Replace ShortConv in Modern Architectures
2deep-rotary-learning. Deep Rotary Learning
2matrix-hidden-states. Matrix Hidden States
2ShortSWA-Ngram-Embedding. Short Sliding Window Attention (ShortSWA) is the next-generation N-gram embedding.
2yifanzhang-pro. Yifan Zhang
1TPA. [NeurIPS 2025 Spotlight] TPA: Tensor ProducT ATTenTion Transformer (T6) (https://arxiv.org/abs/2501.06425)
1efficient-consistent-explanations. [NeurIPS 2023] Trade-off Between Efficiency and Consistency for Removal-based Explanations (https://arxiv.org/abs/2210.17426)
1RPG. Official implementation of Regularized Policy Gradient (RPG) (https://arxiv.org/abs/2505.17508)
1general-preference-model. [ICML 2025] Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment (https://arxiv.org/abs/2410.02197)
1HLA. Higher Attention I: Higher-Order Linear Attention with Parallel Scans
1quantum-lattice.
1lm-theory. A Markov Categorical Framework for Language Modeling
1AutoMathText-2.1.
1