Yifan Zhang

Expert
@yifanzhang-pro

PhD student at Princeton University, Princeton AI Lab Fellow

deep-delta-learning. Official Project Page for Deep Delta Learning (https://arxiv.org/abs/2601.00417)

356

HLA. Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)

102

AutoMathText. [ACL 2025 Findings] Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts (https://huggingface.co/papers/2402.07625)

92

Matrix-SSL. [ICML 2024] Matrix Information Theory for Self-supervised Learning (https://arxiv.org/abs/2305.17326)

34

FALCON. Official Project Page for FALCON: Fast-Weight Attention for Continual Learning (https://yifanzhang-pro.github.io/FALCON)

26

Kernel-InfoNCE. [ICLR 2024] Contrastive Learning Is Spectral Clustering On Similarity Graph (https://arxiv.org/abs/2303.15103)

21

MEA. Official Project Page for MEA: Matrix Exponential Attention

20

quantum-lattice. Official Project Page for "Exact Coset Sampling for Quantum Lattice Algorithms" (https://arxiv.org/abs/2509.12341)

18

lanser-cli. [Lanser-CLI] Official Implementation of "Reinforcement Learning from Compiler and Language Server Feedback" (https://arxiv.org/abs/2510.22907)

18

M-MAE. [ICML 2024] Matrix Variational Masked Autoencoder (M-MAE) for ICML paper "Information Flow in Self-Supervised Learning" (https://arxiv.org/abs/2309.17281)

15

how-to-train-frontier-models.

12

BlueMO. BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series

8

StackMathQA. StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange

7

Revisiting-SVRPG-LLM-RL. HTML

7

ICS. All 8 Labs of Course: Introduction to Computer Systems with textbook CSAPP

6

residual-stream-duality. Official Project Page for Residual Stream Duality in Modern Transformer Architectures (https://arxiv.org/abs/2603.16039)

6

RelationMatch. RelationMatch: Matching In-batch Relationships for Semi-supervised Learning (https://arxiv.org/abs/2305.10397)

5

fast-weight-attention. Official Project Page for Falcon: Fast Weight Attention for Continual Learning

5

syntax-semantics. TemplateMath: Syntactic Data Generation for Mathematical Problems

4

Matrix-LLM. Official Implementation of ICML 2024 paper 'Matrix Cross-Entropy for Large Language Models' (https://arxiv.org/abs/2305.17326)

4

lagrange-superintelligence. Lagrange Points Are the Next Frontier of Superintelligence

3

Rethinking-SWA. Official Project Page for "Rethinking SWA": Why Short Sliding Window Attention Will Replace ShortConv in Modern Architectures

2

deep-rotary-learning. Deep Rotary Learning

2

matrix-hidden-states. Matrix Hidden States

2

ShortSWA-Ngram-Embedding. Short Sliding Window Attention (ShortSWA) is the next-generation N-gram embedding.

2

yifanzhang-pro. Yifan Zhang

1

TPA. [NeurIPS 2025 Spotlight] TPA: Tensor ProducT ATTenTion Transformer (T6) (https://arxiv.org/abs/2501.06425)

1

efficient-consistent-explanations. [NeurIPS 2023] Trade-off Between Efficiency and Consistency for Removal-based Explanations (https://arxiv.org/abs/2210.17426)

1

RPG. Official implementation of Regularized Policy Gradient (RPG) (https://arxiv.org/abs/2505.17508)

1

general-preference-model. [ICML 2025] Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment (https://arxiv.org/abs/2410.02197)

1

HLA. Higher Attention I: Higher-Order Linear Attention with Parallel Scans

1

quantum-lattice.

1

lm-theory. A Markov Categorical Framework for Language Modeling

1

AutoMathText-2.1.

1
34
Apply