Publications
* denotes equal contribution. See my Google Scholar for citation counts.
-
Gradient Flow Polarizes Softmax Outputs Towards Low-Entropy Solutions
arXiv preprint, 2026 [arXiv] [PDF] [Slides] -
(How) Learning Rates Regulate Catastrophic Overtraining
arXiv preprint, 2026 [arXiv] [PDF] -
Learning In-Context n-grams with Transformers: Sub-n-grams Are Near-Stationary Points
International Conference on Machine Learning (ICML), 2025 [arXiv] [PDF] -
Incremental Learning in Transformers for In-Context Associative Recall
EurIPS 2025 Workshop on Principles of Generative Modeling (PriGM) -
Why Do We Need Weight Decay in Modern Deep Learning?
Advances in Neural Information Processing Systems (NeurIPS), 2024 [arXiv] [PDF] [Code] -
SGD vs GD: Rank Deficiency in Linear Networks
Advances in Neural Information Processing Systems (NeurIPS), 2024 [PDF] [OpenReview] -
SGD with Large Step Sizes Learns Sparse Features
International Conference on Machine Learning (ICML), 2023 [arXiv] [PDF] -
On the Spectral Bias of Two-Layer Linear Networks
Advances in Neural Information Processing Systems (NeurIPS), 2023 [PDF] -
Accelerated SGD for Non-Strongly-Convex Least Squares
Conference on Learning Theory (COLT), 2022 [arXiv] [PDF] -
Last Iterate Convergence of SGD for Least-Squares in the Interpolation Regime
Advances in Neural Information Processing Systems (NeurIPS), 2021 [arXiv] [PDF] -
Variants of Homomorphism Polynomials Complete for Algebraic Complexity Classes
ACM Transactions on Computation Theory (TOCT), 2021 — preliminary version in COCOON 2019