Biography
Hi there! I am Hanmin Li, a Ph.D. candidate in Computer Science at KAUST in the Center of Excellence for Generative AI, under the supervision of Prof. Peter Richtárik. My research brings together optimization foundations and scalable learning systems, with an emphasis on the efficient training of large language models (LLMs).
My theoretical work spans convex and non-convex optimization, proximal methods, trust-region methods, linear minimization oracles, Frank-Wolfe methods, and matrix-valued optimization. On the systems side, I work on distributed training and inference. I also work on long-horizon tool-using agents, RLHF and RLVR, LLM evaluation, model compression, activation steering, and MoE pruning.
I am currently an Applied Scientist Intern at Microsoft AI, working on long-horizon LLM agents, MoE model pruning, activation steering, and post-training.
Before starting my Ph.D., I earned my master degress in Computer Science also at KAUST, after completing a B.S. in Computer Science and Technology at the School of the Gifted Young in the University of Science and Technology of China (USTC).
Currently, I am working on:
- Optimization methods for scalable learning, including proximal and trust-region methods, local linear minimization oracles, matrix stepsizes, and geometry-aware optimizer design.
- Distributed training and inference for dense and MoE models on large-scale GPU clusters, using data, tensor, pipeline, expert, and context parallelism.
- Long-horizon agent systems, post-training, and model efficiency, including custom evaluation harnesses, LLM-as-a-judge evaluation, model compression, activation steering, and MoE pruning.
For any inquiries, feel free to contact me at hanmin.li@kaust.edu.sa.
Work Experience
- Applied Scientist Intern at Microsoft AI, June 2026 – Present.
Recent News
-
Starting an Applied Scientist Internship at Microsoft AI
— Jun 01, 2026
I started a new position as an Applied Scientist Intern at Microsoft AI, running from June to September 2026.
-
Talk at the ELLIIT Focus Period in Lund
— May 12, 2026
I gave a talk on Stabilizing Proximal Updates: Trust Regions, Linear Descent, and Connections to Modern ML Optimizers at the ELLIIT Focus Period Optimization for Learning in Lund.
-
Attending NeurIPS 2024
— Dec 16, 2024
This year, I will be attending NeurIPS in Vancouver, Canada.
Papers
-
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
, arXiv preprint. • [paper]
-
Broximal Alignment for Global Non-Convex Optimization
, arXiv preprint. • [paper]
-
Stabilized Proximal Point Method via Trust Region Control
, arXiv preprint. • [paper]
-
The Ball-Proximal (=”Broximal”) Point Method: a New Algorithm, Convergence Theory, and Applications
, arXiv preprint. • [paper] • [BibTex]
-
The Power of Extrapolation in Federated Learning
, NeurIPS 2024. • [paper] • [BibTex]
-
On the Convergence of FedProx with Extrapolation and Inexact Prox
, NeurIPS 2024 OPT-ML Workshop Poster. • [paper] • [BibTex]
-
Det-CGD: Compressed Gradient Descent with Matrix Stepsizes for Non-Convex Optimization
, ICLR 2024. • [paper] • [BibTex]
-
Variance reduced distributed non-convex optimization using matrix stepsizes
, NeurIPS 2023 FL@FM Workshop. • [paper] • [BibTex]
-
SD2: spatially resolved transcriptomics deconvolution through integration of dropout and spatial information
, Bioinformatics. • [paper] • [BibTex]
Talks
-
Stabilizing Proximal Updates: Trust Regions, Linear Descent, and Connections to Modern ML Optimizers
May 12, 2026 — ELLIIT Focus Period Optimization for Learning, Lund, Sweden
-
Poster Presentation of On the Convergence of FedProx with Extrapolation and Inexact Prox
Dec 15, 2024 — NeurIPS 2024 OPT-ML Workshop, Vancouver, Canada
-
Poster Presentation of The Power of Extrapolation in Federated Learning
Dec 11, 2024 — NeurIPS 2024, Vancouver, Canada
-
Talk of Det-CGD: Compressed Gradient Descent with Matrix Stepsizes for Non-Convex Optimization
Jun 27, 2024 — EUROPT 2024, Lund, Sweden
-
Poster Presentation of Det-CGD: Compressed Gradient Descent with Matrix Stepsizes for Non-Convex Optimization
May 07, 2024 — ICLR 2024, Vienne, Austria
Reviewer Services
- NeurIPS 24’, 25’
- NeurIPS OPT-ML 24’
- ICLR 25’
- ICML 25’
- JMLR
- IEEE TNNLS
- IEEE TSP
- Optimization Methods and Software.