Shampoo & Muon: Orthogonalized Stochastic Tensor Optimization Open Access
Becker, Sven (Spring 2026)
Abstract
Adaptive gradient methods have been shown to be some of the most efficient optimizers for stochastic objective functions, and scale to large models very well. Shampoo is a method involving preconditioners formed by second moment estimations, and naive implementations have been shown to perform better than typical adaptive gradient methods in preliminary tests, despite performing many more computations per step. Another optimization method introduced is Muon, involving the use of Newton-Schulz iteration to approximate an orthogonalized gradient. The main computational cost for shampoo is the computation of the root inverse, and the main cost for Muon is the computation of the Newton-Schulz iteration. Variants to both Shampoo and Muon are introduced here with the goal of reducing overall computational cost in performing functions on the matrix of interest. The first variant, for both optimizers, involves randomized projections of the matrix of interest to a smaller subspace to perform the function on this smaller matrix, then projecting the result back to the original size. The second variant, for shampoo only, involves a matrix polynomial interpolation of the root inverse function, utilizing Chebyshev polynomials and a three-term recurrence to efficiently evaluate the function.
Preliminary standardized tests have shown improved training loss for both PyTorch \texttt{resnet} models and an MLP Mixer model compared to Adam for Shampoo, and very similar training loss to Adam with the Muon optimizers. In these tests, the validation accuracy of Shampoo is less than that of Adam, with Muon having very similar validation accuracy to Muon. These experiments could be built on by testing on different types of networks, and the variants can be altered in such ways that other standard optimizers have been altered. Such as introducing a diagonal version for large matrices, decoupled regularization, and grafting methods, but the preliminary tests are provided here.
Table of Contents
Contents
1 Introduction to Stochastic Optimization 1
2 Shampoo 9
3 Analysis 12
4 Shampoo for Tensors 14
5 Implementation 21
6 Muon 24
7 Experimental Results 28
8 Further Directions 38
Bibliography 40
About this Honors Thesis
| School | |
|---|---|
| Department | |
| Degree | |
| Submission | |
| Language |
|
| Research Field | |
| Keyword | |
| Committee Chair / Thesis Advisor | |
| Committee Members |
Primary PDF
| Thumbnail | Title | Date Uploaded | Actions |
|---|---|---|---|
|
|
Shampoo & Muon: Orthogonalized Stochastic Tensor Optimization () | 2026-04-21 15:15:37 -0400 |
|
Supplemental Files
| Thumbnail | Title | Date Uploaded | Actions |
|---|