work / foundational-theory
Neural network foundational theory
Finding an exact, approximation-free relation for the loss change under an arbitrary weight perturbation, and validating it on pruning and quantization.
The question is whether you can compute exactly how much the loss changes when you touch a weight, with no approximation. Pruning and quantization both come down to “how much worse does the loss get if I delete or round this weight?”, and the formulas used in practice (OBS-style second-order, empirical Fisher) are all approximations.
The goal is not to explain why it behaves a certain way. It is to write a relation that answers how much it changes, and to confirm that relation against direct measurement.
Approach
Write the loss change as a path line integral.
By the fundamental theorem of calculus this is exact regardless of the activation function, the size of the perturbation, or how many weights move at once. The skeleton is the same as Integrated Gradients; the new part is using it to compute the combined effect of touching several weights at once (the pruning interaction), and checking that against measurement.
Where second-order information is needed, the Hessian is never formed. Hv is computed without materializing a matrix, matrix-free, and the linear system is solved with conjugate gradient.
What holds so far
- The line integral matches the exact loss change, to float64 precision, in direct measurement.
- The Hessian split holds at a relative error of 2.4e-16 across the whole network.
- Matrix-free HVP + CG:
Hverror 3.7e-16, CG solution error 9.4e-13, working up to GPT-2-xl without ever forming ad×dmatrix. - The same formula applies to quantization, not just pruning: it holds for an arbitrary weight perturbation, grid rounding included, at a median relative error of 0.2%.
Log
- 2026-08-30 · matrix-free HVP + CG. Reshaping
vinto a(width, n_in)matrix lets you computeHvexactly without ever forming thed×dmatrix (the vec-trick). A dense Hessian for a single GPT-2 small MLP layer (~2.36M params) would be 44.5TB; without the matrix, CG converges in seconds. - 2026-09-08 · the line integral does not force an approximation. An earlier working conclusion, that the relation necessarily requires empirical Fisher and a low-rank truncation, turned out to be wrong. The problem was the algorithm choice (building the full inverse), not the nature of the problem.
- 2026-09-09 · agreement to float64 precision. Validation had been stuck at 0.66% because HuggingFace’s built-in loss upcasts logits to float32 before cross-entropy (deliberate, for mixed-precision stability). Computing the loss in float64 directly brought the relative error to 0.000% across K from 3 to 33, the first time full agreement to float64 precision was pinned down.
- 2026-09-09 · CG scaling. Going from gpt2 to medium, large, and xl (1.558B), the CG iteration count barely grows. It depends on the condition number, not the dimension. Beyond a few billion parameters is untested on this hardware.
Open questions (not conclusions yet)
- The error of the OBS second-order approximation itself. An exact
Hand “a pruning decision made with thatH” are different claims. For a finite perturbation the second-order model is an approximation, and at toy scale one case in twenty once missed by a relative error around 600%. At real GPT-2 scale the outliers seem to disappear, but that is not settled. - The Lottery Ticket / Superposition counter-argument. Results flipped several times between toy scale and real transformers. Right now, at real language-model scale, the combination of a magnitude mask and the original initialization wins by a significant margin.
This page publishes only results that have been reproduced and verified. Anything still in flux stays out of the conclusions.