Learning to Unscramble: Simplifying Symbolic Expressions via Self-Supervised Oracle Trajectories

We present a new self-supervised machine learning approach for symbolic simplification of complex mathematical expressions. Training data is generated by scrambling simple expressions and recording the inverse operations, creating oracle trajectories that provide both goal states and explicit paths to reach them. A permutation-equivariant, transformer-based policy network is then trained on this data step-wise to predict the oracle action given the input expression. We demonstrate this approach on two problems in high-energy physics: dilogarithm reduction and spinor-helicity scattering amplitude simplification. In both cases, our trained policy network achieves near perfect solve rates across a wide range of difficulty levels, substantially outperforming prior approaches based on reinforcement learning and end-to-end regression. When combined with contrastive grouping and beam search, our model achieves a 100\% full simplification rate on a representative selection of 5-point gluon tree-level amplitudes in Yang-Mills theory, including expressions with over 200 initial terms.

Paper

References (9)

02Applies a single model action to the reduced sub-expression, accepts the result if it reduces the term count, and reassembles with the remaining termsgroups of 25. Stacked bars show our results: blue indicates forms solved by the greedy contrastive-grouping phase alone, green indicates forms additionally requiring beam search
0325 50 75 -- 100 -- 125 -- 150 -- 175 -- 200 -- Initial number of terms FIG. 6. Solve rate for 5-point Yang-Mills partial amplitudes
04Reverse and find inverse actions following the procedure of Section II B.
05Reverse to obtain oracle trajectory. Reverse the state sequence to get [ s n scr , . . . , s 1 , s 0 ]—a tra-3
06For each group, factors out common spinor brackets shared by all terms, reducing the sub-expression complexity before model evaluation
07For each reference term, selects the most similar neighbors above a threshold, forming groups of up to 25 terms
08Apply a sequence of n scr random identity transformations to produce a complex expression. At each step, select a random part of the expression and apply a randomly chosen identityRecord the sequence of intermediate states: [ s 0 , s 1 , . .
09“ML Polylogarithms: Code and pretrained models for polylogarithm simplification,”github.com/ aureliendersy/ML_Polylogarithms

Similar papers

© 2026 NYSGPT2525 LLC