Thanks for response & Invite open discussion
Thanks for your response. Regarding *"consolidate the "Pattern Collapse" phenomenon, experiment finding, evaluation, etc."*, we have already provided a detailed response in the rebuttal stage, including clarification, statistics, visualizations, and quotations. Here we present some key evidence and conclusions again:
---
### 1. "Pattern Collapse" phenomenon (Evidence: Fact Clarification & Statistics)
First, **we have clarified that the language model's modeling of mathematical expressions is based on tokenization**, so "pattern collapse" occurs at the discrete token level, not at the level of mathematical sense. Understanding "pattern collapse" in terms of the real number line is misconceived, and it goes against the way language models are modeled.
Second, **we have given statistics** on seven different types of math tasks with different domains and difficulty levels, and have compared them to two classic text generation tasks, translation and summarization.
We mainly focus on the number of token types and the duplication rate, which can reflect how much the model collapses at the token level when predicting. We again present the results as follows:
|Types of Tasks|token number in dataset |token type number in dataset |Token Duplication Rate|Vocab Coverage|
|---|---|---|---|---|
|***Mathematical Reasoning***|||||
|Arithmetic (primary difficulty)|16136|14|99.9%|0.04%|
|Arithmetic (middle-school difficulty)|5663|16|99.7%|0.05%|
|Algebra|5234|107|98.0%|0.33%|
|Geometry|2615|75|97.1%|0.23%|
|Counting and probability|2524|43|98.3%|0.13%|
|number theory|2395|71|97.1%|0.22%|
|precalculus|3388|84|97.5%|0.26%|
|*Average* | *5422* | *58* | *98.9%*| *0.18%*|
|***Text Generation***||||
|Translation|2500 |1065 | 57.4% | 3.32% |
||5000 | 1832| 63.3% | 5.10% |
||10000 | 2980 | 70.2% | 9.31% |
|| 15000 | 3494 | 76.7% | 10.61% |
|*Average* | *5833* | *1959* | *66.4%*| *6.12%*|
|***Summarization***||||
||2500 |1265 | 49.4% | 4.01% |
||5000 | 1970 | 60.6% | 6.16% |
||10000 | 3192 | 68.0% | 9.98% |
||15000 | 3876 | 74.1% | 12.11% |
|*Average* | *5833* | *2142* | *63.2%* | *6.69%*|
**The average token duplication rate is up to 99% on all math tasks, and even a staggering 99.9% on some simple arithmetic tasks; The token duplication rate exceeded 97% on all seven math tasks of different difficulties and types.
This demonstrates that the "pattern collapse" occurs on generally all types of mathematical reasoning tasks.**
---
### 2. "Early Stabilization" finding (Evidence: Visualization)
In our GENERAL REBUTTAL, **we have given detailed visualizations** of trajectories across ten datasets of different domains and different difficulty levels in mathematical reasoning.
We have also given the average volatility statistics for layers 1-31 (full layers), 20-31, and 26-31 on each dataset corresponding to the visualizations as follows:
|Dataset |1-31 layers |20-31 layers | 26-31 layers|
| --- | --- | --- | --- |
| *ID Dataset* | | | | |
| MultiArith | 6.53 | 14.84 | 10.89 |
| *Near-shift OOD Dataset* | | | | |
| GSM8K | 8.60 | 20.55| 26.43 |
| SVAMP | 8.02 | 18.82 | 24.68 |
| AddSub | 8.72 | 20.54 | 27.24 |
| SingleEq | 8.14 | 19.26 | 22.42 |
| SingleOp | 7.50 | 17.17 | 21.62 |
| *Far-shift OOD Dataset* | | | | |
| MATH-Algebra | 8.83 | 21.35 | 31.50 |
| MATH-Geometry | 10.00 | 25.27 | 34.14 |
| MATH-Count_and_Prob | 10.30 | 25.77 | 33.70 |
| MATH-Number_Theory | 9.40 | 23.10 | 33.86 |
**In all ten datasets, the phenomenon of "early stabilization" is significantly present, and it can be sufficiently demonstrated that "early stabilization" is universal in mathematical reasoning.**
---
### 3. Experiment Justification (Evidence: Quotation)
**We have explained your misunderstanding** about the purpose of our paper by quoting from the original text:
* The two "challenges" are faced by existing embedding-based methods that cannot be applied to mathematical reasoning, not by us. *As stated in lines 38-39, “However, embedding-based methods encounter challenges under mathematical reasoning scenario”.*
* We aim to circumvent the two challenges faced by existing methods, not to address them. We are not improving on existing methods, but rather two different research ideas. *As stated in lines 47-48, “Therefore, we transform our perspective from the static embedding representation to the dynamic embedding shift trajectory in the latent space".*
---
We strongly respect your concerns about the evidence we gave, and an open discussion always helps us to recognize our limited considerations. Unfortunately, we have not received targeted details of your concerns, can you tell us what it is about our evidence that you do not recognize? Your proposed valuable issues will largely help us to improve our paper quality, and we look forward to a more open discussion with you, thanks!