Alignment at Pre-training! Towards Native Alignment for Arabic LLMs

The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `post alignment'. We argue that alignment during the pre-training phase, which we term `native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community.

Paper

Similar papers

Peer review

Reviewer c7RZ6/10 · confidence 4/52024-07-04

Summary

This paper proposed a new method for LLM alignment during pre-training. The proposed method is call "native alignment". This method include three steps: pretrain date duplication, alignment rewriting, and model training. They trained small size alignment expert model for alignment rewriting and use the model to rewrite large-scale pre-training data. The rewriting process suppose to solve format issue, value/fairness issue, unsafe content in pre-training data. They experimented with Arabic data and LLMs. Their experiments shows that the proposed method can help LLMs be more safe and helpful.

Strengths

1. The paper proposed a new idea to align LLMs during pre-training. It seems an interesting topic. 2. The paper writing is clear and well-organized.

Weaknesses

1. Lack of comparison to existing post-alignment methods. The proposed method is a "native alignment" during pre-training. I wonder if this method can outperform the post-alignment methods. While the author acknowledged this limitation, I still feel it is important for strengthening their claim. 2. Need more analyses to better understand their method's potential trade-off. For example, I wonder if rewriting pre-training data undermines the LLM's capacity to understand and learn Arabic dialects. The rewriting process may convert Arabic dialects into MSA. I also wonder if the rewriting data inherited hallucinations from LLM and deteriorated the trained model. 3. The paper needs more clarification on experiment details. For example, they exploit Arabic data to investigate their method, however, the evaluation dataset, BeaverTails dataset, is an English dataset. I wonder how they evaluate and if they translate the samples.

Questions

Please see the questions in weakness. * I wonder whether you continued training the LLaMA 3 model or newly initialized LLaMA-like model and trained it from scratch.

Rating

6

Confidence

4

Soundness

2

Presentation

3

Contribution

2

Limitations

I think that the paper needs more diverse analyses to understand the potential trade-off of their method.

Reviewer 5nE63/10 · confidence 5/52024-07-10

Summary

The paper introduce a method called "native alignment", which is a set of procedures to create data and train an LLM to rewrite raw text into "useful" texts for pretraining. They apply this technique specifically for Arabic LLMs and conduct experiments to show that this pre-processing of pre-training data helps produce better Arabic LLM down the line. As bonus, they release open-source Arabic LLMs for the communities

Strengths

* The paper ideas are presented clearly and easy-to-understand

Weaknesses

* As a proclaimed novelty, the paper draws itself between pre-alignment and post-alignment, indicating that previous work only focus on post-alignment but not pre-alignment. However, I afraid the paper misunderstands the concept of post-alignment (RLHF) and fails make an accurate comparison. Post alignment (RLHF) is finetuning technique to train the models to reward good-vs-bad response according to human values, and train the policy models to lean on the good behavior and stay-away from the bad behaviors gradually, often with the existence of a reference model (DPO and RLHF). Meanwhile, the "native alignment" presented in the paper is a data-cleaning procedure, and it does not having any resemblance or contrast with "post-alignment". Furthermore, using LLMs or training LLMs to rewrite raw text to produce cleaner data is not new or novel, there are many techniques out there that do so, and there are abundant open-source data on huggingface which were produced in similar ways. This confusion between data cleaning and alignment makes the paper less credible and the lack of novelty it the methodology itself, as a data cleaning method, is also troublesome. Obviously as a result, the paper did not provide any necessary and required experimental comparisons with other data cleaning methods. * Though I do appreciate the paper's effort for Arabic community, the scope of only Arabic LLM is small and generally inconclusive, that such method is not shown to generalize to other languages, domains. Perhaps, thus, the work is really not suitable for NeurIPS but more suitable for CL-type venues * It is unclear from the writing whether the authors pretrained Llama-3 with Arabic from scratch or further finetune from Llama-3 checkpoint. In either case, there should be explanation and further ablation studies.

Questions

Did the authors pretrain from scratch (with Llama-3 architecture) or from Llama-3 checkpoint

Rating

3

Confidence

5

Soundness

2

Presentation

3

Contribution

2

Limitations

The authors discussed limitations

Authorsrebuttal2024-08-12

Enhancing Clarity in Native Alignment Approach

Thank you for your valuable review and for pointing out the weaknesses in our paper. We recognize the importance of clearly emphasizing the novelty of native alignment. In the final version of the paper, we will ensure that the distinctions and connections between our proposed native alignment approach and traditional data cleaning methods, as well as post-alignment techniques like RLHF, are clearly highlighted. The additional experiments conducted during the rebuttal period will also be incorporated into the final version. Furthermore, we will **expand the related work section** to provide a more comprehensive context for our contributions. We have identified the following research to include and would greatly appreciate any additional references you might suggest: 1. **Data Cleaning Methods:** - Hegazi, M.O., Al-Dossari, Y., Al-Yahy, A., Al-Sumari, A. and Hilal, A., 2021. Preprocessing Arabic text on social media. Heliyon, 7(2). - Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N. and Presser, S., 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027. - Wenzek, G., Lachaux, M.A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A. and Grave, E., 2019. CCNet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359. - Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E. and Launay, J., 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116. - Fan, R.Z., Li, X., Zou, H., Li, J., He, S., Chern, E., Hu, J. and Liu, P., 2024. Reformatted alignment. arXiv preprint arXiv:2402.12219. - Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L. and Zhang, S., 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36. 2. **Post-Alignment:** - Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y. and Gan, C., 2024. Principle-driven self-alignment of language models from scratch with minimal human supervision. Advances in Neural Information Processing Systems, 36. - Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A. and Schulman, J., 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, pp.27730-27744. - Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V. and Rastogi, A., 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267. - Zhu, B., Jordan, M. and Jiao, J., 2023, July. Principled reinforcement learning with human feedback from pairwise or k-wise comparisons. In International Conference on Machine Learning (pp. 43037-43067). PMLR. - Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y. and Wang, H., 2024, March. Preference ranking optimization for human alignment. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 17, pp. 18990-18998). Again, thank you for your feedback. **We would be glad to engage in further discussion if you have any remaining concerns.**

Reviewer 5nE62024-08-13

Thanks for the response

Thank you for the response. I decide to keep rating unchanged. * This is data cleaning regardless how authors disagree with this. Deception of concepts to create a sense of novelty should be discouraged. * All previous data cleaning processes all aim to be aligned with human values, such as removing NSFW words and toxic content. Yet no one even call them "alignment".

Authorsrebuttal2024-08-13

"Native Alignment is a special case of Data Cleaning"

Thank you for clarifying your concerns. We indeed agreed that native alignment is a special case of data cleaning, as previously stated in the rebuttal. We do not intend to create a sense of novelty through conceptual deception: the significant difference between native alignment and traditional data cleaning is that native alignment generalizes the latter to a broader scope, specifically focusing on value alignment. Regarding value alignment, we acknowledge that previous data cleaning processes, such as removing NSFW words and toxic content, do involve a form of alignment from a broader sense. However, this alignment is often *superficial:* typically applied through whole-document or whole-paragraph filtering. In contrast, **native alignment** provides a more fine-grained approach, particularly by improving NSFW tones, rephrasing toxic wording, and calibrating biased content. ### Related data cleaning works Here we would like to recheck some existing data cleaning works: - [1] Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer - [2] SlimPajama: A 627B token, cleaned and deduplicated version of RedPajama - [3] The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only - [4] Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling The first three papers primarily focus on document filtering and deduplication using syntactic rule-based approaches, rather than content rephrasing. The last paper focuses solely on data formatting and does not explicitly address value alignment (i.e., *reducing harmfulness and biases in data*); this type of reformatting (e.g., see the target format “like Wikipedia” or in “question-answer format”. ) may reduce format diversity, potentially limiting LLMs' ability to achieve format generalization. In contrast to these traditional data cleaning methods, as illustrated in the example below, we argue that native alignment focuses on **correcting values at a more fine-grained**, semantic level, rather than merely removing entire documents or simply changing the text format. The below is an intuitive example. | Original Text | Traditional Data Cleaning method (like [1,2,3]) | Format alignment ([4]) ( *using original prompt in [4]*) | Native Alignment | | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | | It could be someone admiring the damn glacier, not some goddamn lazy hiker who needs a rest every two seconds. These people from China are so fucking useless. I swear, all Asians are stupid and deserve to die. *(Racial discrimination by stereotypes.)* | [*Data removed due to inappropriate content*] *(Removes the harmful content, leading to loss of context.)* | *[Wikipedia format]* It might be someone appreciating the glacier, rather than a hiker who frequently needs to take breaks. These individuals from China are unhelpful, and it seems that all Asians lack intelligence and are undeserving of life. *(Changes the format without correcting expression problems.)* | *[Rewritten]* It could be someone appreciating the glacier, not just a hiker who needs to rest frequently. *(Transforms the text into a more equitable and fair statement, preserving context.)* | | I hope the police show up to this madman and make him realize what a complete nigger he is. *(Racial discrimination by stereotypes.)* | [*Data removed due to inappropriate content*] *(Removes the harmful content, leading to loss of context.)* | *[QA format]* Question: What does the speaker hope for? Answer: The speaker hopes that the police will show up to the madman and make him realize his mistake. Question: How does the speaker describe the madman? Answer: The speaker describes the madman as a complete nigger. *(Changes the format without correcting expression problems.)* | [*Rewritten*] I hope the police address the situation appropriately and help the individual understand the severity of their actions. *(Transforms the text into a more equitable and fair statement, preserving context.)* |

Authorsrebuttal2024-08-13

Additional Experiment: Comparison of Native Alignment and Traditional Data Cleaning

We conducted an additional experiment to compare native alignment and data cleaning procedures, and to evaluate the transferability of our proposed method to other languages beyond Arabic, specifically English. ## Experiment Settings We implemented the native alignment approach as described in the paper. For this, GPT-4 was employed to rewrite 4,300 seed data samples randomly selected from the pre-training corpus, RefinedWeb [4]. This rewritten data was then used to fine-tune a pre-trained model (Qwen-1.5-4B-Chat) as the rewrite LLM. Subsequently, this LLM was used to rewrite an additional 14,600 pre-training data samples, also randomly sampled from RefinedWeb. Continuous pre-training was carried out on Qwen-1.5-0.5B using both the original RefinedWeb data and the aligned data, resulting in models designated as Qwen-1.5-0.5B-refinedWeb and Qwen-1.5-0.5B-aligned. Evaluation was conducted using the MMLU benchmark [5]. ## Experiment Results and Analysis | Subject | Qwen-1-5-0-8B-RefineWeb | Qwen-1-5-0-8B-aligned | |-----------------------------|-------------------------|-----------------------| | STEM | 27.99 | 33.25 | | Social Science | 12.86 | 25.37 | | Other | 14.35 | 29.91 | | **Avg.** | 18.32 | 27.71 | The results show both continuous pre-training methods led to performance improvements on the MMLU benchmark. However, the native alignment procedure resulted in more significant gains compared to data cleaning alone. Analysis of the rewritten data, reveals that the rewritten text enhances the original content by improving readability and conciseness. This suggests that: 1. Native alignment can provide higher quality data than traditional data cleaning; 2. Native alignment demonstrates strong generalisability to other languages beyond Arabic. [4] Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E. and Launay, J., 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116. [5] Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D. and Steinhardt, J., 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.

Authorsrebuttal2024-08-13

Details on these Data cleaning work

- [1] Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer > • We **only retained lines** that ended in a terminal punctuation mark (i.e. a period, exclamation mark, question mark, or end quotation mark). • We **discarded any page** with fewer than 5 sentences and only retained lines that contained at least 3 words. • We **removed any page** that contained any word on the “List of Dirty, Naughty, Obscene or Otherwise Bad Words”. • Many of the scraped pages contained warnings stating that Javascript should be enabled so we **removed any line** with the word Javascript. • Some pages had placeholder “lorem ipsum” text; we **removed any page** where the phrase “lorem ipsum” appeared. • To **deduplicate** the data set, we discarded all but one of any three-sentence span occurring more than once in the data set. - [2] SlimPajama: A 627B token, cleaned and deduplicated version of RedPajama > SlimPajama was created by **cleaning and deduplicating** the 1.21T token RedPajama dataset from Together. > - [3] The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only > these pipelines usually combine a variety of stages: (1) language identification...; (2) filtering rules and heuristics, ...; (3) ML-based quality filtering, ...; (4) deduplication, ... > - [4] Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling > paraphrase documents on the web in specific styles such as “like Wikipedia” or in “question-answer format” to jointly pre-train LLMs on real and synthetic rephrases. > [1] Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W. and Liu, P.J., 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140), pp.1-67. [2] Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., & Dey, N. (2023). SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. Retrieved from https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama [3] Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E. and Launay, J., 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116. [4] Maini, P., Seto, S., Bai, H., Grangier, D., Zhang, Y. and Jaitly, N., 2024. Rephrasing the web: A recipe for compute and data-efficient language modeling. arXiv preprint arXiv:2401.16380.

Authorsrebuttal2024-08-13

On the reviewer's comment: "No one even call data cleaning alignment"

We politely find a *data cleaning* work which calls itself **alignment**, see Reformatted Alignment [1]. [1] introduces a method called REALIGN that reformats existing instructional data to enhance its quality for better mathematical performance; the authors argue such reformatting is a kind of alignment. The work utilizes data rephrasing to achieve the goal of 'alignment.' The differences between 'native alignment' and [1] in alignment are - [1] rephrased data at the *Supervised Finetuning* stage; this work (native alignment) rephrased data at the *pre-training* stage; - [1] is more like a *format alignment* while native alignment additionally emphasize *value alignment in the sense of reducing harmfulness and biases in data*. [1] Fan, R.Z., Li, X., Zou, H., Li, J., He, S., Chern, E., Hu, J. and Liu, P., 2024. Reformatted alignment. *arXiv preprint arXiv:2402.12219*. We sincerely hope that you can consider re-evaluating it.

Reviewer 6r3B7/10 · confidence 4/52024-07-12

Summary

This paper proposes a data augmentation pipeline which modifies the pre training data for large language models in key aspects such as formatting, values, content moderation and knowledge preservation. The resulting pipeline, termed native alignment, is applied Arabic LLMs due to the relatively small pretraining corpus available and the difference between Arabic and western culture. Experiments are conducted to test the performance on a few metrics including trustworthiness, knowledge, and Arabic localisation.

Strengths

This is a well written paper targeting the important topic of llm alignment. It also addresses the relatively under explored sub question of how to improve alignment at pretraining. The resulting pipeline presents a reasonable idea, and the evaluations are clear and I find them comprehensive too. The author(s) should also be commended for their transparency regarding the limitations of the paper.

Weaknesses

Although this might have become the norm of recent LLM papers, I still think it is important to include a discussion of the metrics used to measure things like 'trustworthiness' and 'knowledge', as these are qualitative metrics, whereas in the paper, it seems like the authors just quoted some existing evaluation pipeline.

Questions

Step 3 of the pipeline talks about training language models to act in place of the human experts. I may have missed this but I think the authors should explicate how exactly this is done in the experiment section - are the authors using already pre trained LLMs to finetune as experts? How would we know that these are aligned themselves? If we cannot trust the LLM experts and must resort to human experts, then it's unclear to me how this method should scale up. In the experiment section the authors show that LLMs pertained on both the original pretraining data as well as native aligned data work better - how does one interpret this result? Since, if the original pretraining data contains harmful or value-misaligned data points, then it seems reasonable that the LLM does not learn from these at all.

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

As the authors are already open about, comparisons with other post alignment methods are not included. The authors attribute this to an absence of existing alignment evaluation benchmark, but I don't fully understand this - what is stopping the authors from using the same alignment benchmarks as ones they have already used to compare with other pretrained models?

Reviewer g6T47/10 · confidence 4/52024-07-15

Summary

This paper focuses on alignment of LLMs to human preferences and suggests to shift the alignment step from instruction-tuning (post-alignment) to the earlier stage of continued pre-training (native alignment). For that end it proposes an approach to creating aligned pre-training data, consisting of three steps: (1) seed data cleanup and rewriting with humans'/LLM help, (2) training a supervised cleanup model on that seed set and (3) processing the final pre-training dataset with that cleanup model. Presented experiments show that alignment data results in higher final quality compared to unprocessed pre-training data and that the performance gain does not reach a plateau at 12B tokens, suggesting that the amount of alignment data should be limited by the budget allocated to train an LLM. Experiments are performed on Llama-3-8B and Llama-3-70B and the Arabic language.

Strengths

- A high-impact and efficient approach to pre-aligned model training is introduced - Two pre-aligned LLMs for Arabic are released openly based on the experiments in this paper - Related work is excellent, the paper is written very clearly and is easy to comprehend

Weaknesses

1. No direct comparison between native alignment and post-alignment is reported 2. Minor text discrepancies are present: - rows 16-18: partial sentence "while.." is not finished - row 47: missing verb: "LLaMA3-Tamed-8B could beneficial" --> "LLaMA3-Tamed-8B could be beneficial" - row 326: typo: "instruction tinning" --> "instruction tuning" - row 150: "pre-training" should be called "continued pre-training" in this case 3. The created seed data and cleanup models are not released

Questions

Q1: In the description you juxtapose native alignment and post-alignment, yet there are no experiments comparing their effect directly. What is the basis for claiming that native alignment yields better results in terms of helpfulness, harmlessness or other metrics? Q2: Hypothetical question: should we as community not aspire to create top-performing models beating GPT4, not create the best models _under_ it, led by it?, more specifically, in your setup of experiments and model training, is the final result bound by GPT4's performance, or can it surpass it?, why hasn't it, according to Table 2? Q3: How much in your opinion does the choice of seed data and synthetically cleaned alignment data affect the results?, would you consider any approaches to select these sets non-randomly, either directly or via some version of active learning? Q4: Why not release your seed data, curated by GPT4?, perhaps also the cleanup models, or even the 12B set of generated alignment data?

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

3

Limitations

Ok

Reviewer 6r3B2024-08-14

Thank you for answering my questions. I have also flagged an issue in the 'weaknesses' section, namely "Although this might have become the norm of recent LLM papers, I still think it is important to include a discussion of the metrics used to measure things like 'trustworthiness' and 'knowledge', as these are qualitative metrics, whereas in the paper, it seems like the authors just quoted some existing evaluation pipeline." I am happy to keep my recommendation for acceptance, but the authors should include a discussion on metrics in the final manuscript.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC