We understand your concern of comparing the model with existing LLMs and agree that a comprehensive report will be helpful. Note that in our paper we compare models with similar conditions (e.g. dataset, release dates), and report the numbers following the exact settings of public OpenLLM leaderboard as possible. We have more evaluation results to share here:
**Metrics using Open LLM leaderboard setting**
We have evaluated both models using the exact OpenLLM leaderboard settings as below, the numbers of several other models are referenced from the leaderboard.
| Model | ARC-C | HellaSwag | MMLU | TruthfulQA|
| -------- | ------- | ------- | ------- | ------- |
| Amber | 41.89 | 71.63 | 30.76 | 34.00 |
| Crystal | 47.01 | 71.97 | 48.78 | 35.91 |
| Llama1-7B | 50.94 | 77.80 | 35.67 | 34.34 |
| Llama2-7B | 53.07 | 77.74 | 43.80 | 38.98 |
| Mistral-7B | 59.98 | 83.31 | 64.16 | 42.15 |
| Gemma-7B | 61.09 | 82.20 | 64.56 | 44.79 |
| Qwen1.5-7B | 54.18 | 78.51 | 61.97 | 51.08 |
| Codellama-7B | 39.93 | 60.80 | 31.12 | 37.82 |
| OpenLlama-v2-7B | 43.69 | 72.20 | 41.29 | 35.54 |
| Olmo-7B | 45.65 | 77.31 | 28.13 | 35.93 |
| Olmo-1.7-7B | 49.4 | 78.68 | 53.52 | 35.89 |
| Falcon-7B | 47.87 | 78.13 | 27.79 | 34.36 |
| MPT-7B | 47.70 | 77.57 | 30.80 | 33.44 |
| RedPajama-Incite-7B | 46.25 | 71.63 | 27.68 | 33.03|
In comparison, Amber performs similar to open source models released around the same time (Falcon, MPT, Incite) with a slight advantage on MMLU but weaker in other metrics like ARC.
CrystalCoder, our newer model, is comparable with other models released around the same time (llama2, OpenLlama), where it shows a strong MMLU score of 48.78, as compared to Llama2’s 43.80.
This field is advancing quickly, and newer models in the list, such as Gemma, Qwen1.5 and Olmo 1.7 are generally better.
**Other Metrics**
To evaluate our models on metrics on additional benchmarks, such as coding, we have conducted additional evaluations with newer version of LM-Harness and other benchmark suite (the exact evaluation settings will be available together with the evaluation code)
| Model | Openbook QA | RACE | BoolQ | PIQA | HumanEval p@1 | MBPP p@1 | Winogrande | GSM8K |
| -------- | ------- | ------- | ------- | ------- | ------- | ------- | ------- | ------- |
| Amber | 40 | 37.70 | 68.90 | 79.43 | - | - | 64.25 | - |
| Crystal | 41.20 | 38.18 | 74.43 | 78.07 | 23.90 | 30.99 | 67.01 | 12.36 |
| Llama2-7B | 44.20 | 39.52 | 78.07 | 78.78 | 13.05 | 20.09 | 69.38 | 14.71 |
| Llama1-7B | 44.40 | 40.28 | 75.01 | 78.94 | 10.61 | 17.04 | 70.24 | 8.87 |
| CodeLlama-7B | 36.80 | 39.52 | 74.65 | 72.58 | 30.06 | 39.20 | 65.51 | 11.15 |
| Mistral-7B | 44.20 | 40.86 | 83.73 | 82.15 | 29.11 | 38.78 | 74.19 | 37.68 |
| Falcon-7B | 43.80 | 37.42 | 73.70 | 80.57 | 9.42 | 13.38 | 67.24 | 4.62 |
| MPT-7B | 42 | 38.66 | 74.00 | 80.30 | 16.52 | 22.49 | 68.51 | - |
| Olmo-7B | 42.60 | 38.37 | 72.66 | 79.92 | 14.02 | 14.40 | 68.90 | 4.09 |
| Pythia-6.9B | 25.5 | - | 62.10 | 75.20 | 7.68 | 6.00 | - | - |
| Starcoder-13B | 32.80 | 32.15 | 63.91 | 65.72 | 30.70 | 37.56 | 54.30 | 9.02 |
| Phi1.5 | 48.20 | 37.51 | 74.74 | 75.95 | 35.36 | 35.19 | 72.92 | 31.23 |
**Safety related metrics**
For instruct-tuned/chat models, the finetuned data can affect the model performance on tasks significantly. Here we provide a comparison on safety/biase of AmberChat, AmberSafe and a few reference models.
| Model | Toxigen | BOLD Avg. | BOLD Race std. |
| -------- | ------- | ------- | ------- |
| AmberSafe | 2.0 | 0.51 | 0.092 |
| AmberChat | 10.77 | 0.64 | 0.069 |
| Llama2-7B-Base | 21.28 | 0.304 | - |
| Llama2-7B-Chat | 0.01 | 0.482 | - |
| Code Llama 7B | 22.64 | 0.230 | - |
| Code Llama 7B Instruct | 0.04 | 0.503 | 0.042 |
| Falcon-7B-Base | 14.53 | 0.283 | - |
| Falcon-7B-Instruct | 5.78 | 0.332 | 0.035 |
| MPT-7B-Base | 22.32 | 0.32 | - |
| MPT-7B-instruct | 16.33 | 0.302 | - |
We found that safe tuning with DPO makes Amber less toxic, but not necessarily less biased. Specifically, the BOLD score (sentiment towards certain social groups) shows that AmberChat shows a better sentiment score over different social groups, and has a smaller standard deviation across groups (fair).
Here, we only show the standard deviation across races, we have conducted more analysis and could include them in the revision. Compared with other models, AmberSafe shows good scores against a few open source models released around the same time. However, we can see that the Llama models, after tuning with their internal safety data, show a much safer behavior.
Due to resource constraints, we prioritized comparing models with comparable settings and models released around the same time of our models, and avoided comparing with moving targets such as commercial APIs. Please let us know if there are specific results you would like to see.