Thank you for reviewing our work, and for your suggestions.
W1: Regarding novelty, we would like to push back on this point. We do not claim to have developed a novel methodology for calculating the environmental impact of training language models. Instead, we aim to set a new standard for reporting the total environmental impact, and encourage other developers in the community to meet this standard going forward. We aim to provide a comprehensive, holistic evaluation, in contrast with many recent technical reports that only evaluate carbon emissions, assume GPUs are always operating at 100% of their maximum power draw, and only report training costs. In other words, we aim to show that this level of detail is feasible to report, and we encourage others to do so as well.
W2: Regarding model size and generalizability, we agree that different model architectures and training setups (hardware, data center locations, etc) would have an impact on the downstream environmental impact. However, the calculations we perform hold at all model sizes, and location-specific variables (PUE, WUE, etc) can be substituted as necessary.
W3 is discussed in our manuscript in both the relevant methods and results subsections, 3.4 and 4.2. To reiterate: Regarding deployment in real world scenarios, we agree that our estimates of the cost of model deployment are limited in comparison to real-world data. However, as we state in the paper, we do not host our own models, and thus do not have access to real world data. Instead, our estimates aim to show potential impact, and we encourage those hosting models at large scales to share similar analyses with their own real-world data in the future. In general, we do not report our precise deployment simulation numbers as part of any central claim we make; instead, we include these results to contextualize the relative costs of training and deployment.
W4 is also discussed in our manuscript (see 5.1, in our paragraph titled “Embodied emissions are still an enigma.”). To reiterate: Regarding embodied emissions, we agree that our estimates likely are not 100% accurate, as we state in the paper. Instead, we aim to provide a better estimate of the embodied emissions compared to previous work, and to highlight how little information regarding embodied emissions is publicly available, which we discuss in Section 5. Additionally, we make many efforts to obtain real information and estimates from our providers (including contacting our data center providers), and we make reasonable, conservative assumptions about information that we were not able to obtain. We aim to be very careful highlighting the aspects of our estimates that are based on assumptions vs. “real” data.
W5: Regarding previously reported environmental impacts, can you explain more about what you mean by “replicating their results,” with regards to OLMo’s and Llama’s carbon emissions? They have not released power consumption data, and thus we must instead take their reported numbers at face value. You do raise a good point though, and we will include estimates of the water consumption *as if their models were trained in our data centers*, as we do not have access to location information for their training runs. In the final version, we will also add comparisons with other models in the deployment estimate section, such as Qwen 2.5.
W6: Regarding model size, we disagree that 7 billion parameter models are not representative of the impact of training larger models. Especially with the growing popularity of deployment-optimized models (such as Gemini Flash, Claude Haiku, GPT-4o mini, etc), we believe that smaller models are only becoming more popular in deployment, especially for on-device settings. However, we have also recently completed training a 13 billion parameter model. We can report that the 13B model, trained to 4 trillion tokens, required about 290 MWh (vs ~157 for the 7B trained to 4T tokens), showing an almost exactly linear trend in training costs as model size grows. We are still calculating other costs for this model, but we will include the full results in the final version.
To answer your questions:
* In the OLMo paper (https://arxiv.org/pdf/2402.00838), they released two separate 7B models, and reported the carbon emissions from both models separately, as they were trained on different hardware in different clusters. We compare against both OLMo models in our paper, but we will make it more clear that these are separate models in the final version.
* We would like to emphasize that training models at the 7B parameter scale and higher is very expensive, in terms of compute, environmental impact, and money. We are training our models to between 2 and 4 trillion tokens, so scaling beyond 7B is very expensive. However, as mentioned above, we have since trained a 13 billion parameter model, and we will include the environmental impact of training this model (also above) in the final version of the paper.