Personalized Adapter for Large Meteorology Model on Devices: Towards Weather Foundation Models

This paper demonstrates that pre-trained language models (PLMs) are strong foundation models for on-device meteorological variables modeling. We present LM-Weather, a generic approach to taming PLMs, that have learned massive sequential knowledge from the universe of natural language databases, to acquire an immediate capability to obtain highly customized models for heterogeneous meteorological data on devices while keeping high efficiency. Concretely, we introduce a lightweight personalized adapter into PLMs and endows it with weather pattern awareness. During communication between clients and the server, low-rank-based transmission is performed to effectively fuse the global knowledge among devices while maintaining high communication efficiency and ensuring privacy. Experiments on real-wold dataset show that LM-Weather outperforms the state-of-the-art results by a large margin across various tasks (e.g., forecasting and imputation at different scales). We provide extensive and in-depth analyses experiments, which verify that LM-Weather can (1) indeed leverage sequential knowledge from natural language to accurately handle meteorological sequence, (2) allows each devices obtain highly customized models under significant heterogeneity, and (3) generalize under data-limited and out-of-distribution (OOD) scenarios.

Paper

Similar papers

Peer review

Reviewer C54F7/10 · confidence 3/52024-07-02

Summary

This paper proposes an approach called LM-WEATHER that utilizes pre-trained language models (PLMs) as foundation models for on-device modeling of heterogeneous meteorological variables. LM-WEATHER enhances PLMs-equipped devices with local weather knowledge through a lightweight personalized adapter. Additionally, it leverages low-rank based transmission to fuse global knowledge among devices, enabling high-efficiency communication.This approach provides an effective solution for modeling real-world, heterogeneous, and continuously updated weather data, without requiring high resource demands. I believe this is a meaningful contribution to the field.

Strengths

This paper extensively compares various time series models across different meteorological variables and weather station. The appendix section of the paper is comprehensive.

Weaknesses

My first concern is that since the personalized adapter is already being used to adapt to the modeling of heterogeneous weather data collected from different weather station devices, what is the point of using the averaging operations during inter-device communication? This requires further explanation. The definitions and formulations of symbols and equations throughout the sections, from Section 3.1 Local Training to Section 4 Theorems, are disorganized and complex. Additionally, there are numerous instances of incorrect singular and plural word usage. It is recommended to rewrite these sections with increased clarity and accuracy, simplifying the content for better comprehension.

Questions

This paper uses GPT-2 as the PLM, which is a very basic pre-trained language model. I am curious why more advanced PLMs like GPT-4 [1], Llama [2], and Vicuna [3] were not considered for better sequence modeling performance. Reference: [1] Achiam J, Adler S, Agarwal S, et al. Gpt-4 technical report[J]. arXiv preprint arXiv:2303.08774, 2023. [2] Touvron H, Lavril T, Izacard G, et al. Llama: Open and efficient foundation language models[J]. arXiv preprint arXiv:2302.13971, 2023. [3] Chiang W L, Li Z, Lin Z, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023.

Rating

7

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

Yes

Reviewer tT2F7/10 · confidence 4/52024-07-05

Summary

This paper proposes LM-WEATHER, which builds upon previous work to explore further the powerful capabilities of PLMs in modelling meteorological variables. By learning sequence modelling capabilities from natural databases and applying them to on-device heterogeneity meteorological variable modelling, a lightweight adapter was developed to provide weather pattern awareness. Additionally, two real-world meteorological station datasets ODW1 & ODW2 support regional weather forecasting. Abundant experiments were designed to verify the feasibility of LM-WEATHER in modelling meteorological variables. The article work is more valuable to further validate the superiority of PLMs in modelling meteorological variables and the provision of weather station data that can be used as a baseline dataset for weather station forecasting.

Strengths

1.This work is more valuable to further validate the superiority of PLMs in modelling meteorological variables. 2.Two real-world meteorological station datasets ODW1 & ODW2, which were collected and compiled from four real-world versatile datasets for on-device meteorological variable benchmark. 3. Exploring the spatio-temporal sequence modeling of weather pattern specificity with high distributional similarity to provide a viable solution paradigm for sparse and heterogeneous weather station data. 4. From the experimental results, this method can further enhance the possibility of maintaining confidence in its future development in the presence of unconstrained data and PLM models.

Weaknesses

1. Figure 1 does not intuitively reflect the contributions of the paper. The Fig.1F mentioned in line 68 seems to be missing; additionally, the explanation for Figure C should be Task Adapter Generation, which is hard to understand. It is recommended to keep the labels a, b, … in the figure consistent with A, B, … in the text. 2. The paper discusses sequence modeling in the time dimension of stations. Can this method be validated for spatiotemporal modeling? 3. The experiments in the paper mainly focus on the accuracy of forecasting time. Can further validation be done on the timeliness and stability of the forecasts? 4. To better present the experimental data, some statistical charts can be drawn to visually demonstrate the advantages of LM-WEATHER in meteorological variable forecasting.

Questions

1. In the Zero-Shot Learning (Out of Distribution Modeling) experiments mentioned in line 282, the domain transfer experiments for regional forecasting on the OWD1 dataset showed that the transfer performance was not as good as the forecasting performance of GPT-4. Can you explain the reason for this? 2. In Theorem 4.2, how is it demonstrated that Low-Rank Matrices ensure privacy? Are there any related research papers on this topic? Can you provide a detailed explanation? 3. Currently, there are many sparse station forecasts. Can interpolation be performed based on sparse station data to further extend the model to spatial dimension forecasting and enhance its generalizability? 4. LM-WEATHER seems to outperform current methods. Can it maintain stability in long-term forecasting, and what is the maximum forecasting time it can achieve? 5. In Table 6, line 305, it appears that having more trainable parameters is better. If the comparison experiment groups had the same number of trainable parameters, what would the results be?

Rating

7

Confidence

4

Soundness

4

Presentation

3

Contribution

3

Limitations

Regarding limitation 1, further validating the impact of increased data scale on model performance using large-scale ERA5 data would be beneficial. Utilizing ERA5 as pre-training data to model invariance in meteorological forecasting and learning fundamental meteorological variable patterns, followed by fine-tuning with real-world data, could be a highly valuable endeavor.

Reviewer mBbn8/10 · confidence 5/52024-07-13

Summary

The paper aims to develop weather foundation models by leveraging the on-device data on many distributed sensors. Specifically, a federated learning and low-rank adaption mechanism have been applied to a time-series-based foundation model training on Meteorology data.

Strengths

1) The paper proposes a new way to develop weather foundation models by leveraging many distributed devices with collected Meteorology data. It is worth noting that weather foundation models are an emerging research domain critically important to geoscience and tackling climate change. 2) The proposed solution is technique sounds. The procedure of time series processing, LoRA-based adaptation, and federated learning are seamlessly integrated to implement its goal. 3) The paper provides sufficient details with an appendix and open-sourced codes. It is an essential part to ensure the reproducibility of this work. 4) The paper conducted a comprehensive experiment for comparison and evaluation of the few/zero-shot learning scenarios on the foundation model.

Weaknesses

1) The paper needs a strong justification to describe the motivation for using a pre-trained language model as a basis for the weather foundation model. Are there any texts in the Meteorology dataset? 2) The paper’s contents organization could be improved. For example, the figure 1 is too complex to understand. In Eq 2, it would be better to write the function as F(theta, D). Because F(theta | D) is easily confusing with the conditional distribution. Moreover, some technique details should not be placed in the appendix. It is a usual assumption that the main paper is self-contained so that the readers can understand the paper without reading the appendix. 3) In the contribution part (lines 73 - 94), the last two items look like the advantages of the proposed method rather than contributions.

Questions

1) Does the proposed method only focus on time series data? 2) Would you please highlight the main research question in plain language? It seems there are many components mixed in the paper. 3) In Section 3, what’s the federated aggregation mechanism? Is there any pseudo code to describe the algorithm? 4) Can you please explain the token embedding for time series data? Is the token to be described as a special pattern of a subsequence of the time series? How to define the vocabulary of the tokens?

Rating

8

Confidence

5

Soundness

4

Presentation

3

Contribution

4

Limitations

N/A

Reviewer Jjnz6/10 · confidence 4/52024-07-13

Summary

This paper introduces LM-WEATHER, a framework leveraging pre-trained language models (PLMs) for on-device meteorological variable modeling. The framework integrates personalized adapters into PLMs, enhancing their ability to handle heterogeneous weather data efficiently. Key contributions include superior performance in forecasting and imputation tasks, minimal communication overhead, and maintaining privacy during client-server interactions. Experiments on real-world datasets demonstrate LM-WEATHER's effectiveness across various scenarios, including few-shot and zero-shot learning.

Strengths

S1. LM-WEATHER introduces a novel approach by integrating personalized adapters into PLMs, allowing for efficient and accurate on-device meteorological modeling. S2. The framework addresses heterogeneity in weather data and demonstrates robust performance in various tasks and scenarios. S3.The approach offers a promising solution for personalized weather modeling on resource-constrained devices, potentially benefiting various applications in meteorology and related fields.

Weaknesses

W1.In the communication section, it is mentioned that only the low-rank matrix parameters are transmitted between the client and server, but specific details about the transmission mechanism and strategy are missing. This includes transmission frequency, bandwidth consumption, and potential issues in practical applications. These missing details may prevent readers from fully understanding the feasibility and effectiveness of the communication strategy. W2. In the parameter adapter section, it describes integrating the generated adapters with the original input through FFN to produce the final weather sequence predictions. However, the specific architecture of the FFN, how its parameters are configured, and why this particular architecture was chosen are not detailed. The lack of these key details might leave readers questioning the mechanism and effectiveness of the parameter adapter. W3. In Section 4, while Theorems 4.1 and 4.2 propose the rationality of time series decomposition and privacy assurance through low-rank matrix exchange. There is no introduction and description about the background and application of these Theorems. Also, the descriptions are too brief, making it difficult for readers to fully understand the theoretical basis. W4. The paper's experimental results are heavily supplemented by a 37-page appendix, which can hinder readability and accessibility of the core content. Important experimental details and results should be better integrated into the main body of the paper to facilitate easier comprehension and evaluation by the readers.

Questions

Q1.Could you provide more detailed proofs and theoretical validation for the decomposition rationality and privacy assurances mentioned in the theorems? Q2.Could you clarify the potential confusion in the notations n¬i and nk in Section 2.2, and ensure the formulas are correctly presented? Q3.Could you provide detailed information on the architecture and parameter configuration of the FFN used in the parameter adapter, and explain why this architecture was chosen? Q4.Could you provide more detailed information about the datasets, particularly whether data from extreme climatic regions(e.g., tropical rainforests and deserts) are included? If not, how do you plan to expand these datasets to cover more global climatic conditions?

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

1 The authors acknowledge the limitations in the diversity of the datasets and the potential challenges in modeling specific climatic conditions not covered in the current study. Future work should focus on expanding dataset coverage and providing more detailed theoretical and empirical validation. Additionally, the practical applications and limitations of few-shot and zero-shot learning capabilities should be more thoroughly explored and documented. 2 It would be much better integrate the essential experimental details and results into the main body of the paper to improve readability and comprehension? 3 I have questioned whether there will be a powerful foundation model for time series prediction. If the assumption is not true, the application of the proposed solution is limited.

Reviewer tT2F2024-08-11

Thank you for your response. This approach offers an efficient solution for modeling heterogeneous and continuously updated real-world weather data, while minimizing resource demands. I believe this is a significant contribution to the field.

Reviewer C54F2024-08-13

I appreciate the authors' response, which addresses most of my concerns. I am happy to raise my rating.

Authorsrebuttal2024-08-14

Dear Reviewer C54F, Thank you for raising the rating of our paper. We are happy to have addressed your concerns. Best regards, Authors

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC