Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting

In this paper, we introduce a novel theoretical framework for multi-task regression, applying random matrix theory to provide precise performance estimations, under high-dimensional, non-Gaussian data distributions. We formulate a multi-task optimization problem as a regularization technique to enable single-task models to leverage multi-task learning information. We derive a closed-form solution for multi-task optimization in the context of linear models. Our analysis provides valuable insights by linking the multi-task learning performance to various model statistics such as raw data covariances, signal-generating hyperplanes, noise levels, as well as the size and number of datasets. We finally propose a consistent estimation of training and testing errors, thereby offering a robust foundation for hyperparameter optimization in multi-task regression scenarios. Experimental validations on both synthetic and real-world datasets in regression and multivariate time series forecasting demonstrate improvements on univariate models, incorporating our method into the training loss and thus leveraging multivariate information.

Paper

References (57)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer EGU14/10 · confidence 3/52024-07-09

Summary

This paper analyzes a linear multi-task regression framework. It derives formulas for asymptotic train and test risk. The formulas provide insights in how the raw data covariances, singal-generating hyperplanes, noise levels, and size of data sets affect the risk. Motivated by the analysis on the linear framework, experiments on multivariate time series forecasting are performed to investigate the effectiveness of MTL regularization.

Strengths

- The paper is clearly written. It outlines main points and is easy to follow. - Throughout analyses, which use random matrix theory, on the multi-task linear regression framework are performed. Specifically, the asymptotic train and test risks are derived for analyzing the behavior of the framework. - The analyses provide useful insights into the behavior of models.

Weaknesses

- The analyses are performed on a simple linear model. They do not apply to general nonlinear models, which are much more common in real practice. - There seems to be a disconnection between the theoretical analysis and the experimental results. The experiments seem to only highlight the effectiveness of MTL regularization and have no connection with the theoretical analysis.

Questions

- Is there any connections between the theoretical analysis and the experiments? E.g. making use insights from the analysis to improve experimental results. - Is splitting the linear operator into a local and global term a novelty of this paper? - Similarly, is this paper the first work to apply this MTL framework (decompose prediction as a sum of global and local term) to nonlinear models? - How are the regularization hyperparameter chosen in the experiments? - In some of the results, incorporating MTL regularization worsen the performance. Why is that?

Rating

4

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

N/A

Reviewer Zuy68/10 · confidence 4/52024-07-12

Summary

The authors analyse the problem of multi-task regression (for a linear model) under the assumption of concentrated random vectors. The results obtained provide a comprehensive description of the performance of the model (including critically its generalisation performance). The authors use this result for hyperparameter optimisation and show on real-world datasets that using the insights from their theoretical model they obtain state-of-the-art performance on multivariate time series forecasting.

Strengths

I feel that this is a really nice achievement. Computing generalisation results are notoriously difficult. Such calculations are usually confined to rather simplistic models. The model chosen is rather simplistic in being linear and make, what seems to me, a very strong assumption about the data. Nevertheless, by choosing the appropriate setting (i.e. a multi-task scenario) and it appears that the concentration assumption seems to innocuous they come up with a useful result. Given how hard it is to do meaningful theory in machine learning I certainly believe this paper deserves to be accepted. It should also be noted, that the paper is technically accomplished.

Weaknesses

The paper will not be everyones taste. The mathematics is not easy to follow, but that is often the nature of making a non-trivial technical contribution. In the equation for $\lambda^*$ immediately preceding section 4.4 there is a missing norm symbol.

Questions

Is there an intuition on why the concentrated random vector assumption justified? Is it that real data satisfies this assumption accurate, or is that the deviations from this assumption don't seem to be relevant? You are tackling real world data using a relatively simple linear model and beating what I assume are sophisticated non-linear models. I find this surprising. Is there an explanation of why you do so well? (Are the datasets you are using very simple--linear?, are the number of examples small compared to the number of features so that more complex models will overfif?).

Rating

8

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

This is fine.

Reviewer 8z1C5/10 · confidence 4/52024-07-12

Summary

The authors characterise the train and test risk of multi-task regression using random matrix theory. Assessment via Figures 2 and 3 shows a good match between theory and empirical. This motivates a regularised objective for learning mutivariate time-series.

Strengths

The theory contribution in section 4 is significant and original, and the assessment in sections 5.1 and 5.2 are significant.

Weaknesses

*Three major weaknesses* A. There are at least two existing works on the theoretical error bounds for multi-task learning [Chai, NeurIPS 2009; Aston & Sollich, NeurIPS 2012]. While these are using Gaussian processes as a basis, comparison or remarks on the general effect of multi-task learning should be discussed with reference to two existing works. This is especially when all these are essentially linear-in-regressors models. B. The proof or at least outline of the proof of Theorem 1 *should* be in the main paper, given that this is *the* major contribution of the paper. C. It is totally unclear the relevance and contribution of the theory itself to section 5.3. Adding a cross-task regularization is already known to help (mostly). *Minor, but affects clarity* - The introduction to the paper is on multi-task multi-response/regression. The multi-response/regression part is not made clear. - Why is there a need to divide by $\sqrt{Td}$ in (1)? Is it to make analysis easier. - There are two different uses of the letter $t$ in Assumption 1. - The setup in section 5.3 needs to relate back to the model in section 3.1 to make clearer their connections.

Questions

1. The analysis of the noise term in section 4.2 says that increase in noise negatively affects transfer. However, it is precisely in the presence of noise that borrowing statistical strength from other tasks is suppose to help. How do we reconcile this? 2. In section 4.3, is it clear that $C_{MTL}$ is always positive, negative or no such conclusion can be made?

Rating

5

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

Authors should include some technical limitations of their current work.

Reviewer 7acX6/10 · confidence 2/52024-07-15

Summary

The authors derived theoretical insights into the train and test risks of the multi task regression loss using Random Matrix Theory.

Strengths

The closed form solution of the optimization parameter using RMT is novel. The authors set a trend in MTR to gain analytical expressions for optimization parameter and error risks.

Weaknesses

1. The paper lacks clarity for instance on line 89, I am not sure what is meant by general optimization and other mathematical tools?

Questions

1. What is $\gamma$ in (3)? 2. Have you studied the implications of the assumptions made to use random matrix theory to derive the expressions of asymptotic train and test errors? 3. Shouldn't it be $\mathbf{Y}\in \mathbb{R}^{Tq\times n}$ in the equations after line 110?

Rating

6

Confidence

2

Soundness

2

Presentation

1

Contribution

4

Limitations

Limitations are not discussed. My guess is that the assumptions made to be able to employ Random Matrix Theory need to be discussed.

Reviewer Zuy62024-08-08

Acknowledgement of feedback comments.

Thank your for addressing questions that I appreciate don't have easy answers. I maintain my believe that your paper represents a strong technical contribution and that using theory as a guide to improving machine learning algorithms is an achievement. Good luck with convincing the other reviewers and don't be discouraged if the outcome is negative. The type of work you are doing is technically challenging and difficult to communicate, but I believe it is beneficial to the field.

Authorsrebuttal2024-08-12

Thank you for your encouraging feedback and for recognizing the technical challenges of our work. We greatly appreciate your support and belief in the value of our approach.

Reviewer 8z1C2024-08-08

C. You wrote "Our results show that the test risk curves for non-linear models follow similar patterns to those predicted by our theory." It is important of show this test risk curves to proof your point. I will increase the score. Content once, it seems that everything is there based on your answers. However, this does suggest some rewrite of the paper which is not insignificant.

Authorsrebuttal2024-08-12

Thank you for your constructive feedback and for increasing the score. We have incorporated your suggested changes into the main paper and believe they greatly enhance its clarity. In addition, please note that we have plotted the test risks in Appendix (Section G.4). Please let us know if you have any other question. Your insights are invaluable to improving the paper.

Reviewer 8z1C2024-08-12

Thanks. For the test risks plot, either put it in the main paper, or point the author to the appendix from the main paper. You cannot assume we read the appendix.

Authorsrebuttal2024-08-13

Official Comment by Authors

Thank you for your advice! We will revise the paper with respect to this point.

Reviewer 7acX2024-08-11

Thank you for addressing my concerns. I understand the paper better now. I have no doubt that the proposed method certainly has novel theoretical contributions. I am raising my score.

Authorsrebuttal2024-08-12

Thank you for your thoughtful feedback and positive evaluation of the paper's theoretical contributions.

Authorsrebuttal2024-08-12

We sincerely appreciate your insightful feedback, which has significantly contributed to the improvement of our paper. We believe that we have carefully addressed your comments and concerns, and we hope our answers meet your expectations. As the reviewer-author discussion period is nearing its end and our window for responding will close soon, we kindly ask you to let us know if you have any additional points that need to be clarified, so that we could engage in further discussion. Thank you again for your valuable review. We look forward to hearing from you and hope our efforts align with your suggestions.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC