In the field of machine learning there is a growing interest towards more\nrobust and generalizable algorithms. This is for example important to bridge\nthe gap between the environment in which the training data was collected and\nthe environment where the algorithm is deployed. Machine learning algorithms\nhave increasingly been shown to excel in finding patterns and correlations from\ndata. Determining the consistency of these patterns and for example the\ndistinction between causal correlations and nonsensical spurious relations has\nproven to be much more difficult. In this paper a regularization scheme is\nintroduced that prefers universal causal correlations. This approach is based\non 1) the robustness of causal correlations and 2) the data not being\nindependently and identically distribute (i.i.d.). The scheme is demonstrated\nwith a classification task by clustering the (non-i.i.d.) training set in\nsubpopulations. A non-i.i.d. regularization term is then introduced that\npenalizes weights that are not invariant over these clusters. The resulting\nalgorithm favours correlations that are universal over the subpopulations and\nindeed a better performance is obtained on an out-of-distribution test set with\nrespect to a more conventional l_2-regularization.\n