We analyze the ability of pre-trained language models to transfer knowledge\namong datasets annotated with different type systems and to generalize beyond\nthe domain and dataset they were trained on. We create a meta task, over\nmultiple datasets focused on the prediction of rhetorical roles. Prediction of\nthe rhetorical role a sentence plays in a case decision is an important and\noften studied task in AI & Law. Typically, it requires the annotation of a\nlarge number of sentences to train a model, which can be time-consuming and\nexpensive. Further, the application of the models is restrained to the same\ndataset it was trained on. We fine-tune language models and evaluate their\nperformance across datasets, to investigate the models' ability to generalize\nacross domains. Our results suggest that the approach could be helpful in\novercoming the cold-start problem in active or interactvie learning, and shows\nthe ability of the models to generalize across datasets and domains.\n