In most settings of practical concern, machine learning practitioners know in\nadvance what end-task they wish to boost with auxiliary tasks. However, widely\nused methods for leveraging auxiliary data like pre-training and its\ncontinued-pretraining variant are end-task agnostic: they rarely, if ever,\nexploit knowledge of the target task. We study replacing end-task agnostic\ncontinued training of pre-trained language models with end-task aware training\nof said models. We argue that for sufficiently important end-tasks, the\nbenefits of leveraging auxiliary data in a task-aware fashion can justify\nforgoing the traditional approach of obtaining generic, end-task agnostic\nrepresentations as with (continued) pre-training. On three different\nlow-resource NLP tasks from two domains, we demonstrate that multi-tasking the\nend-task and auxiliary objectives results in significantly better downstream\ntask performance than the widely-used task-agnostic continued pre-training\nparadigm of Gururangan et al. (2020). We next introduce an online meta-learning\nalgorithm that learns a set of multi-task weights to better balance among our\nmultiple auxiliary objectives, achieving further improvements on end-task\nperformance and data efficiency.\n