We explore best practices for training small, memory efficient machine\ntranslation models with sequence-level knowledge distillation in the domain\nadaptation setting. While both domain adaptation and knowledge distillation are\nwidely-used, their interaction remains little understood. Our large-scale\nempirical results in machine translation (on three language pairs with three\ndomains each) suggest distilling twice for best performance: once using\ngeneral-domain data and again using in-domain data with an adapted teacher.\n