We advocate for a practical Maximum Likelihood Estimation (MLE) approach\ntowards designing loss functions for regression and forecasting, as an\nalternative to the typical approach of direct empirical risk minimization on a\nspecific target metric. The MLE approach is better suited to capture inductive\nbiases such as prior domain knowledge in datasets, and can output post-hoc\nestimators at inference time that can optimize different types of target\nmetrics. We present theoretical results to demonstrate that our approach is\ncompetitive with any estimator for the target metric under some general\nconditions. In two example practical settings, Poisson and Pareto regression,\nwe show that our competitive results can be used to prove that the MLE approach\nhas better excess risk bounds than directly minimizing the target metric. We\nalso demonstrate empirically that our method instantiated with a well-designed\ngeneral purpose mixture likelihood family can obtain superior performance for a\nvariety of tasks across time-series forecasting and regression datasets with\ndifferent data distributions.\n