With the growing use of ML in highly consequential domains, quantifying\ndisparity with respect to protected attributes, e.g., gender, race, etc., is\nimportant. While quantifying disparity is essential, sometimes the needs of an\noccupation may require the use of certain features that are critical in a way\nthat any disparity that can be explained by them might need to be exempted.\nE.g., in hiring a software engineer for a safety-critical application,\ncoding-skills may be weighed strongly, whereas name, zip code, or reference\nletters may be used only to the extent that they do not add disparity. In this\nwork, we propose an information-theoretic decomposition of the total disparity\n(a quantification inspired from counterfactual fairness) into two components: a\nnon-exempt component which quantifies the part that cannot be accounted for by\nthe critical features, and an exempt component that quantifies the remaining\ndisparity. This decomposition allows one to check if the disparity arose purely\ndue to the critical features (inspired from the business necessity defense of\ndisparate impact law) and also enables selective removal of the non-exempt\ncomponent if desired. We arrive at this decomposition through canonical\nexamples that lead to a set of desirable properties (axioms) that a measure of\nnon-exempt disparity should satisfy. Our proposed measure satisfies all of\nthem. Our quantification bridges ideas of causality, Simpson's paradox, and a\nbody of work from information theory called Partial Information Decomposition.\nWe also obtain an impossibility result showing that no observational measure\ncan satisfy all the desirable properties, leading us to relax our goals and\nexamine observational measures that satisfy only some of them. We perform case\nstudies to show how one can audit/train models while reducing non-exempt\ndisparity.\n