Formalising Anti-Discrimination Law in Automated Decision Systems

Algorithmic discrimination is a critical concern as machine learning models are used in high-stakes decision-making in legally protected contexts. Although substantial research on algorithmic bias and discrimination has led to the development of fairness metrics, several critical legal issues remain unaddressed in practice. The paper addresses three key shortcomings in prevailing ML fairness paradigms: (1) the narrow reliance on prediction or outcome disparity as evidence for discrimination, (2) the lack of nuanced evaluation of estimation error and assumptions that the true causal structure and data-generating process are known, and (3) the overwhelming dominance of US-based analyses which has inadvertently fostered some misconceptions regarding lawful modelling practices in other jurisdictions. To address these gaps, we introduce a novel decision-theoretic framework grounded in anti-discrimination law of the United Kingdom, which has global influence and aligns closely with European and Commonwealth legal systems. We propose the “conditional estimation parity” metric, which accounts for estimation error and the underlying data-generating process, aligning with UK legal standards. We apply our formalism to a real-world algorithmic discrimination case, demonstrating how technical and legal reasoning can be aligned to detect and mitigate unlawful discrimination. Our contributions offer actionable, legally-grounded guidance for ML practitioners, policymakers, and legal scholars seeking to develop non-discriminatory automated decision systems that are legally robust.

Paper

Similar papers

Reviewer e8Fx8/10 · confidence 3/52024-06-25

Summary

The paper presents a formalization of fairness metrics intended to ease analysis of discrimination by automated decision making systems in the UK. While there is a relatively applied angle, the bulk of the contribution is intended to be a generic and re-targetable mathematical formalism.

Strengths

This paper shows significant strength in its understanding of nuance with the way law works–something that is sorely missing from the vast majority of CS papers that attempt to handle legal concepts. I was very pleased overall by the mapping the authors performed between relevant legal concepts in the UK and their formal model of fairness. The bulk of the contribution here is in the modelling–which while it results in a simple formulation, should not be taken to undercut the value of the contribution. Non-US legal contexts often get left out of the literature, even common law jurisdictions–yet they impact a significant number of people, and this work takes formalising fairness across that rubicon.

Weaknesses

I do not have any major scientific critiques, though there were areas where the clarity of the paper could improve. Lines 240-274 were written in harder to parse prose than the bulk of the rest of the paper. I had to reread that area multiple times. The case study in Appendix A was actually very useful for understanding the authors' formalism and it is a shape that some of that context was not woven into the paper as concrete examples of how to understand the math. The discussion on proxy discrimination never seemed to finish? I wasn't able to understand its meaning under UK law. Missing a ref to Homer on L299. All these are very minor issues. I'm substantially in favour of accepting this paper.

Questions

Where does proxy discrimination sit under UK law? On L428 the authors make the suggestion to incorporate legitimate features that substantively create the same outcome as the protected features–this seems like a recipe for disaster in a world where we care about the ultimate effects? I'm wondering how we can square the circle there!

Rating

8

Confidence

3

Soundness

4

Presentation

4

Contribution

3

Limitations

Ultimately, adherence to a formalism is *not* what courts generally take into account. While statistical analyses may be used to advance a given line of argument, the standards used are open-textured–and this is an inherent limitation of this line of work. It also would have been good to see where this formalism sits under EU law (or representative EU-member law) or perhaps a discussion of how civil law jurisdictions handle these sorts of issues.

Reviewer Z9767/10 · confidence 1/52024-06-29

Summary

The paper maps existing literature and law on algorithmic fairness onto a decision-theoretic framework. It describes various desiderata (e.g. statistical parity) and legal restrictions (e.g., legitimate aims) in terms of expectations, distributions, estimation error, etc.

Strengths

The paper is well-written and survey a large literature. It appears to state legal tests (particularly under U.K.) with care, while being careful not to overclaim about what its definitions actually establish.

Weaknesses

n/a

Questions

I regret to say that my expertise does not extend to the paper's two principal areas (anti-discrimination law and decision theory). I cannot form a sufficiently educated opinion about the correctness or the novelty of the paper's results. This is my fault, not the authors'.

Rating

7

Confidence

1

Soundness

3

Presentation

4

Contribution

3

Limitations

n/a

Reviewer 4EE73/10 · confidence 2/52024-07-07

Summary

- There is a gap between the definitions of fairness studied in the computer science literature, and the definitions of fairness operationalized by courts adjudicating discrimination claims. This limits the usefulness of the CS definitions. - Amongst work attempting to reconcile legal and computational definitions of fairness, little has focused on anti-discrimination law outside the US. - This paper makes four contributions in this context: - (1) It formalizes elements of anti-discrimination law into a decision-theoretic formalism - (2) If analyzes the legal role of the data-generation process - (3) It proposes conditional estimation parity as a legally-informed target - (4) It provides recommendations on creating SML models that minimize the risk of unlawful discrimination in automated decision-making

Strengths

- The paper’s focus is interesting–the fairness literature is biased towards the US, and I imagine most fairness researchers would be unaware of subtle differences between UK and US anti-discrimination law. - Because UK law is influential around the world, understanding how it regulates fairness in algorithmic systems has global importance.

Weaknesses

- Much of the paper reads like a review of anti-discrimination law. This makes it difficult to parse out (1) what the technical contributions are, (2) why they’re novel, and (3) why they matter. - It’s extremely unclear what the technical payoff of the paper’s modeling choices are. The fairness field is overwhelmed with different definitions/frameworks. Why is the one proposed by the author’s meaningful over others? - It seems like an essential point to the paper’s argument is that prior work hasn’t studied UK anti-discrimination law. But if the paper wants to successfully extend that into an argument about modeling choices, I think it needs to explain why the existing definitions of fairness do not work for UK law. - The recommendations provided are extremely general. Are these new or different from the many recommendations that already exist in the fairness/responsible AI literature?

Questions

The comments in the weaknesses section list the relevant questions!

Rating

3

Confidence

2

Soundness

2

Presentation

2

Contribution

1

Limitations

NA

Reviewer XoDZ6/10 · confidence 4/52024-07-10

Summary

This paper addresses the issues around existing fairness metrics and bias detection/mitigation methods not corresponding with legal notions of fairness, specifically under UK anti-discrimination law. The authors propose a theoretical framework for a data-generating process that aims to formalise the legitimacy of decisions and features in the data. Further, they propose a new metric "conditional estimation parity" which compares estimation errors for different protected groups.

Strengths

1. The paper is well written and coherent. It translates potentially inaccessible legal scholarship and discussions clearly for a technical audience. 2. There is interesting discussion and the paper combines existing literature well. Although these discussions are not particularly novel, UK Equality Law in particular is rarely discussed and the investigations done here are useful to extend the literature for this niche. 3. The work addresses some big limitations in existing literature such as existing fairness metrics not aligning with legal notions of discrimination, particularly under non-US regulations, not considering context of what features are legitimate for an application or considering the estimation errors of decisions.

Weaknesses

1. A lot of the paper is background or a collation of existing literature. The main contribution is the new conditional estimation metric metric but this metric relies on the true DGP and evaluating the estimation error which, as stated, can be complex in practice. This could make it difficult to use the metric in practice. 2. I understand it would be hard to use the metric for evaluating discrimination in existing datasets for the reasons specified above and also due to the inherent context-dependency of the metric (which is a benefit) but it could be useful to include some experimentation or results in a hypothetical scenario to show how it might be used in practice. As there are no results as such to comment on, it is difficult to assess it's significance. 3. The conclusions drawn such as "Assess data legitimacy" or "Build an accurate model", although justified with evidence in the paper, are not novel and are pretty standard, common-sense recommendations. 4. Overall, the main novel contribution is the new metric but this is a small part of the paper. The rest of the paper is a nice collation and narrative of existing literature but I am not sure it significantly advances the field. Other comments: 1. I can't see where SML terminology is introduced - I assume this means supervised machine learning? 2. In Section 1.4, DGPs are mentioned for the first time. It would be useful to have some more background to them before this - what exactly is a DGP? I do not believe it is ever explained.

Questions

1. See weaknesses above. How would you go about using the metric in a real scenario? 2. Do you have any thoughts about how your metric relates to other notions of fairness such as individual fairness metrics? 3. Could you explain "taste-based discrimination" further?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors are honest about the strengths and weaknesses of their work (although some are hidden away and not pointed towards in the checklist). It would be useful to improve the discussion of limitations in Section 1.4 as it only mentions the limitation of applicability only in the UK.

Reviewer iALf7/10 · confidence 4/52024-07-12

Summary

This paper provides a UK-and-European-law-based view of anti-discrimination law as it relates to fair machine learning and automated decision systems. It does a good job laying out the doctrine, arguing correctly that work in this area to-date is very centered on US legal concepts such as disparate treatment vs. disparate impact. Although I am willing to believe that there are subtle differences that drive important aspects of fair ML analysis, as the paper claims, I think the specifics of these differences could be made much clearer and need to be for the paper to have the impact it should. Of particular note, the paper is very well situated in the surrounding literature. Although this contextualization should make the contributions more clearly offset from prior work, as presented I find the opposite: it is difficult to tell what is new as a contribution here. For example, while the contributions are clearly identified in 1.4, I think it would aid the paper if they appeared higher in the intro and were clearer about what is new and why it matters. The example in Appendix A could be used as a running example to show where new concepts are needed and what about existing work does not capture this different legal regime. In particular, after claiming that disparate treatment/disparate impact are distinct to direct & indirect discrimination, the definitions given from 105-114 seem to align tightly to the former. And while I'm not a lawyer, I don't believe that disparate impact claims require a showing of intent under US law either, so I found that distinction somewhat confusing. On the technical level, the discussion of the true data generating process should really be contextualized in the literature on measurement and construct validity, specifically with respect to work by Jacobs & Wallach, which in particular encompasses the material in 2.3 on estimation parity (at least in part). Also, the causal analysis components of the discussion of data generation could cite more of the work of Kohler-Hausmann and also Hu (one paper from these authors is cited, but others are also relevant and speak more directly to causality and counterfactual fairness claims). As a final observation, although the ML community talks in terms of "fair" outcomes, it is often conceptually clearer (and more in line with legal analysis) to use the same techniques as tools for identifying "unfair" activities or outcomes. Phrasing some of the claims this way may condense some arguments and tighten the presentation overall. Related to this, the discussion of these tools as part of an overall practical strategy for risk management is important and should receive more attention. For example, it would be good to discuss how the measures proposed would be used in real legal analysis of an example, such as in litigation or a regulatory proceeding. I was also a bit confused about the analysis of constructed proxies for protected variables in 2.7. I understand that it's necessary to look beyond a formalistic view of whether a specific attribute is considered, but what happens if the proxy for a protected attribute is (say) the sum of two legitimate attributes? Why is it good enough to use only legitimate features? Also, at 393-394 it might be valuable to look at the recent paper on "Less Discriminatory Algorithms" and compare the approaches and outlooks. Incredibly minor: * There is a missing period at 81. * At 284-288, there is a latent call to questions of ecological validity which could be made more explicit

Strengths

* Generalizing beyond the US legal context is important and valuable and this paper does a good job explaining the UK and related legal systems' approach to anti-discrimination law. * The paper is well written and well situated in existing literature

Weaknesses

* Novelty is at times hard to identify. I think it's there, but the claims on what it covers should be clearer. In particular, the discussion of the decision-theoretic framing seems a bit under-attended even though it's potentially very useful. * Some important concepts are missed, notably theories of measurement and construct validity/reliability are at least partially re-invented when they should just be treated as background.

Questions

* Can the example in Appendix A be a running example? * What is new over and above existing literature on ML fairness? How can this novelty be more clearly offset in the presentation? Here I think specifically about the conclusions at 415-430, which seem rather anodyne in light of existing literature.

Rating

7

Confidence

4

Soundness

4

Presentation

3

Contribution

3

Limitations

I believe the limitations are expressed well.

Authorsrebuttal2024-08-06

Rebuttal by Authors (2/2)

> **W2** “Some important concepts are missed, notably theories of measurement and construct validity/reliability are at least partially re-invented when they should just be treated as background.” **R2** We will expand our discussion to include references to relevant work on measurement and construct validity, particularly based on Jacobs & Wallach, incorporating in Section 2.3. We will also incorporate more of Kohler-Hausmann and Hu’s research on causality. Could you please specify which papers by Kohler-Hausmann and Hu you believe are most relevant? We're familiar with the paper by Kohler-Hausmann and Dembroff on causality in the US context, which we could reference in a comparative sense. ## Questions > **Q1** “Can the example in Appendix A be a running example? **A1** Yes, we fully agree with this suggestion and will include specific references to the case study throughout the text. We believe this will also aid the understanding of the practical use of these formalisations and how the measures would be used in real legal analysis. > **Q2** “What is new over and above existing literature on ML fairness? How can this novelty be more clearly offset in the presentation?” **A2** Please see comment in **R1** above. ## Limitations Thank you for confirming that “the limitations are expressed well.” ## Summary We appreciate your thorough feedback and are committed to improving the paper based on your valuable suggestions. We hope you will continue to support our paper for acceptance. Do you have any further questions or comments?

Reviewer e8Fx2024-08-08

Response to Rebuttal

Thanks for the detailed rebuttal. I still maintain my positive score. Regarding A2, if you could really work on clarifying that in the paper, I think it would go a long way to assuaging my concerns. Overall though, nice work!

Reviewer 4EE72024-08-12

Thank you–I really appreciate the response. I will stick with my score however.

Reviewer XoDZ2024-08-13

Thank you for your detailed rebuttal and clarifications and apologies for the late response. I appreciate that this is theoretic work and not applied, however I would expect when talking about the law and fairness there should be some way for the reader to understand how it might be useful in practice. I think adding the existing case study throughout the text will help this. Given this addition and the clarifications on the contributions, I am happy to increase my score as I think this work could be valuable for the Neurips and fairness community. As for adding to the limitations section (Section 1.4) - I would make it clear about the difficulties of finding the true model, or in other words stress the inherent challenges in evaluating the estimation error.

Reviewer iALf2024-08-14

I don't have any other comments or questions, and very much appreciate the responses to the review. I remain very happy with this paper and believe the revision will address the smaller issues raised in all the reviews. I leave as a comment to the area chairs that the issues I raise and the issues raised by reviewer 4EE7 are very similar but our scores are very different. To this, I offer that while the recommendations are very general, using a case study through the paper to answer the critical question of why the difference in legal regimes matters in practice shows why my outlook is positive (also, I believe that contributions need not be a definition or a model, but could be a review of relevant law that unpacks the requirements on how existing work can be applied in a different and understudied context). On the point about making more use of the decision-theoretic framework, adding material about choices between available models and incorporating a running example provides the opportunity to show how the decision-maker's choices can be thought of as decision-theoretic optimization under the legal constraints described here. I believe that change is feasible straightforwardly, since the necessary material is all in the paper (just some is in the appendix). Hu gives commentary on earlier work of Kohler-Hausmann here: https://www.phenomenalworld.org/analysis/disparate-causes-i/. You are likely also aware of another paper focusing specifically on the construction of sensitive groups and the structuring of the data generating process in which they collaborated: https://dl.acm.org/doi/abs/10.1145/3351095.3375674. The general point is that getting at the question of which group differences are causally meaningful vs. which are protected by anti-discrimination law requires either importing an exogenous ontology of protected groups (which the law provides but may not define in as much detail as is needed) or ignoring the contextual construction of subgroup structure in a population for a use case. There are many deep philosophical questions here, and this paper can't reckon with all of them, but it's important to point out where the lines are still blurry to make the scope as clear as possible

Program Chairsdecision2024-09-25

Decision

Reject

© 2026 NYSGPT2525 LLC