contact us
What NYC's Local Law 144 bias audit requirement has actually achieved, what a state audit found about its enforcement, and what real hiring-bias research still turns up.
AI Bias Audits: What They Catch, and What a State Comptroller Just Found They Miss

"This AI system passed a bias audit" gets treated, in a lot of coverage, as roughly equivalent to "this system is fair." NYC's Local Law 144 (the first US law of its kind, requiring annual independent bias audits for automated employment decision tools) is a useful real case study in what that claim actually means, and where it falls short.

What the law actually requires

Local Law 144 requires employers using automated hiring or promotion tools to have an independent auditor conduct an annual bias audit, using historical or test data, and to publicly publish the results, including adverse-impact findings and a description of the data and methodology used. Employers must also notify candidates at least ten business days before using such a tool, and publish a summary of audit results. The New York City Department of Consumer and Worker Protection can impose penalties of $500 to $1,500 per day for violations.

The compliance trend is real

Published NYC bias-audit disclosures rose from zero in July 2023, when the law took effect, to roughly 55 by May 2025. Genuine growth in the number of employers actually publishing audit results, even if the absolute number remains small relative to the number of employers plausibly covered by the law.

What a state audit found about enforcement

A December 2025 audit by the New York State Comptroller concluded that current enforcement of the law is "ineffective," citing problematic complaint-handling processes and inaccurate compliance reviews, and found that in two years, the enforcing agency had received exactly two complaints. That's a striking number given how many employers are plausibly covered by the law, and it's a concrete illustration of a gap we've flagged as a structural risk in bias-audit regimes generally: a law can be well-designed on paper and still fail in practice if the complaint and enforcement infrastructure behind it doesn't function.

What real research keeps finding, audit or no audit

Independent research on AI hiring tools continues to surface exactly the kind of bias these audits are meant to catch. Research published through VoxDev in May 2025 found AI hiring tools systematically favored female applicants over Black male applicants with otherwise identical qualifications, and separate research found LLM-based resume screening disadvantaged Black- and female-associated names even when all other resume content was identical. These aren't hypothetical failure modes. They're documented, current findings from real tools in real use.

The legal cases putting real pressure on this

In May 2025, a federal court certified a collective action in Mobley v. Workday, a case alleging Workday's AI hiring tools disadvantaged applicants based on age, race, and disability, now proceeding as a nationwide collective action rather than an individual claim. Separately, in March 2025, the ACLU of Colorado filed a complaint against Intuit and its AI vendor HireVue, alleging an AI interview tool was inaccessible to deaf applicants and likely to perform worse evaluating non-white applicants, including speakers of dialects like Native American English. California finalized regulations in October 2025 clarifying how existing anti-discrimination law applies specifically to AI hiring tools.

Why this matters beyond hiring specifically

This connects directly to the broader alignment challenge we cover in our explainer on the alignment problem: a bias audit, like red-teaming (covered separately), is one concrete, testable piece of a much larger question about whether a system reliably behaves as intended, not a substitute for that larger question, and not a guarantee that stops mattering once a system has technically passed.

What a more meaningful signal actually looks like

NYC's law shows both halves of this clearly: growing real compliance numbers, and a state-level finding that enforcement behind those numbers isn't yet working as designed. The more trustworthy version of "this system was audited" specifies which categories were tested, how recently, under what enforcement regime, and whether that enforcement regime has itself been independently reviewed, a genuinely higher bar than most public claims meet, and, per the Comptroller's own findings, one even a pioneering law hasn't fully cleared yet.

Share with