Insight

The Robert Williams Case: What Facial Recognition Bias Really Costs

July 27, 2026


In January 2020, Detroit police officers pulled into Robert Williams's driveway, arrested him in front of his wife and two young daughters, and held him overnight without telling him why. The reason, when he finally learned it during questioning, was a facial recognition match: surveillance footage of a shoplifting suspect had been run against a database of driver's license photos, and the system returned Williams's face as the closest hit.

Williams had never set foot in the store. The match was wrong. According to the complaint, when an officer finally placed the surveillance still next to Williams's face — side by side on a steel table — the mismatch was obvious enough that one officer reportedly said something like "Oh, I guess the computer got it wrong." But the computer had not gotten it wrong. It had compared two images, calculated a similarity score, and surfaced a candidate from a database — exactly what it was built to do. The error was not in the algorithm's execution. It was in the chain of human decisions that received a probability score, treated it as enough, and let it travel unchallenged from a vendor's server to a detective's desk to a magistrate's signature to a man's wrists in front of his daughter.

How a Face Becomes a Statistic

Facial recognition systems do not compare photos the way a person eyeballing two pictures would. They convert a face into a set of numerical measurements, then search a database for the closest numerical match. "Closest" is a statistical judgment, not a certainty — the system returns a ranked candidate, not a verified identity.

NIST — the U.S. National Institute of Standards and Technology — defines bias, in its 2022 Special Publication 1270, as "an effect that deprives a statistical result of representativeness by systematically distorting it, as distinct from a random error, which may distort on any one occasion but balances out on the average." Error scatters. Bias points in a direction. A facial recognition system that misidentifies one person, once, in a way unconnected to any group, is producing error. A system whose misidentifications cluster — around women, around older adults, around people with darker skin — is producing bias.

The Bias Was Not a Rumor — It Was Measured

In 2018, MIT Media Lab researcher Joy Buolamwini and Microsoft Research's Timnit Gebru published Gender Shades, testing commercial facial analysis systems already sold to law enforcement, airports, and corporations. Accuracy was not uniform: on the worst-performing system, the error rate for darker-skinned women exceeded thirty percent, while for lighter-skinned men on the same system it was under one percent.

A year later, NIST examined the market directly, publishing Face Recognition Vendor Test Part 3: Demographic Effects — 189 algorithms from 99 developers, drawn from millions of comparisons. False-positive rates varied across demographic groups by factors NIST described as ranging from negligible to two orders of magnitude, and the report added a sentence the agency rarely allows itself: these errors can lead to "severe consequences," including, explicitly, wrongful accusations in law enforcement. As alleged in the Williams complaint, Detroit's facial-recognition policy at the time lacked quality-control standards, peer review, adequate training, and clear rules for corroborating leads — the institutional architecture that could have absorbed the very risk NIST had already mapped simply was not there.

A Lead, Not a Verdict — and Detroit's Own Fix

The most consequential outcome of the Williams case was not the lawsuit, though there was one. It was a policy change. Following the settlement, Detroit's revised policy states that a facial recognition result is an investigative lead only — never a positive identification — and requires independent, reliable evidence before any photo lineup can proceed. The system may suggest; the institution must prove.

That is a governance fix, not a technical one. It does not require a better algorithm. It requires an institution to accept, in writing, that a probabilistic output is not a fact — and to build a process that stops a human being from skipping the verification step because the software sounded confident.

Three Checks Before Trust

Before deploying any algorithmic system whose output touches criminal justice, hiring, lending, housing, or healthcare, the book argues the people responsible should be able to answer three questions — not in principle, but in writing, with evidence.

  • What was this system built to do — and is that still how it's being used? A tool built to surface investigative leads is not the same as one built to make identifications; a model designed to assist a human decision is not the same as one allowed to replace it.
  • When does the number become action — and who can stop it? The threshold that converts a score into a warrant, a hire, or a denial must be documented and open to challenge, and a human reviewer needs real authority to disagree with the machine, not just a button that says approve.
  • Who does the system fail — and who pays when it does? Aggregate accuracy is a marketing number; performance has to be measured across the populations the system actually affects, and the cost of error cannot fall only on the person harmed.

Why This Belongs in a Privacy and AI Governance Conversation

Facial recognition used for law enforcement or identity verification sits squarely inside the "high-risk" category that frameworks like the EU AI Act are built around, and it is exactly the kind of automated processing that triggers a Privacy Impact Assessment obligation under the GDPR, the LGPD, and similar laws. A PIA done before deployment — one that actually asks what the system's error rate is and who it errs against — is the paperwork that could have flagged this risk before Robert Williams was arrested in his own driveway, not after.

This is the pattern this series keeps returning to: an assessment done honestly, before the system goes live, is cheaper than the version of accountability that arrives afterward, in a courtroom, with a family's name attached to it. If your organization is evaluating any AI system that classifies, scores, or identifies real people, the same three checks — known error rate, human verification, contestability — are the starting discipline, and they are the kind of judgment PIA Studio is being built to support at scale.


← Back to Algorithmic Bias