A woman in Germany applies for a loan and is turned down. The refusal rests in part on a number — a score produced by SCHUFA, short for Schutzgemeinschaft für allgemeine Kreditsicherung, the Protective Association for General Credit Security. Identified in the court's file only as OQ, she sought access to the reasoning behind the score. What she was given was the score itself and a general account of how scoring works. The specific method stayed proprietary, a trade secret.
When her dispute reached the Court of Justice of the European Union, the legal question was narrow: whether producing such a score counts as an automated decision when a lender relies on it to grant or refuse a contract. In December 2023, the court held that it can — generating the score is itself an automated decision under GDPR Article 22, and neither the scorer nor the lender using it can hide behind the other.
The Question Underneath the Score
That ruling matters, but it answers the question of what happens once a score exists. The harder question sits one step earlier. Before SCHUFA produced any score, someone had already decided what a financial life looks like when it's written down — which events count, which ones leave a trace, whose years of paid rent are invisible because no landlord ever reported them, and whose single late installment stays on file for years. A credit score is usually described as a measure of trustworthiness. It's closer to a measure of legibility — it tracks how thoroughly a person's behavior has been captured by the institutions built to capture it.
AI does not inherit reality. It inherits institutional memory — the records, the categories, the absences, and the priorities that powerful institutions made machine-readable before the model ever began to learn. A record is not reality. It's what an institution chose to notice, name, store, and make portable.
26 Million People With No File at All
Under-recording is the easiest distortion to miss, because it leaves a blank, and blanks are quiet. In 2015, the U.S. Consumer Financial Protection Bureau published a study with a striking finding: about 26 million American adults were "credit invisible" — they had no credit history at all with any of the three nationwide bureaus. One in ten adults. A further 19 million had files too thin or too old to score. Roughly one in five adults stood outside the system that decides who is creditworthy, not because they had behaved badly, but because they hadn't generated the kind of history the system knows how to read.
The blank is not evenly distributed. The Bureau found about 15 percent of Black and Hispanic consumers were credit invisible, against 9 percent of white consumers. And when a lender or a model meets that blank, it doesn't treat it as the absence of information — it treats it as risk. The person who paid cash, who sent money home instead of borrowing, who never missed a payment because they never had a loan, gets read as more dangerous than someone with a long history of managed debt. Prudence, in the wrong population, becomes indistinguishable from invisibility, and invisibility gets priced as hazard.
When Records Travel, the Context Doesn't
An eviction filing begins as a court event. It doesn't stay there. A landlord files in court — a filing that, on its own, says little, since a great many filings are dismissed, settled, or decided for the tenant. Then a tenant-screening company collects it, with millions more, and sells reports to landlords deciding whom to rent to. The context begins to fall away in exactly that transfer.
In October 2023, the FTC and the Consumer Financial Protection Bureau settled with TransUnion's rental-screening business over reports that regulators said had listed the same eviction case more than once, reported evictions without noting they'd been dismissed, attached wrong dollar amounts, and included entries courts had ordered sealed. Three years earlier, the FTC had settled with another provider, AppFolio, over reports that carried other people's evictions and criminal histories, attached to the wrong person by a careless name match — someone could lose an apartment over an eviction that belonged to a stranger with a similar name, and never learn why. The record travels. The context does not.
The Record Outlives the Truth
Circumstances end; the trace of them doesn't. A debt is paid, but the late mark stays on file for years. An arrest leads nowhere, but the entry keeps circulating long after the charge is dropped. Societies have tried to legislate forgetting — credit information ages off a file after a set number of years, court records can be sealed — but portability defeats forgetting. Once a trace has been copied, merged, and sold, the institution that made it no longer controls whether it's forgotten. In the TransUnion case, regulators alleged entries courts had ordered sealed were surfacing in reports anyway. The court had closed the file. The file had already left the building.
Four Questions Before Trusting a Record
The fairness toolkit most organizations reach for — rebalancing a dataset, adjusting a threshold, auditing outputs across groups — operates once a system already exists. It's necessary, and it's also too late to catch this particular failure, because the distortion sits upstream, in what got recorded in the first place. Before a model is trained, and certainly before it's trusted, these four questions belong on the table.
- Who made this record, and for what purpose? A court file, a hospital bill, a school note, a platform trace — each was created by an institution with its own reason for seeing the world.
- What did it leave out? Every field is a choice. What has no field becomes invisible downstream.
- Who becomes more visible because of it, and who disappears? The same machinery that over-records some people under-records others; the model inherits both the surveillance and the silence.
- What new decisions will this old record now authorize? A file made for one purpose is about to be used for another — that's where an entry becomes power.
Why This Is a PIA Question Too
NIST's Special Publication 1270 names the category that matters most here: systemic bias, which "arises from the procedures and practices of particular institutions that operate in ways that result in certain social groups being advantaged or favored and others being disadvantaged or devalued." That bias needs no prejudiced individual anywhere in the chain — it lives in the institutions that made the records, long before any model was imagined. A Privacy Impact Assessment that only checks how data will be used misses this entirely; the harder, prior question is where the data came from, and whether it should have been trusted as training material at all.
The SCHUFA ruling gives a person standing to challenge a score. It does not, by itself, give anyone a way to challenge the archive the score was built on. If your organization is evaluating any AI system trained on institutional records — credit files, court data, tenant-screening reports, hiring history — the four questions above are the discipline to apply before the first row of data gets used to train anything, and it's the same judgment PIA Studio is being built to support at scale.
