Police departments have been using algorithms far longer than the current AI conversation suggests — automated fingerprint matching dates to the 1970s, which surprises most people I mention it to. What’s new is scope: systems that recognise faces in crowds, forecast where crime will occur, and score individuals for risk. Some of this genuinely helps. Some of it fails in ways that worry me as someone who builds these systems for a living. The honest article is the one that says which is which.
Start with what the technology does well, because it’s real. Machines process volume humans can’t: matching evidence across databases in seconds, finding patterns across thousands of case files, mapping incidents to reveal hotspots and trends an analyst would take weeks to surface. Speed matters in investigations, and consistency matters everywhere — a model applies the same criteria at 4 p.m. and 4 a.m. No tired human can say that.
But each headline application carries a specific, documented failure mode, and they deserve to be named rather than waved at.
Facial recognition’s is unequal error. The systems deployed by US departments perform measurably worse on darker-skinned faces, and a false match in this context isn’t a product bug. It’s a wrongful arrest. The known wrongful-arrest cases in the US trace to exactly this failure, compounded by officers treating a match score as an identification.
Predictive policing’s is the feedback loop, and it’s subtle enough to fool honest people. The models forecast crime where crime was previously recorded. But recorded crime reflects where police patrolled. Send patrols where the model points, generate more records there, retrain: the system is now confirming its own bias and calling it accuracy. “The algorithm says so” launders yesterday’s patrol pattern into tomorrow’s.
Every individual step in that ring is defensible, which is exactly why it gets built by people acting in good faith. Nowhere in the loop is there a measurement of where crime happened.
Risk scoring’s is consequence without explanation. When a model influences bail or sentencing, the person affected deserves to know why. Yet the systems in actual use have often been proprietary black boxes that neither defendant nor judge could interrogate.
Underneath all three runs the privacy question: these tools work by watching everyone to find anyone, and mass collection is a cost even when nothing goes wrong.
So what would using this well look like? Not abstinence; the volume problems are real and the tools help. It looks like the boring governance the vendors never lead with: independent accuracy audits, by demographic, before deployment and continuously after. Explanations for any consequential decision, as a requirement rather than a courtesy. Human decision-makers who treat model output as one input, and are trained on exactly how it fails. And retiring systems that can’t meet that bar, however impressive the demo.
I see the same pattern in every domain we work in. But nowhere are the stakes of a false positive higher, and nowhere is “the computer said so” more corrosive. A tool this powerful in a domain this consequential earns its place through accountability, or it doesn’t belong there.