A family’s account of a mistaken police detention in Moscow includes a technical claim that deserves careful interpretation: officers allegedly said an artificial-intelligence system found an 85–90% match between a 15-year-old schoolboy and a person in a police alert.

The family’s account of that score has not been independently verified, and the system has not been identified. The underlying detention, however, was reported in May. Novaya Gazeta Europe said the boy was stopped on May 4 on Garibaldi Street while walking to school. It cited his mother as saying plainclothes officers knocked him down, bound him and took him to a police unit, where it became clear that he had been mistaken for someone else.

The later Ostorozhno Media report says the father was told the teenager had been suspected of drug distribution. According to him, police eventually told the mother that an AI system had produced the 85–90% figure.

In biometric systems, a similarity score is not a universal probability. Facial-recognition algorithms convert images into mathematical representations and compare them. In a one-to-many identification search, one probe image is searched against a gallery containing many identities. The system ranks candidates or returns candidates above a threshold.

NIST’s Face Recognition Technology Evaluation measures two key classes of errors in this setting. A false positive identification occurs when a search that should not return a person nevertheless produces one or more candidates above the threshold. A false negative identification occurs when the correct person is in the gallery but fails to be returned above the threshold. The threshold itself is selected by the operator or system designer to balance risks for a particular use case.

This is why a reported value such as 85% or 90% cannot be interpreted in isolation. One vendor might express similarity on one numerical scale and another on a different scale. The operational question is whether the score exceeds a threshold calibrated to a known false-positive rate, under conditions comparable to the actual camera imagery.

Moscow has deployed face recognition in public transport. The city’s official transport portal says the Sfera system in the Metro generates biometric keys, checks them against wanted-person databases and alerts police. But no evidence ties Sfera specifically to the Garibaldi Street detention.

Image quality is another variable. Lighting, camera angle, resolution, motion and partial occlusion can alter performance. In real-world one-to-many searches, the size and composition of the gallery also matter because the system is comparing the probe against many possible identities.

The human consequences in the Moscow case are the reason these technical distinctions matter. Novaya Gazeta Europe reported that the teenager was diagnosed with a concussion after release. Ostorozhno Media later cited medical records describing additional injuries. The family says officers used excessive force.

A police response quoted by the family says the teenager resisted, injured officers and tried to flee, leading to the use of physical force. The family disputes that account. Complaints were filed with police and the Investigative Committee, and the review remains unresolved from the family’s perspective.

If investigators confirm that facial recognition initiated the stop, a scientifically meaningful review would need more than the headline percentage. It would require the system’s identity, its threshold, the source image, the candidate list, the false-positive performance at that threshold and the human-review procedure used before officers acted. Without those elements, “85–90%” is a number without enough context to establish identity.

Jordan Quincy

Author

Technology Reporter

Jordan Quincy covers public affairs, politics, business, culture and daily news for Science Official. The role focuses on verification, context, and clear explanations for readers.