Images from police body camera footage with red and blue squares over important parts of the image.

AI tools can miss the mark on police bodycam footage

AI is increasingly being used to help analyze police bodycam recordings, but new research from Princeton Engineering shows that in 1 out of 4 cases the models can’t identify basic details, like whether an officer handcuffed someone or drew a gun. 

AI models don’t perform this task reliably because they are trained on video that is staged, well-lit and clear, said Olga Russakovsky, an associate professor of computer science and lead researcher on the paper, which is being presented at the European Conference on Computer Vision in September. Actual body camera footage is grainy, blurry and fast-moving.

Accurate AI review would be very useful for police departments, since most lack the resources to analyze the petabytes of footage collected by officers every year. But while current models are rapidly improving at tasks like coding and transcription, they have not substantially improved at analyzing real-world body camera video.

“One goal of the paper is to provide a way for researchers and police departments to evaluate off-the-shelf tools for processing bodycam recordings,” said Brandon Stewart, professor of sociology at Princeton and a senior author on the paper. “The general-purpose tools we tried like ChatGPT and Gemini don’t work well enough to be deployed in agencies today.”

Police departments are already using specialized AI tools for a variety of tasks, according to the researchers, including detection of facial expressions, body movement and use of force. The researchers did not evaluate any of these commercial products.

Olga Russakovsky and Brandon Stewart
Olga Russakovsky and Brandon Stewart. Photo of Russakovsky by Sameer Khan/Fotobuddy, photo of Stewart by Melissa Kelly

Instead, they created a computer-vision benchmark to test the accuracy of current AI models. The benchmark consists of a dataset with 185 hours of public footage from police departments in Illinois, California, Texas and the District of Columbia. Thirty-three students annotated 1-second, 10-second and one-minute clips from this footage, noting whether a police officer ran, handcuffed someone, drew their weapon or provided medical attention.

The researchers tested 12 AI models using this benchmark, including Gemini 2.5 Flash, GPT 4.1, Llama-VID and Qwen 2.5VL. The most accurate models correctly identified what was happening in a one-minute video about 77% of the time. The least accurate were correct 11% of the time.

For a typical computer vision task, 77% accuracy would be a decent result, said Jihoon Chung, a doctoral student in computer science and co-first author on paper. For identifying dog breeds, for example, perfect accuracy isn’t expected. “But for this specific high-stakes task, even if we get 95% accuracy, that’s still not good enough to be deployed in the real world,” he said.

The problem is not a 5% failure rate, but establishing trust, said Max Gonzalez Saez-Diez, co-first author on the paper and a former master’s student in computer science. “If this system fails, even 1 in 20 cases, then people will always question its accuracy and question whether it’s actually helping.”

Body-worn cameras were first introduced about fifteen years ago to provide greater oversight and transparency in policing, according to the Urban Institute. But there is simply too much footage collected for any department to review. Rules about storing footage vary across states and jurisdictions but according to the Police Executive Research Forum most departments discard recordings after 90 days unless it’s involved in an investigation.  

“No one’s going to be crawling through all of this footage unless really extreme events happen,” said Stewart. “For those extreme events it does matter that there’s a record of what happened, but for lower-level incidents it just kind of falls away because there’s no way to work with the data.”

AI tools could give this enormous collection of data a new purpose, helping humans to filter and sort through thousands of hours of video and detect patterns. “AI can reduce by a hundredfold and more the amount of time needed to review the footage,” said Russakovsky.

The ideal reviewing system for body camera footage would always have a human in the loop, she added, and would be able to pair video data with other sources, like police reports. The team is currently working on research that outlines how a system like this might function in a real-world setting.

The goal is to create an AI tool that can aid human judgment, not replace it. “A model might tell you there’s a fight going on, or there’s a gun, or the civilian is running away,” said Gonzalez Saez-Diez. “But figuring out who’s to blame? I think that’s a human problem, not a technical one.”


The paper, EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage, will be presented at the European Conference on Computer Vision in Malmo, Sweden on September 10. In addition to Chung, Gonzalez Saez-Diez, Russakovsky and Stewart, authors include Adam D. Wolsky and Jonathan Mummolo from Princeton University and Gregory Lanzalotto and Dean Knox at the University of Pennsylvania.

Related Faculty

Olga Russakovsky

Related Departments

Computer Science

Computer Science

Leading the field through foundational theory, applications, and societal impact