Georgetown study: AI video summaries distort human memory, even when users know source

Even when told the summary was AI, people's memories were still distorted.
Participants exposed to misleading AI summaries had their recall corrupted regardless of whether they knew the source was artificial.
Mark

So the study shows that bad AI summaries change what people remember. But how bad are we talking? Are these summaries obviously wrong, or are they subtly misleading?

Mimi

Both, actually. The AI systems omitted more than half of the central details on average. In the car-pedestrian scenario, 95 percent of summaries left out the collision itself—the main event. That's not subtle. But the experiment also tested whether people could catch the error if they knew it came from AI. They couldn't. Even when told the summary was machine-generated, people's memories were still distorted.

Luke

Wait—the study used animated videos, not real footage. Does that matter? Are people's memories more malleable with simplified animations than with actual video?

Mimi

That's a fair question. The researchers chose animations deliberately, based on classic memory studies, to control variables. But you're right that it's a limitation. They want to move to real-world footage next—actual police body camera video, actual medical records. That's where the real risk lives.

Mark

Why does it matter whether people know it's AI-generated? Shouldn't that make them more skeptical?

Mimi

You'd think so. But the study found no difference. People who were told "this is AI" had their memories distorted just as much as people who thought a human wrote it. The source label didn't protect them.

Luke

That's interesting, but I want to be careful here. The study tested whether people *believed* the summary was AI or human. Did they actually test whether *knowing* something is AI makes people more skeptical in general? Or did they just tell people it was AI and then measure memory?

Mimi

They told people the source before showing the summary. But you're right—that's different from testing whether people who are generally skeptical of AI perform differently. The study doesn't isolate that.

Mark

What about the AI systems themselves? Were they using the latest, most powerful models?

Mimi

They used ChatGPT and Google's Gemini—the consumer-facing versions available at the time. But Eiger, one of the researchers, pointed out that police departments and hospitals might be using enterprise versions designed for organizational use, which the team didn't have access to. So there's a gap between what was tested and what's actually being deployed.

Luke

That's a crucial gap. We don't know if enterprise AI is better or worse at summarization. And we don't know what happens when these systems are fine-tuned for specific domains. The study is testing general-purpose AI on a controlled task. Real-world deployment is messier.

Mark

So what's the takeaway? Should we stop using AI to summarize videos?

Mimi

The researchers aren't saying ban it. They're saying we need to think deeply about whether and how to use it in high-stakes settings. The concern is that humans-in-the-loop—the idea that people will catch AI errors—doesn't work if the errors corrupt people's memories before they can catch them.

Luke

But the study doesn't test what happens if you show people both the original video and the summary. It tests memory 24 to 48 hours later, after the original video is gone. In a real police investigation, would officers have access to the original footage alongside the summary?

Mimi

That's another open question. The study is designed to isolate the effect of the summary on memory. But you're right—real-world workflows might include safeguards the study doesn't model.

  • AI video summarization tools are already in use across law enforcement and medicine, yet a new study reveals they omit more than half of central details on average — and missed the collision entirely in 95% of tested summaries.
  • Exposure to a single misleading AI summary was enough to nearly halve participants' ability to correctly recall what they had just watched, collapsing accuracy from 84% to 45%.
  • Perhaps most alarming: knowing the summary was AI-generated offered no protection — human memory was reshaped by the errors regardless of the source label.
  • Researchers warn that the assumption of human oversight — that people will catch AI mistakes — may be fundamentally broken, because the act of reading a flawed summary corrupts the memory it was meant to describe.
  • The study's authors are calling for rigorous testing before AI summarization tools are deployed in high-stakes settings, and plan to extend their research into actual police body camera footage and medical records.

In the long human struggle to hold accurate witness to events, a new distorting force has emerged: artificial intelligence systems that summarize video footage not only miss what matters, but quietly rewrite what observers believe they saw. Researchers at Georgetown University and the University of Washington have demonstrated that AI-generated summaries containing errors can corrupt human memory of the original event — reducing correct recall from 84 percent to 45 percent — and that this distortion persists even when people know the summary was machine-made. The finding arrives at a moment when such tools are already being deployed in courtrooms, hospitals, and police precincts, where the distance between what happened and what is remembered can determine a life's course.

When an AI system summarizes a video, its errors don't simply sit alongside human understanding — they burrow into it. That is the central and unsettling finding of a study led by Mattea Sim of Georgetown's McCourt School of Public Policy, conducted with colleagues at the University of Washington.

Participants watched animated videos of a car approaching a traffic sign before colliding with a pedestrian, then read either accurate or misleading summaries generated by ChatGPT or Google's Gemini. The results were unambiguous: those who read accurate summaries recalled the correct sign 84% of the time; those who read misleading ones fell to 45%. Across 331 participants and two sessions spaced a day or two apart, the pattern held firm.

Before testing memory, the researchers catalogued the AI systems' errors. On average, summaries omitted 51.6% of central details. More striking still: 95% of all summaries failed to mention the collision — the very event the video depicted. Beyond omission, the systems also fabricated details that never appeared in the footage at all.

What made the findings particularly difficult to dismiss was their indifference to awareness. Even participants told explicitly that they were reading AI-generated text showed the same memory distortion. The common assumption — that human oversight will catch and correct AI errors — appears to break down at the point of reading: by then, the memory has already been altered.

Ph.D. candidate Yael Eiger, whose research focuses on technology in criminal justice, said she was struck by how poor the summaries were even at this stage of AI development, and voiced concern that police departments may be adopting these tools without understanding their risks. The implications extend to prosecutors, judges, and juries who may encounter AI-processed body camera footage as evidence.

Sim and her colleagues, including computer scientist Yoshi Kohno, are calling for serious scrutiny of AI use in critical documentation contexts. The study will be presented at the AAAI/ACM Conference on AI, Ethics and Society in October 2026, with follow-on research planned using real police footage and medical records — the domains where the cost of a corrupted memory is highest.

Researchers at Georgetown University and the University of Washington have documented something unsettling: when artificial intelligence systems summarize video footage, the errors they introduce don't just sit passively in the background of human understanding. They actively reshape what people remember about the events they watched.

The study, led by Mattea Sim, an assistant research professor at Georgetown's McCourt School of Public Policy, tested how AI-generated summaries affect memory by having participants watch animated videos of a car approaching an intersection with either a stop or yield sign, then colliding with a pedestrian. The researchers then showed some participants accurate summaries of what happened and others misleading ones—all generated by ChatGPT or Google's Gemini—and measured how well people could recall the original event. The results were stark: 83.6 percent of participants who read accurate summaries answered correctly when asked what sign the car had approached, while only 44.8 percent of those exposed to misleading summaries got it right. The gap between truth and distortion was nearly 40 percentage points.

What made this finding particularly troubling was that it held true regardless of whether participants knew they were reading AI-generated text. Even when told explicitly that the summary came from an artificial system, people's memories were still warped by the inaccurate information. The researchers tested this with 331 participants across two sessions separated by 24 to 48 hours, controlling for variables and ensuring that all summary text came from the same AI sources to maintain experimental consistency.

Before measuring memory effects, the team analyzed what kinds of errors the AI systems actually produced. They found that summaries omitted 51.6 percent of central details on average. More strikingly, 95 percent of all summaries failed to mention the most important element of the video: the collision itself. The AI systems didn't just miss minor points. They systematically left out the core event. Beyond omission, the summaries also misconstrued information and generated details that were entirely fabricated—what researchers call hallucination in AI terminology.

Yael Eiger, a Ph.D. candidate at the University of Washington whose research focuses on technology in the criminal justice system, expressed particular concern about the implications for policing. "I was struck by how bad the summaries were, even at this stage in AI development," Eiger said. "It worries me that police departments may be using video summarization technologies without rigorous testing and without an awareness of how incorrect AI-generated summaries could be." The concern is not hypothetical. AI video summarization tools are already being deployed in real-world settings—from news organizations creating article previews to workplaces transcribing meetings to law enforcement agencies processing body camera footage.

The researchers deliberately did not use actual police footage for this study, instead relying on controlled animated videos based on materials from classic human memory research. But the implications point directly toward those high-stakes applications. If an AI system summarizes body camera footage and omits or misrepresents a critical detail, that summary could shape how officers, prosecutors, judges, and juries understand what actually occurred. The study suggests that human oversight—the common assumption that people will catch and correct AI errors—may not work as intended. When people read a flawed summary, their memory of the original event gets corrupted. They are less likely to notice the error because their recollection has already been altered.

Sim and her colleagues, including Yoshi Kohno, the McDevitt Chair in Computer Science, Ethics, and Society at Georgetown, are calling for deeper scrutiny of whether and how AI should be used to summarize information in critical contexts. "AI is a new method of delivering misinformation, and it has the potential to create these false memories for people who are reading that information," Sim said. The study will be presented at the Ninth AAAI/ACM Conference on AI, Ethics and Society in October 2026. The researchers plan to extend this work by examining real-world examples—actual police body camera footage, actual medical records—to understand whether the memory distortion effect holds in the domains where the stakes are highest.

AI is a new method of delivering misinformation, and it has the potential to create these false memories for people who are reading that information.
— Mattea Sim, Georgetown University
It worries me that police departments may be using video summarization technologies without rigorous testing and without an awareness of how incorrect AI-generated summaries could be.
— Yael Eiger, University of Washington
Quer a matéria completa? Leia o original em News-Medical ↗
Fale Conosco FAQ