As artificial intelligence grows capable of fabricating voices, images, and narratives at scale, researchers are turning the same technology inward — training it to recognize the patterns of its own deceptions. From university labs in North America and Europe, a new generation of AI tools is being developed not to replace human judgment, but to extend it, flagging suspicious claims before they harden into belief. The promise is real, the limitations are honest, and the lesson ancient: no tool is wiser than the hand that wields it.
AI Tools Show Promise in Fighting Misinformation, But Human Oversight Remains Essential
We cannot just rely only on the AI. We need to observe and correct.
Why would we trust AI to catch misinformation when AI is what's creating it in the first place?
Because the same language understanding that lets AI generate convincing lies also lets it recognize patterns of falsehood. It's not about trusting the machine—it's about using it as a filter so humans can focus their limited time on what matters most.
But the accuracy numbers seem shaky. Fifty-five percent for Grok? That's barely better than a coin flip.
True, but that's not how it's meant to work. The tool isn't supposed to make the final call. It flags something as suspicious, and then a human journalist or fact-checker investigates. It's triage, not judgment.
What about the study where ChatGPT convinced people to abandon conspiracy theories? That seems almost too good to be true.
It does, and the data had some issues that the authors corrected. But the core finding held: when an AI system has infinite patience and uses evidence-based reasoning, it can actually persuade people. Humans get tired, frustrated, impatient. Machines don't.
So the real value isn't detecting lies—it's understanding the shape of misinformation as it spreads?
Exactly. One person's false claim about voter fraud is noise. Ten thousand posts echoing the same narrative, with variations and mutations—that's a pattern you can map, understand, and counter at scale.
And yet you're saying we still need humans watching over all of this.
Always. The moment we treat these systems as autonomous truth-tellers, we've already lost. They're tools, not oracles.
Il Polso
- AI-generated robocalls impersonating Joe Biden suppressed voter turnout in New Hampshire in 2024, illustrating how synthetic media has already crossed from nuisance into democratic threat.
- Detection models trained on human-verified claims achieve between 55 and 90 percent accuracy — a wide and telling range that exposes how uneven and context-dependent the technology remains.
- Large language models hallucinate confidently, inherit human biases, and struggle with ambiguity, meaning the very systems built to fight misinformation can themselves become sources of it.
- Researchers are engineering safeguards — models that admit uncertainty, bots that withhold answers when evidence is thin, and tools that analyze tone and rhetorical manipulation rather than content alone.
- A landmark 2024 study found AI chatbots reduced belief in conspiracy theories by 20 percent on average, outperforming traditional psychological interventions through sheer patience and persistence.
- Experts and platforms alike are converging on a hybrid model: AI as a triage and flagging system, with human journalists and fact-checkers retaining final authority over what is true.
As artificial intelligence grows capable of fabricating voices, images, and narratives at scale, researchers are turning the same technology inward — training it to recognize the patterns of its own deceptions. From university labs in North America and Europe, a new generation of AI tools is being developed not to replace human judgment, but to extend it, flagging suspicious claims before they harden into belief. The promise is real, the limitations are honest, and the lesson ancient: no tool is wiser than the hand that wields it.
The same artificial intelligence that generated a fake Joe Biden voice urging New Hampshire voters to stay home in 2024 is now being recruited to fight the misinformation it helped unleash. Researchers across North America and Europe are developing AI systems capable of identifying false claims, mapping how they spread, and even persuading people to abandon them — a counterintuitive but increasingly serious line of defense.
Early machine learning models trained on human-verified claims could spot COVID-19 misinformation with roughly 90 percent accuracy, but they were fragile — useful only within the narrow datasets and time periods they were built on. The shift toward large language models, the same architecture behind ChatGPT, brought far greater flexibility. These systems understand context, summarize narratives, and can engage users in sustained, evidence-based arguments. In one striking 2024 study published in Science, a version of ChatGPT reduced belief in conspiracy theories — including the faked Moon landing — by an average of 20 percent, outperforming conventional psychological interventions.
But the same depth of language understanding that makes these models useful also makes them unreliable. They hallucinate — generating confident, fluent falsehoods when faced with ambiguous or outdated information. An early 2025 version of the chatbot Grok agreed with human fact-checkers only 55 percent of the time, below even the 64 percent rate at which human fact-checkers agree with each other. Researchers are responding with targeted fixes: models trained to acknowledge uncertainty, bots that tell users when evidence is insufficient, and European tools that scan for 42 rhetorical signatures of deliberate disinformation rather than evaluating claims directly.
Beyond detection, AI is being used to chart the architecture of misinformation itself. Researchers like Jevin West at the University of Washington track how false narratives emerge, mutate, and proliferate across social media — work that language models handle well, quickly characterizing the overarching story beneath thousands of individual posts so fact-checkers can prioritize their responses.
Experts are unanimous on one point: none of this works without human oversight. AI tools inherit the biases of their training data, and blind trust in their outputs risks compounding the very problems they are meant to solve. The emerging consensus treats AI as a triage system — capable of flagging suspicious content at scale — while leaving final judgment to journalists and professional fact-checkers. YouTube removed over eleven thousand videos for misinformation violations in late 2025 using exactly this hybrid approach. Whether other major platforms are doing the same remains unclear, though mounting legal liability may soon force the question.
The same technology that floods social media with fake animal videos and AI-synthesized voices impersonating politicians is now being enlisted to fight the very misinformation it helped create. In 2024, thousands of New Hampshire voters received robocalls featuring an artificial voice of Joe Biden urging them to skip the primary election. That same year, AI-generated content—from fabricated images to doctored videos depicting violence in the Middle East—proliferated across platforms designed to harvest clicks and advertising revenue. Yet researchers at universities across North America and Europe are betting that artificial intelligence, properly constrained and supervised, could become a powerful tool for identifying and understanding false claims before they spread.
The logic is counterintuitive but compelling. Machine learning systems have long been trained to spot falsehoods by learning patterns in human-verified claims—the overuse of capital letters, exclamation points, emotionally charged language. When researchers at American University trained such a model to identify COVID-19 misinformation during the pandemic, it agreed with human fact-checkers roughly 90 percent of the time. But these older systems were brittle, trained on narrow datasets from specific time periods and platforms, leaving them useless in the messy real world. Researchers have shifted toward large language models—the same systems powering ChatGPT—which absorb vast amounts of internet text and learn the relationships between words, concepts, and contexts. These models understand human language deeply enough to analyze claims, summarize narratives, and even persuade people to abandon false beliefs.
Yet the same capability that makes them useful also makes them dangerous. Large language models are fundamentally language-imitation machines, not lie detectors. When given ambiguous or outdated information, they confidently generate false answers in a process researchers call hallucination. An early 2025 version of Grok, a chatbot designed to search the web for current information before answering, agreed with human fact-checkers only 55 percent of the time when asked to verify claims—worse than human fact-checkers agreeing with each other, which happens 64 percent of the time. The problem deepens when claims are genuinely ambiguous. The statement "Mark Carney is prime minister" is true only in Canada; a model given insufficient context might flounder or misinterpret contradictory evidence.
Researchers are developing workarounds. Dorsaf Sallami at McGill University trained her model to recognize ambiguity and ask users for clarification rather than guessing. The Dubawa fact-checking bot, launched in 2024 by a Nigerian nonprofit and accessible via WhatsApp, explicitly tells users when there is insufficient evidence for a claim instead of attempting an answer. A European collaboration called AI4Trust developed tools that analyze not what is said but how it is said—scanning for 42 common characteristics of deliberate disinformation, from allusions to secret conspiracies to emotionally manipulative language. When compared with human fact-checkers, that system agreed 70 percent of the time, enough to flag suspicious claims for journalists to investigate further.
Beyond detection, researchers are using large language models to map how misinformation spreads. Jevin West at the University of Washington and his colleagues track clusters of social media posts that propagate false narratives, watching how stories emerge, proliferate, and mutate over time. After Joe Biden's 2020 election victory, the "stop the steal" conspiracy theory spawned thousands of posts mixing outright fabrications with real but misleading information—videos of poll watchers being denied entry, for instance, without context that they were later admitted. Summarizing the larger narrative beneath such noise is challenging, but West finds that language models excel at this labeling task. When crisis managers and fact-checkers lack resources to address each individual claim, an AI system can quickly characterize the big-picture story so it can be evaluated and, if necessary, debunked.
Perhaps most striking, researchers have found that AI can actually change minds. In a 2024 study published in Science, nearly 2,200 Americans who believed conspiracy theories—that the Moon landing was faked, for instance—chatted with a version of ChatGPT instructed to persuade them otherwise. The chatbot reduced belief in the conspiracy theory by an average of 20 percent, a higher success rate than psychological interventions targeting the underlying psychology of conspiracy adherence. Thomas Costello of Carnegie Mellon University, who co-led the study, attributes this to the infinite patience of language models: they can spend whatever time and effort is necessary to conduct fact-based arguments, something human fact-checkers cannot do at scale.
Yet experts are adamant that these tools should never operate without human oversight. "We cannot just rely only on the AI," says Thanh Thi Nguyen of the University of the Sunshine Coast, comparing the challenge to raising a child—you want autonomy, but you must observe behavior and correct course. Large language models inherit the biases of the human-compiled data they were trained on, making blind trust in their outputs a recipe for the same misinformation problems they're meant to solve. Most researchers see AI primarily as a triage system: flagging content that journalists, fact-checkers, and social media platforms can then investigate thoroughly. YouTube reported removing 11,337 videos for misinformation violations in the last quarter of 2025, using a combination of automated detection and human review. Whether other platforms are deploying such tools remains unclear—TikTok, Meta, and X declined to comment—though recent court cases holding Meta and Google liable for harm to young users may prompt companies to intensify their efforts. The technology shows genuine promise, but only when humans remain in control.
Citazioni salienti
We should fight fire with fire.— Jevin West, University of Washington misinformation researcher
They're actually incredibly good at using reason and evidence to talk someone out of a particular belief.— Thomas Costello, Carnegie Mellon University, on language models' persuasive capacity