AI Tool 'Waldo' Detects Hidden Health Risks in Social Media Posts

Health experiences shared online are valuable safety signals invisible to traditional systems
Researchers argue that social media posts reveal real-world harms that FDA's voluntary reporting system never captures.
Mark

So Waldo is scanning Reddit posts and finding reports of harm from cannabis products. But how do we know those Reddit posts are actually describing real adverse events and not just people complaining or exaggerating?

Mimi

That's exactly why the researchers manually validated a sample. They took a random subset of the 28,832 posts Waldo flagged and had humans review them. Eighty-six percent checked out as genuine adverse events. So there's real harm being reported, not just noise.

Luke

But that also means 14% of what Waldo flagged was false positives. And we're only talking about a sample validation, not the entire dataset. So the actual number of true adverse events could be somewhat lower than 28,832.

Mimi

True, but even at 86%, that's still roughly 24,000 confirmed harm reports that the FDA's voluntary system would never have captured. These are people describing real health problems they experienced.

Mark

Why does the FDA's current system miss these? Isn't that its job?

Mimi

The FDA system depends on doctors and manufacturers voluntarily submitting reports. But most people who buy cannabis products or supplements don't go to a doctor about side effects—they just post about it online. And manufacturers have no incentive to report harms. So the system is blind to a huge population of users.

Luke

That said, we should be careful about one thing: Reddit users are not representative of all cannabis users. They're people who use Reddit, which skews younger, more tech-savvy, maybe more likely to discuss health issues online. The adverse events Waldo finds are real, but they might not reflect the full spectrum of harms across the entire user population.

Mimi

Fair point. But it's still a signal that traditional systems are missing entirely. Waldo isn't claiming to be perfect—it's claiming to be better than nothing, which it clearly is.

Mark

And the fact that it outperformed ChatGPT—does that surprise you?

Mimi

Not really. ChatGPT is a generalist. Waldo was trained specifically on adverse event detection. It's like comparing a general practitioner to a specialist.

Luke

Though we should note that the comparison was on a specific task with a specific dataset. ChatGPT might perform differently on other types of health data or other platforms. The study tested one tool against one chatbot on one type of product.

Mark

What happens now that it's open-source?

Mimi

Anyone can use it. Regulators could monitor social media for safety signals. Clinicians could track what their patients are experiencing. Researchers can apply it to other products that lack oversight.

Luke

The question is whether they actually will. Open-source doesn't guarantee adoption. And there are privacy and ethical questions about scraping social media data at scale that the paper doesn't deeply address.

  • Millions of people use cannabis products and dietary supplements with no mandatory safety reporting in place, meaning serious harms can go undetected for years while the market keeps growing.
  • Waldo, trained on the specific language of human suffering shared online, achieved 99.7% accuracy identifying adverse events — outperforming ChatGPT and validating 86% of nearly 29,000 flagged harm reports from Reddit.
  • The FDA's voluntary reporting system, built for prescription drugs and approved devices, was never equipped to catch safety signals from the unregulated consumer health landscape that now surrounds it.
  • By releasing Waldo as open-source software, the UC San Diego team is handing regulators, clinicians, and researchers a ready-made surveillance tool — no licensing fees, no gatekeepers, no waiting.
  • The project reframes social media not as noise but as a living record of real-world health experience, one that specialized AI can now read with clinical precision.

In the vast, largely unmonitored territory where consumer health products meet everyday human experience, a team at UC San Diego has built a quiet sentinel. Their AI tool, Waldo, listens to the informal confessions people make on social media about what products have done to their bodies — filling a silence that formal regulatory systems were never designed to hear. At a moment when cannabis derivatives and dietary supplements reach millions without mandatory safety reporting, this act of attentive listening may represent a meaningful shift in how societies learn from harm.

Researchers at UC San Diego have built an AI system called Waldo that scans social media to detect health harms people report about consumer products — harms that traditional safety monitoring has largely failed to capture. Tested on Reddit discussions about cannabis-derived products, Waldo achieved 99.7% accuracy compared to human reviewers. Across more than 437,000 posts, it flagged nearly 29,000 potential harm reports, and manual review confirmed 86% were genuine adverse events.

The tool addresses a real structural gap. The FDA's adverse event reporting system depends on voluntary submissions from doctors and manufacturers — a process designed for prescription drugs and approved medical devices. But cannabis derivatives, dietary supplements, and similar products now reach millions of consumers with no mandatory safety reporting attached. Serious harms can persist undetected for months or years. Waldo was built to treat what people share online as legitimate real-world evidence about what products do to human bodies.

What makes Waldo technically notable is its specialization. Built on a machine learning model called RoBERTa and trained specifically to recognize the language people use when describing health problems, it substantially outperformed ChatGPT on the same task — suggesting that focused training beats general sophistication. The researchers released Waldo as open-source software, making it freely available to regulators, clinicians, and public health officials worldwide.

Lead author Karan Desai framed the work as a way to surface voices that traditional reporting never hears. Senior researcher John Ayers described it as a demonstration of how digital tools can reshape post-market surveillance. Though the immediate application was cannabis products, the team sees broader reach: any consumer health category that escapes regulatory oversight could potentially be monitored this way. The open-source release was a deliberate choice — an invitation to collaborative, transparent science that widens the safety net for people using products that fall outside conventional regulatory channels.

Researchers at UC San Diego have developed an artificial intelligence system capable of sifting through social media posts to identify health harms that people report about consumer products—a gap that traditional safety monitoring has largely missed. The tool, called Waldo, was tested on Reddit discussions about cannabis-derived products and achieved 99.7% accuracy when compared against human reviewers who manually identified adverse events in the same posts. When the team ran Waldo across a dataset of 437,132 Reddit posts about cannabis products, it flagged nearly 29,000 potential reports of harm. When researchers manually checked a random sample of these flagged posts, they confirmed that 86% represented genuine adverse events.

The motivation behind Waldo reflects a real gap in how the United States monitors product safety. The FDA's current system for tracking adverse events relies on voluntary reports submitted by doctors and manufacturers—a process that works reasonably well for prescription medications and medical devices that have gone through formal approval. But the market for consumer health products has expanded rapidly in recent years: cannabis-derived products, dietary supplements, and other items sold without FDA oversight now reach millions of people. These products generate no mandatory safety reporting, which means serious harms can persist undetected for months or years. Waldo was designed to capture the safety signals that people naturally share online, treating social media not as noise but as a legitimate source of real-world evidence about what products do to people's bodies.

The technical achievement matters. Waldo is built on a machine learning model called RoBERTa that was carefully trained to recognize the specific language people use when describing health problems. When researchers compared it to ChatGPT—a general-purpose AI chatbot—Waldo substantially outperformed the larger, more famous model. This suggests that specialized training on a specific task produces better results than relying on a general system, no matter how sophisticated. The researchers have released Waldo as open-source software, meaning regulators, clinicians, researchers, and public health officials can download it and use it without paying licensing fees or waiting for permission.

Karan Desai, the study's lead author, framed the finding this way: the health experiences people share online contain valuable information about safety. By capturing these voices systematically, researchers can surface real-world harms that traditional reporting systems never see. John Ayers, a senior researcher on the project, emphasized that the work demonstrates how digital tools can reshape post-market surveillance—the ongoing monitoring of products after they reach consumers. Vijay Tiyyala, another team member, noted that the accuracy of the model was encouraging, suggesting that careful training can produce tools that outperform state-of-the-art alternatives.

The immediate application was cannabis products, but the researchers suggest Waldo's approach has broader reach. Any consumer health product that lacks regulatory oversight could potentially be monitored this way. The team's decision to make the tool open-source reflects a deliberate choice to democratize access to safety monitoring technology. Rather than keeping Waldo proprietary or limiting it to a single institution, they've made it available to anyone who wants to use it. This move is intended to accelerate what researchers call open science—collaborative, transparent research that spreads knowledge and tools widely—while simultaneously improving the safety net for patients who use products that fall outside traditional regulatory channels.

The health experiences people share online are not just noise, they're valuable safety signals. By capturing these voices, we can surface real-world harms that are invisible to traditional reporting systems.
— Karan Desai, lead author
This project highlights how digital health tools can transform post-market surveillance. By making Waldo open-source, we're ensuring that anyone, from regulators to clinicians, can use it to protect patients.
— John Ayers, UC San Diego
Contattaci Domande frequenti