An LLM generated fake research themes, and 20 researchers didn't notice
Something unexpected happened in a recent workshop.
My workshops follow a certain process:
Everyone runs their own independent LLM analysis of the same research dataset, producing themes and citations.
Everyone checks and compares everyone else’s LLM-generated themes.
Lastly, everyone checks the cited quotes that support the themes, comparing them back to the original transcripts.
In this particular workshop, we went through steps 1 & 2 just fine–everybody in the workshop reviewed everyone else’s themes, and nobody saw anything wrong. They all looked like real, legitimate themes.
But in step 3, something strange happened for one of my participants: Even though the quotes looked real, they couldn’t find them in the original transcripts.
It took a lot of digging before they realized what had happened: The transcripts had never made it into the LLM at all! So the entire report–themes and all–was fabricated, based solely on the LLM’s “guesses” from the prompt.
But no one noticed until they manually checked the quotes!
Bear in mind that all of the reported themes had already passed the “gut checks” of 20 other researchers. It was only through extensive manual checking that anyone flagged something as wrong.
This should be a wake-up call for all of us using AI as a research tool.
We can’t trust that our familiarity with the data is enough to protect us from making mistakes.
LLMs are so good at producing plausible-looking output that completely fabricated themes can fly under our radar if we fail to conduct appropriate checks.
👉 Does your team need a process for evaluating LLMs and their output on research tasks? Join me at my next Rosenfeld Media workshop