Wire flash
Tokyo Metropolitan University study: Pangram AI detector misses 79.8% of abstracts rewritten by Meta's Muse-Glimmer
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A new study from Tokyo Metropolitan University reveals that Pangram, an AI-text detection tool, failed to identify 79.8% of scientific abstracts that had been rewritten by Meta's Muse-Glimmer language model. The research found that Pangram's detection accuracy varied significantly depending on which AI model was used to rewrite the text. While the detector caught 93.5% of abstracts rewritten by GPT-5, it missed the vast majority of those rewritten by the newer Muse-Glimmer model. Notably, Pangram falsely flagged only 1 out of 5,000 human-written abstracts as AI-generated, indicating a very low false positive rate. The findings highlight the challenge of reliably detecting AI-generated text as language models continue to evolve and improve.
Source report
A new study from Tokyo Metropolitan University reveals significant limitations in the AI-text detection tool Pangram. The detector missed 79.8% of scientific abstracts rewritten by Meta's Muse-Glimmer model, while incorrectly flagging only 1 out of 5,000 human-written abstracts.
Key Findings
- Detection performance varies by model: Pangram's miss rate depends heavily on which large language model (LLM) performed the rewriting.
- GPT-5 rewrites caught effectively: The detector successfully identified 93.5% of abstracts rewritten by GPT-5.
- Muse-Glimmer rewrites largely missed: In contrast, 79.8% of abstracts rewritten by Meta's Muse-Glimmer went undetected.
The researchers conclude that Pangram's accuracy is not uniform across AI models, with its miss rate determined primarily by the specific LLM used for rewriting.
Source
rohanpaul_aiNeutral / independent