Scam2Prompt Framework Reveals Severe Security Vulnerabilities in Production LLMs
A new research paper introduces Scam2Prompt, a scalable automated auditing framework designed to evaluate security risks in Large Language Models (LLMs). The study highlights the danger of LLMs absorbing malicious content from uncurated web datasets, citing a November 2024 incident where a user lost $2,500 due to phishing code generated by ChatGPT. Scam2Prompt identifies scam intents and synthesizes developer-style prompts to test if LLMs generate malicious code. Initial tests on four models, including GPT-4o and Llama-4-Scout, showed a 4.24% failure rate. A benchmark called Innoc2Scam-bench was created with 1,377 prompts that consistently triggered malicious outputs. Further testing on seven additional LLMs released in 2025 revealed severe vulnerabilities, with malicious code generation rates ranging from 12.9% to 47.3%. The researchers concluded that existing safety measures, such as state-of-the-art guardrails and RAG-based agents, are insufficient to prevent these security breaches, indicating a critical need for improved auditing and safety protocols in AI development.
Wire timeline
Scam2Prompt Framework Reveals Severe Security Vulnerabilities in Production LLMs
A new research paper introduces Scam2Prompt, a scalable automated auditing framework designed to evaluate security risks in Large Language Models (LLMs). The study highlights the danger of LLMs absorbing malicious content from uncurated web datasets, citing a November 2024 incident where a user lost $2,500 due to phishing code generated by ChatGPT. Scam2Prompt identifies scam intents and synthesizes developer-style prompts to test if LLMs generate malicious code. Initial tests on four models, including GPT-4o and Llama-4-Scout, showed a 4.24% failure rate. A benchmark called Innoc2Scam-bench was created with 1,377 prompts that consistently triggered malicious outputs. Further testing on seven additional LLMs released in 2025 revealed severe vulnerabilities, with malicious code generation rates ranging from 12.9% to 47.3%. The researchers concluded that existing safety measures, such as state-of-the-art guardrails and RAG-based agents, are insufficient to prevent these security breaches, indicating a critical need for improved auditing and safety protocols in AI development.
cs.AI updates on arXiv.org