Toggle light / dark theme

AI MetaHacking: How 3 Microsoft Copilot Flaws Exposed A New Attack Method

Varonis Threat Labs found a critical Microsoft Copilot vulnerability, CVE-2026–24301, without reverse-engineering any code. Researchers simply asked Copilot to explain its own security limitations, and it did. Microsoft patched the flaw, nicknamed CoSnitch, on August 18. It’s the third Copilot vulnerability Varonis has disclosed this year, and the AI meta-hacking discovery method may matter more than the bug itself.

AI meta-hacking just proved an AI assistant can be talked into exposing its own attack surface. Varonis Threat Labs disclosed CoSnitch, a chain of three vulnerabilities in Microsoft Copilot Personal, discovered through what the firm calls meta-hacking: prompting Copilot to explain why certain actions shouldn’t be possible, then using its own answers to map the exact boundary of what actually was possible, according to The Register’s coverage of the disclosure. Microsoft assigned the flaw a CVSS severity score of 8.8 and shipped a fix on August 18, 2026.

Varonis researcher Håkon Måløy and colleagues didn’t start by probing code. They started by asking Copilot conversational questions about its own guardrails, and the assistant’s technical explanations revealed an undocumented URL parameter,?autorun=1, that could auto-execute a malicious prompt the moment a victim clicked a crafted link, according to The Hacker News. That single click gave an attacker’s prompt the ability to act inside the victim’s authenticated session, retrieving data from any connected app, Gmail, Google Drive, calendars, chat history, using Copilot’s own existing permissions. AI meta-hacking as a discovery method didn’t require finding a coding error at all; it required getting the AI to describe its own limits out loud.

Leave a Comment

Lifeboat Foundation respects your privacy! Your email address will not be published.

/* */