Security researchers reveal how Copilot was manipulated into exposing self-hacking details

Cybersecurity researchers at Varonis Threat Labs discovered a now-fixed vulnerability dubbed 'CoSnitch' in Microsoft's Copilot, which allowed them to manipulate the AI into revealing internal architecture details and an undocumented URL parameter.

Security researchers at Varonis Threat Labs have identified a now-fixed vulnerability in Microsoft's Copilot AI tool, which they dubbed "CoSnitch." According to the research team, the system wasn't technically breached through traditional code exploits; instead, it was manipulated through continuous questioning into revealing sensitive details about its own architecture and how to hack itself.

The cybersecurity team initiated the process by asking Copilot how to execute an automatic prompt without user interaction. Initially, the AI responded that user intent was required and that prompts could not be enacted independently. Rather than accepting the refusal, the researchers continued responding with follow-up questions. Each time Copilot offered a technical justification for its refusal, the team used the explanation to map the system's internal architecture, treating the AI's resistance as an invitation to probe further.

This technique, referred to as meta-hacking, eventually caused Copilot to reveal an undocumented URL parameter known as "autorun=1," along with information regarding the protections put in place to disable it. After testing the parameter, the researchers created a malicious URL designed to force Copilot to load into an authenticated session via a browser, trigger an auto-prompt execution, and process the results without explicit user action. The team noted that this method could potentially exfiltrate data gained from connected applications like Gmail, OneDrive, and Calendar using Copilot's built-in URL-fetch capability.

Varonis disclosed the security issue to Microsoft in December of the prior year, and the vulnerability was successfully patched out on August 18. Although the specific flaw has been addressed, the research team warns that this meta-hacking technique can potentially be applied to any agentic AI platform featuring a natural language interface, and further research on the topic is expected.

Further reading

Sources

  1. PC GamerEstablished publication · recorded Aug 19, 2026
    Copilot was bamboozled into revealing how to hack itself, security researchers claim: 'Copilot wasn’t breached; it was played'