New malware campaign tricks AI scanners with fake nuclear weapon prompts — malicious code triggers safety failsafes so scanners skip the payload
The Hades malware campaign weaponizes AI safety filters by injecting biological/nuclear weapon prompts to halt security scans before payload detection—a novel adversarial ML technique.
The Hades Malware Campaign: When AI Safety Guards Become Attack Vectors
A sophisticated malware campaign dubbed Hades has emerged with a novel and deeply concerning tactic: it deliberately injects text about biological and nuclear weapons into AI chatbot conversations to trigger safety filters that halt security scans before the malicious payload is ever analyzed. Discovered by researchers at ANY.RUN, this campaign represents a significant evolution in adversarial machine learning, turning the very guardrails designed to protect users into a shield for malware.
How the Attack Works
The Hades campaign distributes its payload through malicious email attachments—typically disguised as invoices, shipping notifications, or job applications. When opened, these documents execute a multi-stage infection chain that ultimately deploys an information stealer capable of harvesting browser credentials, cryptocurrency wallets, and system information.
What makes Hades unique is its AI evasion technique. During execution, the malware injects prompts referencing the creation of biological weapons, nuclear materials, and other extreme harm scenarios into any AI-powered security analysis tools that might be monitoring the system. These prompts trigger the "refusal" or "failsafe" mechanisms built into major AI models—causing them to stop analysis, refuse to engage, or truncate their output before the actual malicious behavior is documented.
In effect, the malware weaponizes AI safety alignment. By flooding the context window with prohibited content, it forces the AI to "look away" at the critical moment when the payload would otherwise be detected and flagged.
Technical Details from ANY.RUN Analysis
ANY.RUN's interactive sandbox observed the campaign in action across multiple samples. The infection chain typically begins with a .lnk shortcut file or malicious Office document that launches PowerShell commands. These commands download additional stages from compromised legitimate websites or cloud storage services, making network-based detection more difficult.
The final payload is a variant of the Vidar stealer, a well-known malware-as-a-service offering that targets:
- Browser autofill data, cookies, and saved passwords
- Cryptocurrency wallet files (Electrum, Exodus, Atomic, and others)
- Two-factor authentication databases (Authy, Google Authenticator exports)
- System hardware and software inventory
- FTP and VPN client credentials
Critically, the AI-triggering text is not merely present in the malware's code—it is actively injected into runtime logs and memory at the moment security tools would capture behavioral telemetry. This suggests the operators understand exactly how AI-driven sandboxes and endpoint detection and response (EDR) systems ingest and process telemetry data.
Implications for AI-Driven Security
The Hades campaign exposes a fundamental tension in modern cybersecurity: as defenders increasingly rely on LLMs to triage alerts, analyze sandbox output, and generate detection rules, attackers gain a new attack surface. The same prompt injection techniques that plague consumer chatbots can now be deployed against security infrastructure.
Several defensive implications follow:
- Sanitization pipelines are essential. Telemetry fed to AI analyzers must be stripped of adversarial prompts before analysis.
- Deterministic, rule-based detection cannot be fully replaced. Heuristic and signature-based layers remain critical as a backstop when AI analysis is subverted.
- Red-teaming AI security tools must include prompt injection scenarios modeled on real malware behavior, not just academic benchmarks.
- Human-in-the-loop review for high-severity alerts where AI analysis was truncated or refused.
Industry Response and Mitigation
ANY.RUN has updated its sandbox to detect and neutralize the Hades AI-evasion technique, and other sandbox vendors are reportedly deploying similar countermeasures. However, the cat-and-mouse dynamic is clear: as long as AI models have hard-coded refusal triggers tied to specific keywords or concepts, malware authors can weaponize those triggers by ensuring those keywords appear at strategically critical moments.
For organizations, the immediate recommendations are practical:
- Block execution of .lnk files and scripts from email attachments via Group Policy or endpoint controls.
- Deploy application control (AppLocker, WDAC) to restrict PowerShell and scripting engine usage to signed, approved scripts.
- Ensure EDR telemetry collection is not solely dependent on AI-based analysis for final verdicts.
- Conduct phishing simulations that include .lnk and HTML smuggling techniques.
A Harbinger of Things to Come
Hades is likely not an isolated innovation. As AI becomes more deeply embedded in security operations—from SOC analyst assistants to autonomous response systems—the incentive for adversaries to develop anti-AI techniques will only grow. We can expect to see:
- More sophisticated prompt injection payloads tailored to specific security vendors' AI prompts
- Training data poisoning targeting open-source models used in security tooling
- Model extraction attacks to reverse-engineer detection prompts
- Adversarial examples crafted to evade AI-based static analysis
The Hades campaign serves as a stark reminder that AI safety mechanisms, when deployed in adversarial environments without robust input validation, can be repurposed as offensive weapons. The cybersecurity industry must treat AI not just as a force multiplier for defense, but as an attack surface requiring its own threat modeling, red teaming, and hardening.
Source: Tom's Hardware — "Hades Malware Campaign Now Tricks AI Bots by Injecting Text About Biological and Nuclear Weapons"


