The paper introduces ToolHazard, a framework that automatically synthesizes adversarial environments to test the security of LLM-based agents that use tools. Built from three components — an Environment Simulator, an Attacker Agent, and a User Simulator — the framework generates environment-specific payloads to discover injection vulnerabilities at scale. Using the resulting ToolHazard-Bench, researchers found that attack timing significantly affects effectiveness, and showed that alignment data derived from the framework improves agent security while preserving performance on legitimate tasks.
