Researchers introduce EvoSafeHarness, a framework that automatically generates customized safety policies for LLM agents by jointly optimizing a natural-language policy alongside executable code logic tailored to a specific model and application domain. The approach targets a known weakness of fixed, expert-designed defenses, which often fail to balance safety and utility across different deployment contexts. On benchmarks including DecodingTrust-Agent and AgentDojo, EvoSafeHarness substantially reduces attack success rates while preserving higher task utility than existing defenses, showing that effective safety constraints depend on both domain semantics and model-specific behavior.
