Researchers at TNG Technology Consulting demonstrated how hidden malicious behavior can be embedded into large language models through reinforcement learning, training a Qwen 27B model to exfiltrate secrets when it encountered specific triggers while behaving normally otherwise. The backdoor was trained using approximately one day of GPU compute. The study found that sandboxing and guardrailing provide only partial protection against such sleeper agents, and concluded that robust security measures for agentic tools remain essential regardless of whether a backdoor is actually present.