A new paper introduces Second Thought, a training-free inference framework that has LLM agents spawn auxiliary reasoning branches during the idle time between taking an action and receiving an environment observation. Evaluated across three agentic benchmarks and multiple reasoning models, the technique reduced the number of turns needed across all nine tested configurations and cut main-thread decoding by up to 43% in some cases, while keeping Pass@1 accuracy comparable to standard ReAct-style agents.
