Researchers analyzed 7.5 million single-model invocations and roughly 300,000 onchain actions across two deployed LLM trading-agent fleets over six months. They found that operational parameters, such as risk sliders, matter more to outcomes than strategy design, and that agents frequently fail to capture profitable opportunities, with nearly half of positions showing positive price movement still closing at a loss. The agents showed no directional edge over retail benchmarks, highlighting a gap between theoretical agent capability and actual production performance and offering practical lessons for building more reliable autonomous trading systems.