S3Gym is a new interactive benchmark testing whether LLM agents can turn accumulated experience into genuine self-improvement, evaluating three coupled capabilities — self-testing, self-judging, and self-improvement — across seven text-based games with executable verifiers. The study finds self-improvement is neither automatic nor uniform: summarized experience helps when strategies compress into reusable rules but can underperform raw history otherwise, and parameter-level training yields large gains on some tasks alongside unstable results on others.
