A paper introduces LOPD, a method that makes an agent’s own privileged context end-to-end learnable from experience rather than relying on hand-crafted supervision, by retrieving relevant past experiences and compressing them into latent tokens that guide a self-teacher during training. On tool-use and code-generation tasks, LOPD outperformed RLVR, OPSD, SDPO, and Skill-SD baselines while using less than 30% of the rollout budget required by GRPO and Skill-SD.