The paper presents SHAPER, a framework that lets embodied AI agents improve their performance without updating any model weights. With the underlying vision-language model kept frozen, the system evolves two external components instead — reusable procedural skills and context-construction code — using feedback gathered from interactions with the target environment. Evaluated on the VLABench and ESI-Bench benchmarks across different action interfaces, SHAPER achieved improvements comparable to or exceeding supervised fine-tuning baselines, suggesting that skill-and-harness optimization is a practical route to self-improving embodied agents when retraining is expensive or unavailable.