StarHarness is a framework that optimizes the executable environment surrounding a fixed language model by evolving prompts, tool interfaces, skills, and agent-loop configurations through stratified task sampling. The approach separates proposer-visible search tasks from proposer-hidden selection tasks while reserving held-out evaluation sets to measure genuine generalization. Across three enterprise benchmarks (ITBench, EnterpriseOps-Gym, and AutomationBench), the method achieved performance improvements of 20-35 percentage points using only 4-12 accepted modifications per environment, with gains transferring to unseen tasks and different model families without re-evolution.