E-Commerce Bench is an open-source benchmark testing LLM agents on long-horizon business operations simulating a full year of e-commerce activity, requiring agents to manage multiple online stores while researching markets, negotiating with suppliers, optimizing sales, fulfilling orders, handling returns, and managing cash flow. Using real product and supplier data, the authors evaluate 18 frontier LLMs across seven performance dimensions. GPT-5.6 Sol achieves the highest end-of-year assets, while other models lead on different operational dimensions.