EM^2Mem is an event-centric multimodal memory framework for long-video question answering that binds multimodal evidence to event anchors during memory construction rather than retrieving isolated captions or frames. The resulting “generation-ready” memory cells align text, video, temporal context, and relational data. Across three video QA benchmarks the method improves accuracy and evidence recall while cutting inference latency and token consumption.