This paper studies how visual understanding and generation tasks interact inside unified multimodal models through controlled experiments. The authors propose a task-decoupled architecture that specializes conflicting computations while preserving shared semantic knowledge, and show end-to-end unified models outperform modular planner-executor pipelines on complex tasks. Experiments at the representation, task, and system level identify the conditions under which task coexistence becomes genuine synergy rather than interference.
