Researchers present SenseNova-U1.5, an 8-billion-parameter multimodal model that performs visual understanding, reasoning and generation within a single unified architecture, eliminating the separate encoder and VAE components typically used for these tasks. The system uses optimized specialized experts and multi-expert distillation to improve image fidelity, text rendering and complex visual composition. The authors report the model advances image fidelity, text rendering, complex composition, multi-reference editing and interleaved generation, and it generalizes to complex visual instructions despite limited structured training data.