Researchers introduce LLaDA-Image, an open framework pairing a 6B diffusion transformer with a frozen vision-language module built on the LLaDA2.0-Mini backbone, trained via image-only pre-training and the Muon optimizer rather than early paired image-text data. The model produces photorealistic images that closely follow fine-grained editing instructions, and a distilled variant, LLaDA-Image-Turbo, enables fast inference in just 2-4 sampling steps. On the Qwen-Image-Bench benchmark it sets a new state-of-the-art among open-source models on both English and Chinese tracks, and the team has released model weights, training code, and detailed recipes.