VLAct is a vision-language-action model using representation-centric continued pre-training on diverse robot data, applying VLM-prior preservation, multi-head action co-supervision, and cross-embodiment action layouts to transfer knowledge across robot morphologies. It achieves competitive results on benchmarks including LIBERO-Plus and RoboTwin, and demonstrates strong zero-shot transfer to unseen embodiments such as humanoids.