TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Researchers developed TLive-Omni, a multimodal model that processes image, video, audio, and text inputs into a unified representation space for…
Researchers developed TLive-Omni, a multimodal model that processes image, video, audio, and text inputs into a unified representation space for…
Researchers introduce ConceptEdit-12M, a 12-million-pair dataset covering more than 1,000 fine-grained edit concepts to improve image editing models, alongside a…
Researchers introduce Block3D, a text-to-3D generation framework that partitions discrete shape-token sequences into blocks, generates them sequentially, and jointly denoises…
Ant Group researchers propose the User Behavioral Densing Law, which quantifies how tokenization capacity should scale with data size for…
Researchers present RISE, an adaptive framework for world action models that dynamically decides when to stop imagination rollouts based on…
Researchers present ReWorld, an interactive world model that addresses the competing demands of real-time control and long-horizon memory through architectural…
Researchers introduce GameXpert-Bench, an evaluation framework with three benchmark tracks assessing coding agents across the game-development lifecycle: initial creation (GameGen,…
Researchers introduce ARC (Advantage Regularization via Conditioning), a training methodology addressing fairness issues when comparing diverse interaction behaviors in reinforcement…
Researchers present AutoResearch, a two-stage autonomous research system combining idea generation with idea execution. In the generation phase, the system…
Researchers present R2-OPD, a method addressing misalignment between teacher-derived rewards and actual reasoning progress in on-policy distillation for language models.…
Researchers introduce LongWoF-Bench, a benchmark of 778 machine-verifiable tasks spanning code generation, agent synthesis, mathematical reasoning, and rule-following, alongside EvoMap,…
Researchers developed AstroPT, a transformer trained on millions of galaxy images, as a controlled testbed for interpretability research. By probing…