WeMM-Embedding: WeChat’s Multi-Modal Embedding Technical Report
Tencent researchers describe WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes trained through large-scale…
Tencent researchers describe WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes trained through large-scale…
Microsoft researchers present AutoSaddler, a framework that treats agent-harness improvement as an offline learning problem using failure-trace diagnosis, structured patch…
ByteDance Seed researchers present DiffusionOPSD, an on-policy self-distillation framework that converts image-level rewards into intermediate training targets for diffusion models.…
Researchers present Secure On-Policy Distillation (SecOPD), a defensive fine-tuning approach that provides token-level feedback to guide training rather than the…
Researchers introduce CyberFactory, an open-source framework that integrates data construction, trajectory synthesis, and model training across three cybersecurity tasks: proof-of-concept…
Princeton University researchers present Recuris, an architecture for long-horizon agent tasks built on two coupled memory systems: a Working Memory…
Researchers introduce OPDVR, a method integrating on-policy distillation with verifiable reward signals for language model training. The approach reformulates on-policy…
Researchers introduce BPCO, a reinforcement learning recipe combining decoupled PPO, bounded value predictions, Monte Carlo targets, and adaptive advantage estimation…
Researchers developed a full-stack framework comprising GameUI-Taxonomy and G2WEngine that automatically extracts reusable UI assets from real gameplay footage and…
A research survey organizes the smart-glasses field around seven interdependent foundational capabilities and proposes an L0-L5 framework spanning capture, reactive…
Researchers propose Meta^n, a system where a fixed meta-operation is recursively applied to its own outputs, with each layer reading…
Researchers introduce a large-scale open video dataset containing 80 million videos totaling 10 million hours of content for multimodal learning…