LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
Researchers introduce a large-scale open video dataset containing 80 million videos totaling 10 million hours of content for multimodal learning…
Researchers introduce a large-scale open video dataset containing 80 million videos totaling 10 million hours of content for multimodal learning…
Tencent Hunyuan researchers introduce CAFE, a framework in which a shared-parameter model alternates between agent and critic roles to build…
Researchers present AgentRoom, a real-time collaborative editing protocol that enables concurrent multi-agent LLM coding through a shared filesystem built on…
Researchers convert agent execution traces into compact finite-state machines containing 7 to 43 states that reconstruct held-out data with 0.997…
Researchers present Entropy-Valley, a training-free technique for selecting target canvas lengths in masked diffusion machine translation that scores candidate lengths…
Alibaba researchers describe DREAM, an agentic control architecture that adds a perception-aware policy layer above existing recommendation pipelines, comprising an…
Researchers investigate how intermediate language artifacts in multi-stage LLM agent workflows can inadvertently weaken operational constraints, evaluating 1,296 synthetic episodes…
Researchers propose MoTE, a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal…
Kai Zhao presents TorchMorph, a lightweight PyTorch extension implementing 22 morphological operators as fused CUDA kernels that operate directly on…
Researchers describe the Station, an autonomous multi-agent mathematical discovery system where AI agents from different model families collaboratively pursue research…
Researchers developed LAWA, a world action model architecture that represents future intentions using compact latent actions rather than generating explicit…
Researchers developed EchoWM, an omnimodal world model that generates synchronized 720p video, environmental sound, music, and speech while responding to…