Writing Down the Line Between Luck and Skill
FINAL-Bench launched a live financial forecasting benchmark that scores entrants against a “luck ceiling” – the 95th-percentile return of 20,000…
FINAL-Bench launched a live financial forecasting benchmark that scores entrants against a “luck ceiling” – the 95th-percentile return of 20,000…
EdgeFirst released a Model Zoo of YOLO object-detection and segmentation models, spanning YOLOv5 through YOLO26, validated across multiple edge hardware…
IBM released granite-speech-5.0-470m-turboctc, a pair of compact, encoder-only speech recognition models built from a stack of 16 Conformer blocks with…
Researchers introduced TAVR, a talking-avatar generation method that uses short video clips rather than a single reference image to preserve…
Tencent researchers describe WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes trained through large-scale…
Microsoft researchers present AutoSaddler, a framework that treats agent-harness improvement as an offline learning problem using failure-trace diagnosis, structured patch…
ByteDance Seed researchers present DiffusionOPSD, an on-policy self-distillation framework that converts image-level rewards into intermediate training targets for diffusion models.…
Researchers present Secure On-Policy Distillation (SecOPD), a defensive fine-tuning approach that provides token-level feedback to guide training rather than the…
Researchers introduce CyberFactory, an open-source framework that integrates data construction, trajectory synthesis, and model training across three cybersecurity tasks: proof-of-concept…
Princeton University researchers present Recuris, an architecture for long-horizon agent tasks built on two coupled memory systems: a Working Memory…
Researchers introduce OPDVR, a method integrating on-policy distillation with verifiable reward signals for language model training. The approach reformulates on-policy…
Researchers introduce BPCO, a reinforcement learning recipe combining decoupled PPO, bounded value predictions, Monte Carlo targets, and adaptive advantage estimation…