Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
Researchers present R2-OPD, a method addressing misalignment between teacher-derived rewards and actual reasoning progress in on-policy distillation for language models.…
Researchers present R2-OPD, a method addressing misalignment between teacher-derived rewards and actual reasoning progress in on-policy distillation for language models.…
Researchers introduce LongWoF-Bench, a benchmark of 778 machine-verifiable tasks spanning code generation, agent synthesis, mathematical reasoning, and rule-following, alongside EvoMap,…
Researchers developed AstroPT, a transformer trained on millions of galaxy images, as a controlled testbed for interpretability research. By probing…
Researchers developed RIBOSPAN, a 1.61-billion-parameter RNA foundation model using dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to process…
Researchers introduce TPO, a face-free presentation-attack-detection dataset built from recordings of vegetables subjected to print, replay, and recapture processes. Using…
Researchers evaluate a hybrid quantum-inspired Kolmogorov-Arnold network (HQKAN) against a standard multilayer perceptron for arrhythmia classification using federated learning on…
Researchers introduce WorldToken, a robotic imitation learning approach that organizes multimodal inputs through a time-first architecture, fusing multiview images, proprioception,…
Researchers present EXPL-FR, a method using a lightweight adapter to align vision-language model encoders with frozen face recognition embedding spaces,…
Prime Agent is an open-source agent framework built around a persistent IPython REPL that implements the “Recursive Language Model” abstraction,…
Researchers introduce a lightweight, training-free audit for prefix invariance — the property that a sequence model’s representation at any position…
The paper proposes Quantization-Aware Healing (QAH), which recovers accuracy lost when a large language model is both structurally pruned and…
Researchers from Kakao Corp introduced a compute-efficient hyperparameter transfer framework for large-scale Mixture-of-Experts models, adapting Maximal Update Parameterization (μP) to…