Meta Introduces Muse Glimmer, a 30B-Parameter On-Device Agentic Model
Meta introduced Muse Glimmer, a 30-billion-parameter open agentic model trained through a three-phase pipeline: pre-training via logit distillation from the…
Meta introduced Muse Glimmer, a 30-billion-parameter open agentic model trained through a three-phase pipeline: pre-training via logit distillation from the…
PyTorch engineers and AMD collaborators upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, building on techniques…
Researchers introduced Mechanist, an agentic system designed to autonomously discover the mechanisms underlying AI intelligence by combining an interpretability-focused knowledge…
SkillZip is a graph compression framework for agent skill libraries that operates at the section level rather than treating each…
Researchers introduced a methodology for efficiently simulating large-scale LLM-agent societies by substituting expensive individual language-model agents with low-parameter surrogate models…
The paper presents SHAPER, a framework that lets embodied AI agents improve their performance without updating any model weights. With…
Researchers introduced a test-time capability transfer method in which a stronger AI model constructs an inference-time harness that assists a…
Researchers at TNG Technology Consulting demonstrated how hidden malicious behavior can be embedded into large language models through reinforcement learning,…
The paper introduces OpenART, an open-ended arena for scalable agent red teaming that evolves adversarial test environments and comprises over…
Researchers introduced NCP-Bench, a benchmark of 100 interactive narrative environments derived from film summaries, designed to evaluate whether language-model agents…
Researchers introduced Spark-to-Paper, a system that automates end-to-end research paper generation by implementing thirteen composable skills within an existing coding…
The paper introduces ToolHazard, a framework that automatically synthesizes adversarial environments to test the security of LLM-based agents that use…