MemTrapBench Finds That Giving LLM Agents Memory Can Backfire
Researchers introduce MemTrapBench, a benchmark showing that retrieved memories can actually harm large language model task performance even when the…
Researchers introduce MemTrapBench, a benchmark showing that retrieved memories can actually harm large language model task performance even when the…
Tencent researchers present SkillEvo, addressing the problem that AI agent skills typically fail to improve from interaction failures beyond a…
Researchers at Shanghai Jiao Tong University present Repo0, a system for generating complete software projects with proper modular architecture directly…
Z.ai’s GLM-5.3 model was trained with post-training scaling improvements aimed at vulnerability detection, achieving 84.5% on the CyberGym benchmark for…
Researchers address a weakness in graph-based retrieval-augmented generation systems, where automatically constructed knowledge graphs often suffer from thematic irrelevance, logical…
Liquid AI released Q4_0 GGUF checkpoints for four LFM2.5 models using Quantization-Aware Distillation (QAD), a technique that distills a high-precision…
Dharma AI describes a constraint-aware GPU allocation system that improved cluster utilization by up to 33 percentage points compared to…
Sentence Transformers v6.0 introduces a MultiVectorEncoder that enables ColBERT-style late interaction retrieval, keeping one vector per token instead of compressing…
IBM Research’s ALTK-Evolve framework examines how AI agents benefit from memory depending on their capability level, testing the approach across…
Hugging Face organized a large-scale hackathon in which 1,221 community members used AI agents to reproduce claims from 2,226 ICML…
Amazon and Hugging Face describe an end-to-end robotics workflow that combines Strands Agents, LeRobot, and Hugging Face storage buckets. The…
OpenAI introduced ChatGPT for Teens on August 18, 2026, an age-gated version of the chatbot that automatically enrolls users it…