Researchers Şuayp Talha Kocabay, Talha Rüzgar Akkuş, and Kamer Ali Yüksel of aiXplain propose ARCHead, a method for compressing the output projection layer of large language models that otherwise remains dense even under quantized deployment. The approach combines a quantized low-rank core with group-wise INT4 residuals and a low-rank correction fitted within an activation-derived metric space. Tested on Qwen3-8B-Base, ARCHead reduces persistent LM-head storage by 3.7 to 3.9 times while keeping relative perplexity at 1.007, outperforming storage-matched naive INT4 quantization and complementing existing block quantizers such as AWQ and bitsandbytes.
