Back to papers
May 14, 2026cs.LGcs.AR

A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models

Categories

cs.LG, cs.AR