The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchAugust 12, 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Post-training quantization pipelines routinely leave the softmax output layer in high precision. Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all parameters, so a nominal ``2-bit'' model with an fp16 head can store several times as many bits per weight. We pose softmax-laye...

Read Original Article →

Source

http://arxiv.org/abs/2608.12026v1