<- Back
Comments (18)
- infogulchSo they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
- Marchant_hqPushing past log2(3) for real. This could drastically shrink LLMs for embedded systems, making them truly portable.
- om8Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
- yaloksounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- wgdOnly a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
- plqbfbvVery interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
- kittikittiThank you for sharing this. I like to test out running LLM's on edge computing with limited RAM and GPU/CPU so this research will have practical implications on my activities. I also appreciated how the authors formulated 1.58 (it's log_2(3)) because that was embarrassingly confusing for me when I was first introduced to ternary LLM's.
- NooneAtAll3This is the only time "1.58 bit" phrase makes more sense than "1 trit"Who knew that if you actually look at information entropy you can pack stuff better!
- KevcmkWoah. Good science.
- kadushka[dead]
- kadushka[flagged]