<- Back
Comments (27)
- androiddrewSo I'd like to see a Nemotron3 Ultra converted to a 1.58bit format then have them retrain on the open dataset.
- kamranjonThere's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format.PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsaiI wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.
- moinismThe content on that page is too AI-generated to make sense to me; I don't understand what the model is for.
- codeduckSo... not a new subatomic particle discovery then.
- secultThere is not a single person mentioned on the website, github created 3 days ago, no real contact, everything hidden. Completely anonymous. Domain owner hidden.
- HavocCan’t say I’m a fan of containers for this. A big chunk of local LLM gains come (imo) from the open modular nature of llama.cpp and friends. Easy to modify. Easy to experiment.Containers are the proprietary binary blob in hardware world equivalent
- yborgLargely outperformed by Ternary-Bonsai-8B by their own chart, doesn't seem clear what their special sauce is here.
- sparse-MatrixUnfortunately it crashed out 'no space left on device' while installing the python demo/quickstart.Only problem was there is plenty of space on the device. PLENTY (not quite 750gb).
- bmiekreThis all sounds like middle-out
- NetOpWibbyThere’s a new announcement every other day wrt models. How do y’all keep track of them all same know what’s decent? Good grief!And if it’s decent today, it’s shit in eight months! I tool hop as much as the next dev but this is a bit much.
- madhu_ghalameThe real strength of an 8B model is efficiency. It will be interesting to see the balance between performance and inference cost.
- Alien1BeingAI slop site with AI slop research...Blog populated with incoherent PR material generated by Yet Another AI.Sigh...
- drbsclSlop article, slop site... slop model?
- anonundefined
- runtime_lens[flagged]