Need help?
<- Back

Comments (27)

  • androiddrew
    So I'd like to see a Nemotron3 Ultra converted to a 1.58bit format then have them retrain on the open dataset.
  • kamranjon
    There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format.PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsaiI wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.
  • moinism
    The content on that page is too AI-generated to make sense to me; I don't understand what the model is for.
  • codeduck
    So... not a new subatomic particle discovery then.
  • secult
    There is not a single person mentioned on the website, github created 3 days ago, no real contact, everything hidden. Completely anonymous. Domain owner hidden.
  • Havoc
    Can’t say I’m a fan of containers for this. A big chunk of local LLM gains come (imo) from the open modular nature of llama.cpp and friends. Easy to modify. Easy to experiment.Containers are the proprietary binary blob in hardware world equivalent
  • yborg
    Largely outperformed by Ternary-Bonsai-8B by their own chart, doesn't seem clear what their special sauce is here.
  • sparse-Matrix
    Unfortunately it crashed out 'no space left on device' while installing the python demo/quickstart.Only problem was there is plenty of space on the device. PLENTY (not quite 750gb).
  • bmiekre
    This all sounds like middle-out
  • NetOpWibby
    There’s a new announcement every other day wrt models. How do y’all keep track of them all same know what’s decent? Good grief!And if it’s decent today, it’s shit in eight months! I tool hop as much as the next dev but this is a bit much.
  • madhu_ghalame
    The real strength of an 8B model is efficiency. It will be interesting to see the balance between performance and inference cost.
  • Alien1Being
    AI slop site with AI slop research...Blog populated with incoherent PR material generated by Yet Another AI.Sigh...
  • drbscl
    Slop article, slop site... slop model?
  • anon
    undefined
  • runtime_lens
    [flagged]