<- Back
Comments (14)
- steevehttps://x.com/harshagundal/status/2100044305536889015?s=20> They were building in stealth for 2 years, I was building in stealth for 2 hours…> Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
- razsterYou can ask this Redditor saying he made it. https://old.reddit.com/r/LocalLLaMA/comments/1wihgum/i_liter... I think.
- mmastracAny diffusion model is potentially a Jev in disguise: https://github.com/vllm-project/vllm/pull/57250Runs ~0.2s per decision on my DGX Spark. 10/10 programming language detection 9/10 human language detection 10/12 unit magnitude comparison All incorrect answers are marked with low-P.It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.
- vrcOut of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
- rochansinhaCan play Doom too - https://x.com/vinnylarouge/status/2100281651930513460
- tomrodI like it! I suspect Jev may have more going on under the hood, but I like the idea of efficient universal transformers
- _superposition_That was super quick.