Need help?
<- Back

Comments (35)

  • VoidWhisperer
    https://github.com/nex-agi/Nex-N2/issues/4Seems that they didn't make/train a new novel model, they did a mix of two existing models and then gave it an instruction to say it was 'Rio, trained by Rio AI Labs'
  • mettamage
    https://xcancel.com/ZenMagnets/status/2065796012820848699Correct me if I'm wrong but reading through the comments of the thread this seems to be post training/fine tuning.
  • adrian_b
    > Post-trained from Qwen 3.5 397BModel Card:https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B
  • Aurornis
    A city government funding a fine-tune of a model is interesting.As for the benchmarks: If you spend any time playing with fine tunes of published models you know that benchmarks are gamed so much that they're a useless indicator of performance for models from small teams. It's too easy to fine tune a model to perform well on the benchmarks, release it, put a line on your resume saying you released a model that beat the major labs on benchmarks, and then try to use that to jump into a new job. The temptation is high.There are a lot of fringe models and fine tunes that claim to have better performance on some benchmark. Then you try to use them and find they're often worse at general tasks than the base model.I would wait and see if these results hold across other benchmarks. It's cool that the city is doing something with AI, but this is something where extraordinary claims require extraordinary evidence. I doubt a small, previously unknown team has unlocked something secret that the team who made Qwen couldn't figure out. It's more likely it was fine tuned for a specific outcome (possibly these benchmarks) and performance in other areas was reduced as a consequence.
  • HeliumHydride
  • arjie
    Benchmaxxing is the new “have a crypto trading strategy”. No one is impressed by it except non practitioners.
  • anon
    undefined
  • anon
    undefined
  • mrandish
    > Rio de Janeiro's city government model...Because... lack of a good open weight LLM is a pressing need high on the municipal priorities list for Rio de Janeiro citizens?
  • pelasaco
  • xbar
    Sexy.
  • anon
    undefined
  • cuzezzzbbfofai
    [flagged]
  • hmokiguess
    Never let them know your next move
  • ramon156
    Every day I'm reminded why I don't spend time on twitter. What use does it have to claim "X is better than Y in benchmark Z, disagreeing with that means disagreeing with me"Information is power, dick measurements are not.