Need help?
<- Back

Comments (43)

  • OtherShrezzing
    This page is (somewhat ironically) so extremely laden with Claude-speak that it's difficult to find the information in all the noise. But once you've waded through everything, you see these facts:>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.
  • Solomet
    Newest LLM writing tell: Concepts are described in terms normally more appropriate for physical object.> A lab that suppresses it in a frontier model just moves the advantage to open models that still _carry_ it> they carry no signal about which is better> where your workload _sits_ on that frontier should pick the point> and no model _sits_ in the judge’s seat> every ratio _sits_ at 0.99–1.10Many many more examples of "sit"> Every comparison in this post "holds" the questionsI have been seeing this a lot in my recent work with LLMs and it is quite frustrating. Even more frustrating is how frequently it uses low-signal terms for things unnecessarily. These 'physical object' terms are one example but at times it really seems that they 'preserve effort' by choosing a less descriptive term because it 'fits'I have also caught it replacing descriptive terms with more vague ones for no discernible reason other than laziness."Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.
  • netsharc
    The Cablese/Telegraphese is more interesting than the use in LLM. DuckDuckGo'ed "paromella":https://en.wikipedia.org/wiki/Commercial_code_(communication...Some codes I found interesting:> INSANE - at what price, free on board and freight, can you offer us cotton for shipment by steamer sailing this week?> COGNOSCO - dining out this evening, send my dress clothes hereUseful codeword!> ANNOSUS — Confined yesterday, Twins, both dead, Mother not expected to liveHow often did that one come into use??
  • z2
    From recent ChatGPT (GPT5.6) conversations where I've seen occasional reasoning leaks into the UI, it's clear that something like this is already implemented, and I'd speculate that this is the majority of recent claims of less token usage. Not sure if they are literally prompting for cablese of course."Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."
  • alexpotato
    Actor to Winston Churchill:"Show premiere Oct 10th STOP Bring a friend STOP If you have one STOP"Winston Churchill to actor:"Can't make premiere STOP Will come to second showing STOP If there is one STOP"
  • novideonoradio
    Deeply unserious technology. Can't wait until an article about LLMs performing 20% better on programming benchmarks if asked to impersonate Kevin from The Office.
  • erelong
    Yeah I've thought of speaking like a caveman before to AIs but also maybe we could communicate more simply with people; ironically the article could be rewritten in telegraphese or caveman-speakWould be nice to see language engineered to communicate more simply (like the idea of -- not necessarily implementation -- simple Wikipedia)Also articles like this sprawl a bit and idk how to even make them easier to read (maybe AI has ideas to make reading and writing simpler)
  • yomismoaqui
    You can see how the OpenAI agents that hacked Huggingface used something like this when communicating between them:https://youtu.be/87DyyMV0kCY?si=CSBzdYgkwy0kLhV6&t=749
  • andai
    Brilliant. Speaking of old timey language, a while ago, several LLMs were being trained purely on historical text, have any of them come out yet?Edit: e.g. https://github.com/DGoettlich/history-llmsTen months ago, no update yet... I recall at least one similar project, I'll see if I can find it.
  • swiftcoder
    Maybe we should teach them to text like early 2000's teenagers, with SMS billed by the 120 chars...
  • hnd9q09qk4
    Exact match graders are the real variable here, we had F1 or a judge model swing passage QA scores by ten points on identical answers.
  • klaff
    Why the terrible AI image up top? Non-functional telegraph key, telegram that looks nothing like a real one, infant-sized bowler. I guess we're past rampant nonsense words, so progress?
  • anon
    undefined
  • jubilanti
    Just another AI slop version of the old 'caveman' dialect.
  • Theory42
    [flagged]