• MangoCats@feddit.it
    link
    fedilink
    English
    arrow-up
    5
    ·
    17 hours ago

    What I wonder is: how long before “bot speak” becomes unintelligible to humans.

    Already: AI slop is so voluminous, redundantly repetitive, thoroughly complete that it defies complete comprehension due to soporific effects.

    LLM agents are writing English (or Spanish, or Chinese, Hindi, whatever… and I wonder which they’re “best” at) documents, ostensibly for human review and consumption, but 90%+ of what I have my LLM agents write is exclusively consumed by other agents, and I’m constantly encouraging them to make their writings more easily comprehended and accessed by other agents - it seems like the English translation is pretty useless for that layer - what would a token stream look like instead?

    • boonhet@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      4
      ·
      10 hours ago

      The tokens ARE language. They’re just words or parts of words.

      If you’re using a western LLM, English IS its strongest language. Otherwise it might be Chinese or English.

      There’s no “LLM native language” to convert to for efficiency. Maybe pseudocode or actual code for things where you need to disambiguate.

      • MangoCats@feddit.it
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 hours ago

        There’s no “LLM native language” to convert to for efficiency.

        Yet. With 1000x as much bot-content being generated on the web as human generated content, the “bot spoken” training set will grow rather quickly, and I see no reason for it not to evolve in a similar way to how human spoken languages evolve.

    • Logi@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      11 hours ago

      LLM agents are writing English (…) documents, ostensibly for human review and consumption, but 90%+ of what I have my LLM agents write is exclusively consumed by other agents, … and accessed by other agents - it seems like the English translation is pretty useless for that layer …

      But remember, the LLM is trained on human language. It doesn’t understand what it is talking about, it just simulates humans talking about it. I don’t think there is a more fundamental token stream there to be uncovered.

    • eicker@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      12 hours ago

      Humans may end up seeing a translated audit layer while the actual machine conversation looks more like APIs than language.

    • gandalf_der_13te@feddit.org
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      15 hours ago

      “bot speak” becomes unintelligible to humans.

      at some point we will give a name of this phenomenon of internet users who are completely undecipherable for normal people … and call it autism (/hj)

      • MangoCats@feddit.it
        link
        fedilink
        English
        arrow-up
        3
        ·
        14 hours ago

        Example of something I had to lookup to understand: (/hj)

        Back when I was in school, Autism was a 1/10,000 Dx. 20 years ago it was crashing down from 1/100 to 1/50, checking now… 1/31 today. They’ve lumped so much into the Autism diagnosis that it’s meaningless anymore; it covers so many varied conditions and severities, and the stereotypes don’t fit most recipients of the Dx anymore.