CDN and cybersecurity giant Cloudflare held its Q2 earnings call on Thursday, during which chief financial officer Thomas Seifert gave a concerning prediction about what the web...
What I wonder is: how long before “bot speak” becomes unintelligible to humans.
Already: AI slop is so voluminous, redundantly repetitive, thoroughly complete that it defies complete comprehension due to soporific effects.
LLM agents are writing English (or Spanish, or Chinese, Hindi, whatever… and I wonder which they’re “best” at) documents, ostensibly for human review and consumption, but 90%+ of what I have my LLM agents write is exclusively consumed by other agents, and I’m constantly encouraging them to make their writings more easily comprehended and accessed by other agents - it seems like the English translation is pretty useless for that layer - what would a token stream look like instead?
There’s no “LLM native language” to convert to for efficiency.
Yet. With 1000x as much bot-content being generated on the web as human generated content, the “bot spoken” training set will grow rather quickly, and I see no reason for it not to evolve in a similar way to how human spoken languages evolve.
LLM agents are writing English (…) documents, ostensibly for human review and consumption, but 90%+ of what I have my LLM agents write is exclusively consumed by other agents, … and accessed by other agents - it seems like the English translation is pretty useless for that layer …
But remember, the LLM is trained on human language. It doesn’t understand what it is talking about, it just simulates humans talking about it. I don’t think there is a more fundamental token stream there to be uncovered.
Example of something I had to lookup to understand: (/hj)
Back when I was in school, Autism was a 1/10,000 Dx. 20 years ago it was crashing down from 1/100 to 1/50, checking now… 1/31 today. They’ve lumped so much into the Autism diagnosis that it’s meaningless anymore; it covers so many varied conditions and severities, and the stereotypes don’t fit most recipients of the Dx anymore.
What I wonder is: how long before “bot speak” becomes unintelligible to humans.
Already: AI slop is so voluminous, redundantly repetitive, thoroughly complete that it defies complete comprehension due to soporific effects.
LLM agents are writing English (or Spanish, or Chinese, Hindi, whatever… and I wonder which they’re “best” at) documents, ostensibly for human review and consumption, but 90%+ of what I have my LLM agents write is exclusively consumed by other agents, and I’m constantly encouraging them to make their writings more easily comprehended and accessed by other agents - it seems like the English translation is pretty useless for that layer - what would a token stream look like instead?
The tokens ARE language. They’re just words or parts of words.
If you’re using a western LLM, English IS its strongest language. Otherwise it might be Chinese or English.
There’s no “LLM native language” to convert to for efficiency. Maybe pseudocode or actual code for things where you need to disambiguate.
Yet. With 1000x as much bot-content being generated on the web as human generated content, the “bot spoken” training set will grow rather quickly, and I see no reason for it not to evolve in a similar way to how human spoken languages evolve.
But remember, the LLM is trained on human language. It doesn’t understand what it is talking about, it just simulates humans talking about it. I don’t think there is a more fundamental token stream there to be uncovered.
Humans may end up seeing a translated audit layer while the actual machine conversation looks more like APIs than language.
at some point we will give a name of this phenomenon of internet users who are completely undecipherable for normal people … and call it autism (/hj)
Not too keen on the ableism, but a handjob is a handjob ¯\_(ツ)_/¯
Example of something I had to lookup to understand: (/hj)
Back when I was in school, Autism was a 1/10,000 Dx. 20 years ago it was crashing down from 1/100 to 1/50, checking now… 1/31 today. They’ve lumped so much into the Autism diagnosis that it’s meaningless anymore; it covers so many varied conditions and severities, and the stereotypes don’t fit most recipients of the Dx anymore.
The spectrum is a big place.
A big place with very varying “special needs” ranging from none all the way through fully supported living.