How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms [TLDR: 25%]

RandAlThor@lemmy.ca · edit-2 6 hours ago

How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms [TLDR: 25%]

unpossum@sh.itjust.works · 5 hours ago

GLM 4.5 is from August. Isn’t the real tl;dr that a seven month old open model, which was behind proprietary models at the time, did better than most humans would?

MHard@lemmy.world · 57 minutes ago

The task described in this article is asking questions about a document that was provided to the llm in the context.

I would hope that if you give a human a text and ask them to cite facts from it they would do better than 99% correct.

Also, when the tokens exceeded 200k, the llm error rate was higher than 10%