Skip to content

LLMs give confidently wrong answers

LLM apps10 posts from 10 people9d active+44 posts in the last 7 days, 4 the 7 days before (steady)

Posts per day

Posts per day10 posts, Sep 2 to Sep 25

The posts behind it

10, newest first
PostDate
Again, no one on this thread is claiming that AI "doesn't work." But it's not delusional to say it isn't working in some area, like the legal area. It's hallucinating cases and the law. We see it every week. I had a client tell me they used AI last week and the info it gave them...Bluesky search@bwkemper.bsky.socialSep 223 days ago
Non-technical COO "discovered" AI coding, generated an entire dashboard in a single HTML file, and now wants me to review and maintain it....politically without setting your career on fire, while also not inheriting a maintenance nightmare? .r/sysadminu/boyrokSep 223 days ago
And they hallucinate errors in the millions and struggle with financial data that is in a layout that isn’t in the training data. Ie balance sheet etc. Been trying to add more AI to my workflow but it just doesn’t work (yet) - not in the same way as vibe coding does The...Hacker News commentsHavocSep 214 days ago
...occurrence. Anti-AI folks have ignorantly painted this as the models being fundamentally unreliable, but you also have a non-zero probability of being struck by lightning or eaten by a shark.Hacker News commentsCuriouslyCSep 196 days ago
...experience has been a constant fractal of hallucination, in which I have to constantly struggle to keep my grip (and Claude's grip) on actual reality, because it's so willing to make up surface-plausible nonsense. ______ P.S. Also, in another task today, it thought pip wasn't...Hacker News commentskragenSep 178 days ago
...LLMs for a classification task People out there are so resigned to the models being unreliable that they are really doing things like hallucinating deliberately, and then matching the hallucinations to embeddings - https://softwaredoug.com/blog/2026/08/10/hypothetical-classi......Hacker News commentszenlikethatSep 1510 days ago
And you tested this, that the presence of the last full stop flips the outcome with a large SOTA model? Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.Hacker News commentsuser43928Sep 1411 days ago
/hallucination problems that the autoregressives solved ~2 years ago. So it's not ready yet but improving. Also, FFS why is Grok the only model that knows how to do parallel tool calls? Such a useful ability and nobody else trains it in. Or if they do it just doesn't work.Hacker News commentsoctoberfranklinSep 1312 days ago
Spanda: Sub-microsecond LLM epistemic uncertainty in Rust...(Nature 2024) is great at detecting hallucinations, but the quadratic NLI cross-encoder bottleneck (90ms GPU overhead with DeBERTa) makes it unusable for real-time production serving. We found that exact-match normalized entropy (R_sc) achieves the same discriminative AUROC on...Ask HNliquidngasSep 112 weeks ago
...themselves are saying that gaming the peer-review system with barrage of AI is extremely frustrating and pointless: Peer review is unpaid work that we do (often on nights and weekends) because peer review on our own work is so valuable. Spending hours going through a submission...Hacker News commentsdelis-thumbs-7eSep 23 weeks ago

Companies and products named

Company or productPosts naming it
GitHub1
PyPI1
Claude1
Grok1
ChatGPT1
Power BI1
About this problem

Evidence

10 posts from 10 people in 4 places, about 3 a week over 21 days. Mostly on Hacker News comments, Ask HN, r/sysadmin. Tools named alongside: GitHub, PyPI, Claude, Grok.

Frustration Frustration 2 of 3· Seen on Hacker News, Reddit, Bluesky

How it was grouped

Posts that state a pain and match the "hallucination" rule. First post Sep 2, 2026, latest Sep 22, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Momentum: fading to steady
  • 2026-09-25 Added: LLMs give confidently wrong answers (9 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.