No reliable way to evaluate LLM app quality
LLM apps29 posts from 28 people15d active+33 posts in the last 7 days, 3 the 7 days before (steady)
Posts per day
The posts behind it
29, newest firstCompanies and products named
About this problem
Evidence
29 posts from 28 people in 7 places, about 6 a week over 30 days. Mostly on ggml-org/llama.cpp, Hacker News comments, mlflow/mlflow. Tools named alongside: llama.cpp, NVIDIA, Hugging Face, Qwen.
Frustration Frustration 1 of 3· Seen on GitHub issues, Hacker News, Reddit
How it was grouped
Posts that state a pain and match the "evals" rule. First post Aug 27, 2026, latest Sep 25, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method
History
- 2026-09-25 Added: No reliable way to evaluate LLM app quality (26 posts)