Skip to content

LLM token costs are hard to predict and control

Cloud and tool cost5 posts from 5 people5d active+11 post in the last 7 days, 1 the 7 days before (steady)

Posts per day

Posts per day5 posts, Sep 2 to Sep 25

The posts behind it

5, newest first
PostDate
...inference bill look like if this was the case? The premise of your ARC-AGI-2 comparison is broken. > 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass. Ironic to accuse...Hacker News commentsAurornisSep 214 days ago
[Performance][ROCm][gfx1100] compressed-tensors silently enables fp8 KV cache, which is far slower than bf16 for paged decode attention on RDNA3### Summary On a Radeon RX 7900 XTX (gfx1100, RDNA3), a compressed-tensors checkpoint that carries a `kv_cache_scheme` gets its KV cache silently switched to `fp8`, even though `--kv-cache-dtype` is left at its default `auto`. In this configuration decode attention is...vllm-project/vllmyouyoulyzSep 1510 days ago
[Feature Request]: Idea: token-budgeted connected subgraph retrieval### Feature Description Hi! First of all, thanks for building LlamaIndex! I maintain [SteinerPy](https://github.com/berendmarkhorst/SteinerPy), a Python package for Steiner tree problems and related graph optimization problems. While looking at PropertyGraphIndex, I wondered...run-llama/llama_indexberendmarkhorstSep 72 weeks ago
Add prompt caching support for Mistral gateway endpoints...reduces our Mistral API spend and improves end-user latency. #### Why is it currently difficult to achieve this use case? The MLflow AI Gateway does not expose Mistral's `prompt_cache_key` parameter. Users cannot enable prompt caching without bypassing the gateway and calling...mlflow/mlflowarankeparthSep 52 weeks ago
...looks way more expensive from my quick calculations. Has anyone else calculated it? Is there any way to get Kimi K3 or Qwen 3.8 Max at a similar cost to what we're all paying by subscription for Claude or Codex? If not, I'd push back and say these Chinese models are actually...Hacker News commentsesperentSep 23 weeks ago

Companies and products named

Company or productPosts naming it
Claude1
Codex1
Qwen1
Anthropic1
LlamaIndex1
GitHub1
About this problem

Evidence

5 posts from 5 people in 4 places, about 2 a week over 20 days. Mostly on Hacker News comments, run-llama/llama_index, vllm-project/vllm. Tools named alongside: Claude, Codex, Qwen, Anthropic.

Frustration Frustration 0 of 3· Seen on GitHub issues, Hacker News

How it was grouped

Posts that state a pain and match the "llm-cost" rule. First post Sep 2, 2026, latest Sep 21, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: LLM token costs are hard to predict and control (4 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.