Skip to content

Context windows run out on real workloads

LLM apps13 posts from 13 people8d active+66 posts in the last 7 days, 0 the 7 days before (rising)

Posts per day

Posts per day13 posts, Aug 28 to Sep 25

The posts behind it

13, newest first
PostDate
Why doen't AI companies agree on a single standard?...them. I'm new to local models after using cloud models for years which "just works". I hate how many different engines, api formats, harnesses, tools I had to go through to optimise my setup. Maybe I'm missing something..r/LocalLLaMAu/accelerate_to_asiSep 24yesterday
Mac Studio M5 Ultra 96GB vs M5 Max 128GB for local LLMs?...it?. If you're running Flash-Next on a 96GB machine, how's it working in practice? Any issues with memory pressure or long contexts?. Is the speed of the Ultra worth giving up the extra memory, or will 96GB feel tight?. Thanks!.r/LocalLLaMAu/MxmtmSep 24yesterday
The token limit does not workWhere does the bug appear (feature/product)? Cursor IDE Describe the Bug In the “Agent Window” mode, you can choose to limit tokens to 256k, but the context window will grow to 500k (in the case of grok 4.7). In the cas…Cursor ForumMaksim_ErmakovSep 232 days ago
...from within Codex and tell it when to use it. > There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing Not true, it works again[1]. I confirm that it works both on 5.6 Sol and...Hacker News commentsnoname120Sep 223 days ago
ToolStrategy retries structured-output validation failures without any cap### Submission checklist - [x] This is a bug, not a usage question. - [x] I added a clear and descriptive title that summarizes this issue. - [x] I used the GitHub search to find a similar question and didn't find it. - [x] I am sure that this is a bug in LangChain rather than...langchain-ai/langchainSBerikSep 223 days ago
[Bug]: Refine response synthesizer refines repacked sub-chunks in reverse order### Bug Description When `Refine` processes a chunk that no longer fits the prompt once the existing answer is included, the loop repacks the chunk into smaller sub-chunks and pushes them back to the front of a deque: `deque.extendleft` reverses the order of its argument. The...run-llama/llama_indexHarsh23KashyapSep 214 days ago
...they work). And of course that can happen from random chance, and of course the LLM has no way to externally verify the extent to which randomness is in play (or the ground-truth probabilities). The paper makes clear that they used pre-trained, frontier models - in other words,...Hacker News commentszahlmanSep 92 weeks ago
Can not inccrease Context to 1 M limited 256K### What is the issue? Ollama already have many 1 M Context Size Models ! When we will be able to adjust Setting Slider past 256K Context Window Size? When can we switch to 1 M. Please Correct the Setting slider to allow to use Full COntxt Window Size example GLM 5.3 FLash and...ollama/ollamacomputerworxdevSep 92 weeks ago
However, this doesn't change the fact that you are pumping more and more tokens to a static model 's context window, even if you do compaction, the model is not more intelligent than previous turn. Nature doesn't work that way.Hacker News commentsbayindirhSep 43 weeks ago
What is the general design of these new math solving systems? [D]...and see if it can answer a question I have about higher dimensional geometry. I'm struggling to find a meaningful way to compose larger ideas from smaller ones. I can imagine it's relatively simple if you know what to do. . What things have you seen? Do you have any ideas you...r/MachineLearningu/tough-danceSep 43 weeks ago
Misc. bug: /slots restore yields no KV reuse on hybrid/recurrent and SWA models (context checkpoints are not persisted)### Name and Version `b94041a` (2026-08-16); behaviour still present in `36b1015` (2026-09-01) - the code path is unchanged and no commit in between touches slot-save checkpoints. ### Operating systems Windows ### GGML backends SYCL ### Hardware 2x Intel Arc Pro B70, oneAPI...ggml-org/llama.cppky095nSep 13 weeks ago
> Where are all these verbal tics coming from and why is it so hard to get rid of them? It’s a side effect of post-training for effectiveness and efficiency at technical tasks. Over time the models learn to pack as much information as possible into their available context...Hacker News commentsjiggawattsSep 13 weeks ago
[Bug]: GLM-5.3-Flash — ModelOpt NVFP4 checkpoints emit invalid UTF-8 byte tokens on SM120, while a compressed-tensors NVFP4 of the same model is clean### Your current environment Environment ### 🐛 Describe the bug Two ModelOpt-produced NVFP4 conversions of GLM-5.3-Flash emit **invalid UTF-8 byte-token sequences** during generation on SM120. A compressed-tensors NVFP4 conversion of the *same* model, served from the same image...vllm-project/vllmshing100Aug 284 weeks ago

Companies and products named

Company or productPosts naming it
Ollama2
vLLM2
Cursor1
Grok1
ChatGPT1
Codex1
About this problem

Evidence

13 posts from 13 people in 9 places, about 3 a week over 28 days. Mostly on Hacker News comments, r/LocalLLaMA, Cursor Forum. Tools named alongside: Ollama, vLLM, Cursor, Grok.

Frustration Frustration 1 of 3· Seen on GitHub issues, Hacker News, Reddit, Community forums

How it was grouped

Posts that state a pain and match the "context-window" rule. First post Aug 28, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Context windows run out on real workloads (13 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.