Skip to content

Running LLMs locally hits memory and speed limits

MLOps5 posts from 5 people5d active+22 posts in the last 7 days, 1 the 7 days before (steady)

Posts per day

Posts per day5 posts, Sep 6 to Sep 25

The posts behind it

5, newest first
PostDate
How much space and memory would the next generation ngram models take?...to SSDs and RAM. This could free up VRAM memory significantly. But, then wouldn't the bottleneck become the RAM and SSD storage? Considering just how much information these models are trained upon, wouldn't even the small 27B models take up hundreds of GB RAM or terabytes of...r/LocalLLaMAu/accelerate_to_asiSep 24yesterday
Experimenting with hypersurface-constrained dynamic weight updating [P]...test an approach to reduce the number of the model's training parameters, since the main bottleneck in training is VRAM.. The core idea is very similar to Universal Transformer. Let's take just a single decoder block and iteratively pass the input through it L times in a loop....r/MachineLearningu/manila_danimalsSep 196 days ago
Most vector quantizers are the same 6 primitives strung in a different order [R]...art. So this summer we set out to survey it. We found more papers than we expected, and no way to compare them: every paper measured different metrics, on different datasets, tuned for different hardware. We could not find a single systematic evaluation of the leading methods...r/MachineLearningu/soryx7Sep 178 days ago
MLX compile-cache CHECK failed spams logs on non-MLX hardware (Windows, no CUDA/Apple Silicon)On every `ollama` command (including `ollama list`, `ollama serve`, `ollama run`), the following error is printed to the console before the actual output, regardless of which command was run: Sep 3 2026 16:31:11 - ERROR - generated.c:2650 - CHECK failed: mlx_compile_cache_new_...ollama/ollamaznanovikSep 72 weeks ago
[Bug]: CPU MoE kernel produces NaN logits under torch.compile for Qwen3_5MoeForConditionalGeneration — reproduces with AND without GPTQ-Int4 quantization, fixed by --enforce-eager...mode worth flagging on its own - a healthy pod/process gives zero signal that generation is broken. **Workaround: `--enforce-eager` completely fixes it, on both checkpoints.** With the identical checkpoint, config, and request, eager mode produces correct, coherent generations...vllm-project/vllmixcansSep 62 weeks ago

Companies and products named

Company or productPosts naming it
GitHub2
Pinecone1
PyTorch1
OpenAI1
vLLM1
Hugging Face1
About this problem

Evidence

5 posts from 5 people in 4 places, about 2 a week over 19 days. Mostly on r/MachineLearning, r/LocalLLaMA, vllm-project/vllm. Tools named alongside: GitHub, Pinecone, PyTorch, OpenAI.

Frustration Frustration 0 of 3· Seen on Reddit, GitHub issues

How it was grouped

Posts that state a pain and match the "local-llm" rule. First post Sep 6, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Running LLMs locally hits memory and speed limits (5 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.