Skip to content

Running llama.cpp locally hits memory and speed limits

MLOps7 posts from 7 people6d active+11 post in the last 7 days, 1 the 7 days before (steady)

Posts per day

Posts per day7 posts, Aug 27 to Sep 25

The posts behind it

7, newest first
PostDate
Epyc for inference...at some point plus enough RAM to overflow bigger models than I can fit in VRAM.. I’m struggling to find clear information on whether this would actually be worth it for the cost.. I am currently able to run Qwen3.8-flash-next in two rather painful configurations (not using the...r/LocalLLaMAu/mrgreatheartSep 24yesterday
Need Kimi K3 model architecture support### Model description hey there, i saw transformers still havent support kimi k3 architecture yet, while llama.cpp for inference already. i have trained new tiny model from scratch using the kimi k3 architecture (0.5B & 1.35B), it also fully compatible with llama.cpp gguf. but...huggingface/transformersimezxSep 169 days ago
support Spark-X2.5-4B-GGUFFailed to load model: This model is not supported yet. Try a different model. (Original error: llama.cpp does not support this GGUF's model architecture ('spark2_5'). The file is valid, but this model type cannot be run with llama-server.)ggml-org/llama.cppkof8855Sep 33 weeks ago
bug: Qwen3.8 DFlash/MTP speculative emits OOB token id == n_vocab (248320) on Vulkan## Summary On AMD Strix Point (`gfx1150`) Vulkan with llama.cpp **b10724**, Qwen3.8-27B + z-lab DFlash2 (or ggml MTP draft) fails decode with: **248320 equals `tokenizer.ggml.tokens` array length** - valid token ids are `0..248319`. This looks like a classic off-by-one in draft...ggml-org/llama.cppmaciejdzialoSep 13 weeks ago
Feature Request: `--n-cpu-mode` FFN band selection for `--n-cpu-ffn`### Prerequisites - [x] I am running the latest code. Mention the version if possible as well. - [x] I carefully followed the [README.md](https://github.com/ggml-org/llama.cpp/blob/master/README.md). - [x] I searched using keywords relevant to my issue to make sure that I am...ggml-org/llama.cppJohn-194Aug 293 weeks ago
qwen4exp: slow decode on Pascal (sm_61) — 18 tok/s vs 29.5 for same-class qwen3.5moe with more active params### Summary On 4× Tesla P40 (sm_61), `qwen4exp` (Qwen3.8-Flash-Next, 125B / 6B active) decodes at **~18 tok/s**, while `qwen3.5moe` (Qwen3.5-122B-A10B, 122B / **10B** active) on the same GPUs and the same build decodes at **~29.5 tok/s**. This is the opposite of what the...ggml-org/llama.cpppodolchakagencyAug 293 weeks ago
qwen4exp (PR #27742): multi-segment prompts degrade to '//////' on gfx1151 — deterministic repro...usage already works great here (thinking mode included, after the chat-template workaround below). ### Environment - Strix Halo laptop APU: Ryzen AI MAX+ 395 + Radeon 8060S (`gfx1151`), ROCm 7.2, kernel 6.19.8 - Tested both #27774 head (`abdc7a0`) and #27742 head (`6c5afc86`) -...ggml-org/llama.cppranxiangleiAug 274 weeks ago

Companies and products named

Company or productPosts naming it
llama.cpp7
NVIDIA2
Hugging Face1
GitHub1
Tesla1
Unsloth1
About this problem

Evidence

7 posts from 7 people in 3 places, about 1 a week over 29 days. Mostly on ggml-org/llama.cpp, r/LocalLLaMA, huggingface/transformers. Tools named alongside: NVIDIA, Hugging Face, GitHub, Tesla.

Frustration Frustration 0 of 3· Seen on GitHub issues, Reddit

How it was grouped

Posts that state a pain and match the "local-llm" rule about llama.cpp. First post Aug 27, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Running llama.cpp locally hits memory and speed limits (7 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.