Running DeepSeek locally hits memory and speed limits
MLOps2 posts from 2 people2d active+11 post in the last 7 days, 1 the 7 days before (steady)Unverified
Posts per day
The posts behind it
2, newest first| Post | Date | |
|---|---|---|
| Feature Request: Add support for K2 Horizon (0.9B, 3.7B, 7B, 32B, 36B MoVA)### Prerequisites - [x] I am running the latest code. Mention the version if possible as well. - [x] I carefully followed the [README.md](https://github.com/ggml-org/llama.cpp/blob/master/README.md). - [x] I searched using keywords relevant to my issue to make sure that I am...ggml-org/llama.cppbitalov | ggml-org/llama.cppbitalov2 reactions, 2 replies | Sep 25today |
| Add DeepSeek-V4.1-Flash (deepseek_v41)### Model description **DeepSeek-V4.1-Flash** (released by DeepSeek, [`deepseek-ai/DeepSeek-V4.1-Flash`](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash), ~510 GB, fp8 / packed-fp4 weights). Text backbone of the V4 line with three additions over `deepseek_v4`: - **CSA2 -...huggingface/transformersmalaiwah | huggingface/transformersmalaiwah5 replies | Sep 1312 days ago |
Companies and products named
About this problem
Evidence
2 posts from 2 people in 2 places, about 1 a week over 13 days. Mostly on huggingface/transformers, ggml-org/llama.cpp. Tools named alongside: Hugging Face, Llama, llama.cpp, NVIDIA.
Frustration Frustration 0 of 3· Seen on GitHub issues
How it was grouped
Posts that state a pain and match the "local-llm" rule about DeepSeek. First post Sep 13, 2026, latest Sep 25, 2026. Unverified: fewer than 3 posts, or one person or place only. Method
The same problem elsewhere
- Running Ollama locally hits memory and speed limits (7)
- Running llama.cpp locally hits memory and speed limits (7)
- Running LLMs locally hits memory and speed limits (5)
- Running Qwen locally hits memory and speed limits (4)
- Running vLLM locally hits memory and speed limits (3)
- Running OpenAI locally hits memory and speed limits (2)
Other DeepSeek problems
History
- 2026-09-25 Added: Running DeepSeek locally hits memory and speed limits (2 posts)