Skip to content

Running Ollama locally hits memory and speed limits

MLOps7 posts from 7 people7d active+22 posts in the last 7 days, 2 the 7 days before (steady)

Posts per day

Posts per day7 posts, Aug 28 to Sep 25

The posts behind it

7, newest first
PostDate
Feature request: yield idle model VRAM under GPU-memory pressure## Summary Please add an optional, best-effort policy for an idle Ollama model to release its VRAM early when another GPU application needs memory. ## Why The keep-alive timer is useful, but it does not account for a game, video editor, renderer, or other GPU-heavy application...ollama/ollamaGrymgrinSep 232 days ago
0xc0000005 access violation loading ANY model on Vulkan (AMD RX 6800 XT, driver 32.0.21045.5002)## Summary Any model (tested with both a plain-transformer 7B and a hybrid SSM+transformer 27B) crashes `llama-server` with `exit status 0xc0000005` (access violation) when loaded on the Vulkan GPU backend. CPU inference works fine. This looks like the same underlying issue as...ollama/ollamaValarAULESep 205 days ago
Feature Request: Support Prism PQ2_0 (GGML type 142) and PTQ1_0 (type 143) used by Ternary-Bonsai-2### Prerequisites - [x] I am running the latest code. Mention the version if possible as well. - [x] I carefully followed the [README.md](https://github.com/ggml-org/llama.cpp/blob/master/README.md). - [x] I searched using keywords relevant to my issue to make sure that I am...ggml-org/llama.cppmikestaubSep 187 days ago
Add MLX support for Bonsai's low-bit (1-bit/2-bit) quantized weights to the new 0.19 MLX backend.Add MLX support for Bonsai's low-bit (1-bit/2-bit) quantized weights to the new 0.19 MLX backend. The 2-bit ternary format already runs on stock MLX, so this may only require letting ollama pull fetch arbitrary MLX repos rather than registry-only models. The 1-bit format needs...ollama/ollama2tqkfqv2yy-cmykSep 178 days ago
Multi-GPU VRAM accounting uses discovery device names instead of child's log names### What is the issue? ## Summary On a multi-GPU system, the scheduler's VRAM accounting diverges from actual because three lookup sites read device names assigned by discovery (e.g. `CUDA1`) while the maps they read from are keyed by the child llama-server's log names (e.g....ollama/ollamaAlimar777Sep 92 weeks ago
Does Ollama support IQ3_S quantization for Qwen3.8-27B-GSQ-RCO-GGUF? Returns empty content## Question Does Ollama currently support running the `IQ3_S` quantization of `Qwen3.8-27B-GSQ-RCO-GGUF` from hf.co? ## Problem When running `hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_S` via Ollama: - Generation completes (`done_reason: "stop"`) - `content` field is always...ollama/ollamaSaber120Sep 72 weeks ago
llama-server malloc heap grows with request volume on macOS/Metal: 6.5 GB paged to swap while KV cache stays resident (0.32.15, Apple Silicon)...the pipeline idle but the model still loaded, swapouts were **exactly 0** over 90 s. ### Workaround Evicting the model (`{"keep_alive": 0}`) between units of work terminates that `llama-server` process, and the heap goes with it. With eviction every source, same 8192 context and...ollama/ollamaYassfive20Aug 284 weeks ago

Companies and products named

Company or productPosts naming it
Ollama7
NVIDIA2
GitHub2
LM Studio1
Docker1
llama.cpp1
About this problem

Evidence

7 posts from 7 people in 2 places, about 2 a week over 27 days. Mostly on ollama/ollama, ggml-org/llama.cpp. Tools named alongside: NVIDIA, GitHub, LM Studio, Docker.

Frustration Frustration 1 of 3· Seen on GitHub issues

How it was grouped

Posts that state a pain and match the "local-llm" rule about Ollama. First post Aug 28, 2026, latest Sep 23, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Running Ollama locally hits memory and speed limits (7 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.