PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)
For agentic workflows with llama.cpp, increasing the -cram parameter beyond its default of 8192 MB can significantly improve speed, especially with large context lengths and multi-turn projects. While this uses more RAM, it prevents the entire context from needing reprocessing on every turn when the context length exceeds the default cache size. For instance, 20480 MB has been effective with Qwen 27B 3.8 at 262K context, without increasing VRAM usage.
This report highlights a specific optimization for llama.cpp agentic workflows by increasing -cram beyond its default, unlike general advice that often overlooks this parameter's impact on large contexts.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 25, 2026, 03:01 UTC
- Ingested
- Sep 25, 2026, 03:01
- Source type
- Dev community
Full text isn't available here.
Read at source →