Skip to content
RCreddit.com·
Not on the current live radar

PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)

AI summary

For agentic workflows with llama.cpp, increasing the -cram parameter beyond its default of 8192 MB can significantly improve speed, especially with large context lengths and multi-turn projects. While this uses more RAM, it prevents the entire context from needing reprocessing on every turn when the context length exceeds the default cache size. For instance, 20480 MB has been effective with Qwen 27B 3.8 at 262K context, without increasing VRAM usage.

Why this one

This report highlights a specific optimization for llama.cpp agentic workflows by increasing -cram beyond its default, unlike general advice that often overlooks this parameter's impact on large contexts.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 25, 2026, 03:01 UTC

Ingested
Sep 25, 2026, 03:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com