Back
RCreddit.com
14
·13 hr ago·RSS
Not on the current live radar

Android Studios native Gemma 4 runs on llama.cpp

View original
LlamaModel release

Heat trend

New
Latest 24h versus previous 24h · 7-day curve

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

Llama model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.

AI summary

Android Studios' native Gemma 4, running on llama.cpp, is speculated to be the Vulkan and QAT versions. It supports multi-GPU configurations and the 31B model boasts a maximum context length of 128k. When fully loaded, Gemma 4 utilizes 34 GB of VRAM. Currently, there are no visible options to adjust the context length or display PP/TG speed within the application.