Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs
Slow Vale, an LLM-driven life simulation, now features a Chinese server with over 800 AI residents in a continuously running city. The developer shared insights on managing concurrent decisions, dynamic action spaces, context caching, and inference costs for this persistent multi-agent system. User retention for the 225 August entrants was 68.9% on Day 1, 56.0% on Day 7, and 40.9% on Day 30. An English browser version is available at slowvale.com, requiring only an email for registration.
This report details the engineering challenges and solutions for running 800+ persistent agents in an LLM-driven simulation, unlike many theoretical discussions.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 8, 2026, 17:00 UTC
- Ingested
- Oct 8, 2026, 17:00
- Source type
- Dev community
Full text isn't available here.
Read at source →