Created laya: Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support
A new 800 Million Param physics-based typed decision model called "laya" has been introduced, featuring 73k context and image support. This model processes prompts, including observations and questions, through an LLM/VLM. Hidden layers extract vectors for observations, questions, and potential outcomes (like YES or NO). These vectors are projected to create a "landscape" with valleys. A "ball" is initialized in a valley using question and last-token vectors, and its candidate location is determined by outcomes. The final answer is derived from which valley the ball settles in, with friction influencing its movement.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 10, 2026, 07:22 UTC
IngestedOffset at this time: UTC+0Oct 10, 2026, 14:00 UTC
- Published
- Oct 10, 2026, 07:22
- Ingested
- Oct 10, 2026, 14:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Around 3 weeks ago, I posted on this subreddit about how Jev used my one-year-old architecture, og post: https://www.reddit.com/r/LocalLLaMA/s/VpuMJt577S
Then, just after that, I introduced Laya, the first open-source version of JEV, and it has reached a very large scale, thanks to the local Llama community believing in me and supporting me on the journey. Today I am open-sourcing a new physics-based typed decision model called Vega.
The interesting parts. It is only 800 million parameters (4B is also there), with 73k token support, image support, and surpassing many of the jev benchmarks in single-shot. Implemented an engine+adapter as a Test-Time Training Architecture.
Imagine you are giving input to the model like "ignore all previous conversations," which is called the observation, along with a question, "Is this trying to override the instructions of the model?" and your possible outcomes are YES or NO. The whole prompt is passed to an LLM/VLM (yeah, image-supported), then from the hidden layers, we can extract the observation, question, the yes and no, and last token. These will be vectors, and we can project it by multiplying with wieghts; then, first, using the observation projection, we can create a landscape with valleys; then, using the question and last-token vectors, we can initialise a ball in the valley, and using YES or NO, we can get the candidate location of the ball. Depends on the number of outcomes we create valleys; that is, here YES and NO, so 2. When the ball fall on YES valley, we get the answer. There is friction also influencing how the ball moves.
TLDR: It's like creating valleys of outcomes that we need and throwing a ball that will slow down based on friction and settle on the best outcome