跳到正文
RCreddit.com·

Created laya: Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support

AI 摘要

A new 800 Million Param physics-based typed decision model called "laya" has been introduced, featuring 73k context and image support. This model processes prompts, including observations and questions, through an LLM/VLM. Hidden layers extract vectors for observations, questions, and potential outcomes (like YES or NO). These vectors are projected to create a "landscape" with valleys. A "ball" is initialized in a valley using question and last-token vectors, and its candidate location is determined by outcomes. The final answer is derived from which valley the ball settles in, with friction influencing its movement.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月10日 07:22 UTC

收录当时偏移:UTC+02026年10月10日 14:00 UTC

发布
2026年10月10日 07:22
收录
2026年10月10日 14:00
来源类型
开发者社区
档位
社区
信源状态
同步延迟

档位是按信源手工设定的编辑判断,不是逐条打分。

Around 3 weeks ago, I posted on this subreddit about how Jev used my one-year-old architecture, og post: https://www.reddit.com/r/LocalLLaMA/s/VpuMJt577S

Then, just after that, I introduced Laya, the first open-source version of JEV, and it has reached a very large scale, thanks to the local Llama community believing in me and supporting me on the journey. Today I am open-sourcing a new physics-based typed decision model called Vega.

The interesting parts. It is only 800 million parameters (4B is also there), with 73k token support, image support, and surpassing many of the jev benchmarks in single-shot. Implemented an engine+adapter as a Test-Time Training Architecture.

Imagine you are giving input to the model like "ignore all previous conversations," which is called the observation, along with a question, "Is this trying to override the instructions of the model?" and your possible outcomes are YES or NO. The whole prompt is passed to an LLM/VLM (yeah, image-supported), then from the hidden layers, we can extract the observation, question, the yes and no, and last token. These will be vectors, and we can project it by multiplying with wieghts; then, first, using the observation projection, we can create a landscape with valleys; then, using the question and last-token vectors, we can initialise a ball in the valley, and using YES or NO, we can get the candidate location of the ball. Depends on the number of outcomes we create valleys; that is, here YES and NO, so 2. When the ball fall on YES valley, we get the answer. There is friction also influencing how the ball moves.

TLDR: It's like creating valleys of outcomes that we need and throwing a ball that will slow down based on friction and settle on the best outcome

来源·reddit.com