跳到正文
RCreddit.com·

Have URMs and UTs been integrated into frontier models? Or did they disappear into the dustbin of forgotten papers? [D]

AI 摘要

UT-based small models have shown superior performance compared to most standard Transformer-based Large Language Models (LLMs), even when trained from scratch without internet-scale pre-training. This suggests that the integration of such models, potentially utilizing architectures like "H t+1 ← LayerNorm( H t+1 + Transition(H t+1 )), t = 0,..., T − 1", could be beneficial for frontier models. The question remains whether these Universal Representations and Transformers (URMs and UTs) have been adopted or overlooked.

为什么是这条

This report highlights that UT-based small models consistently outperform most standard Transformer-based LLMs, unlike the common assumption that larger, pre-trained models are always superior.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月8日 20:28 UTC

收录当时偏移:UTC+02026年10月9日 01:00 UTC

发布
2026年10月8日 20:28
收录
2026年10月9日 01:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

暂无对比
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

UT-based small models, despite being trained from scratch on these tasks without internet-scale pre-training, consistently outperform most standard Transformer-based Large Language models (LLMs) by a significant margin. 👈

The Universal Transformer (UT) extends the standard Transformer by introducing recurrent computation over depth. Instead of stacking L distinct layers, the UT applies a single transition block repeatedly to refine token representations. For an input sequence x with embedding matrix H 0 ∈ R n×d, the UT updates states as

H t+1 = LayerNorm(H t + MHA(H t ))

followed by a shared position-wise transition function

H t+1 ← LayerNorm( H t+1 + Transition(H t+1 )), t = 0,..., T − 1

where Transition is either a feed-forward network or separable convolution. To encode both position and refinement depth, UT adds 2-D sinusoidal embeddings at each step.

https://i.imgur.com/S8TAh3f.png

(below) Figure 2: Illustration of our Universal Reasoning Model (URM) architecture. The left shows a standard Transformer layer stack, while the right illustrates the URM with fixed loops, ACT loops, and the ConvSwiGLU module. For illustrative purposes, components such as embeddings, residual connections, RMSNorm, positional encodings, and other modules are omitted, x in right figure represents the first x loops of the inner loop in forward-only mode, TBPTL represents our proposed Truncated Backpropagation Through Loops.

https://i.imgur.com/uEuhAeC.png

Have frontier labs at big tech companies already implemented the enhancements of URMs and UTs? Or is this cutting-edge research still waiting for its day in the sun?

Below are the original paper, a simplified blog, and a talk.

Paper: https://arxiv.org/pdf/2512.14693

blog: https://bdtechtalks.substack.com/p/inside-urm-the-architecture-beating

youtube talk: https://www.youtube.com/watch?v=RxNPFCYCFBU

来源·reddit.com