Learning to Learn a Language: in-context learning of natural language from a synthetic non-linguistic prior [R]
A new paper, "Learning to Learn a Language," extends the concept of prior-fitted networks to structured sequences like natural language. The researchers propose a prior over languages, where each training sequence originates from a randomly sampled recurrent causal model, creating a new synthetic "language." A 300M-parameter byte-level transformer, trained solely on these synthetic sequences, demonstrates in-context learning of real languages. When given Wikipedia text with frozen weights, its next-byte predictions improve significantly across six tested languages (English, Chinese, Hindi, Arabic, Japanese, Korean), reducing from 8 bits per byte to 0.9–2.4 after a million bytes.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月6日 14:00 UTC
- 收录
- 2026年10月6日 14:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →