RCreddit.com
16
·1天前·RSS
暂不在当前实时榜单
snkii/Sori-1B: Audio-Grounded LM Trained From Scratch (No Text-Only Pretraining)
模型发布
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
Sori-1B is a 1B-parameter audio-language model developed by a single SNU researcher, notable for its decoder being trained entirely from scratch on audio-paired text, without text-only pretraining or pretrained-LM initialization. This approach aims to ground its answers in audio, unlike typical AF3-style models that rely on text-only priors. Sori-1B reuses NVIDIA’s frozen Audio Flamingo Next encoder and includes a custom “auditory-ontology” tokenizer. It supports MCQ, open QA, captioning, and ASR modes, with weights gated under a non-commercial/academic-only license.