RCreddit.com
16
·1 days ago·RSS
Not on the current live radar
snkii/Sori-1B: Audio-Grounded LM Trained From Scratch (No Text-Only Pretraining)
Model release
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Sori-1B is a 1B-parameter audio-language model developed by a single SNU researcher, notable for its decoder being trained entirely from scratch on audio-paired text, without text-only pretraining or pretrained-LM initialization. This approach aims to ground its answers in audio, unlike typical AF3-style models that rely on text-only priors. Sori-1B reuses NVIDIA’s frozen Audio Flamingo Next encoder and includes a custom “auditory-ontology” tokenizer. It supports MCQ, open QA, captioning, and ASR modes, with weights gated under a non-commercial/academic-only license.