Back
RCreddit.com
12
·16 hr ago·Dev community · RSS

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

View original
Hugging Face

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen and Gemma models in terms of performance.

Should have a massive tokens/sec on most systems. I quite like tiny MoE's conceptually.

Edit: looks like the model card actually reports speeds:

With FP8, Ling-3.0-tiny reaches around 100-105 tokens/s on DGX Spark and 86-90 tokens/s on an M4 Pro MacBook, with approximately 8.34 GiB peak memory usage at an 8K context length.

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face · BuzzRadr