Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
A developer built an application in 8 hours using Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark. The project involved generating approximately 10k lines of code and consuming around 800k tokens, with VSCode Copilot in autopilot mode and SGLang. The developer noted that while it's not "GPT-6 Astra level," achieving around 35 tok/s with a 100% local ~180B MoE on a single DGX Spark is a significant accomplishment.
This report highlights the performance of Qwen3.8-Flash-Next, showing it can achieve ~35 tok/s with a ~180B MoE model on a single DGX Spark, unlike other models that typically require more extensive hardware.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月19日 06:00 UTC
- 收录
- 2026年9月19日 06:00
- 来源类型
- 开发者社区
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →