Skip to content
RCreddit.com·
Not on the current live radar

Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark

AI summary

A developer built an application in 8 hours using Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark. The project involved generating approximately 10k lines of code and consuming around 800k tokens, with VSCode Copilot in autopilot mode and SGLang. The developer noted that while it's not "GPT-6 Astra level," achieving around 35 tok/s with a 100% local ~180B MoE on a single DGX Spark is a significant accomplishment.

Why this one

This report highlights the performance of Qwen3.8-Flash-Next, showing it can achieve ~35 tok/s with a ~180B MoE model on a single DGX Spark, unlike other models that typically require more extensive hardware.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 19, 2026, 06:00 UTC

Ingested
Sep 19, 2026, 06:00
Source type
Dev community

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com