The expensive part of AI video isn't generating it. It's regenerating it.
Building a product on video generation models revealed that the primary cost and frustration stem from regenerating, not initial generation. A capacity incident with an upstream model provider caused timeouts, highlighting the need to design for model slowdowns to prevent product outages. The product, Melodious.ai, converts songs into music videos, offering a free first storyboard. The developer is seeking insights on handling retry problems, such as capping spend, previewing, or accepting redo rates.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 09:28 UTC
收录当时偏移:UTC+02026年10月9日 14:00 UTC
- 发布
- 2026年10月9日 09:28
- 收录
- 2026年10月9日 14:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
I have spent the last months building a product on top of video generation models, and the thing that surprised me most is where the money and the frustration actually go.
When people think about the cost of AI video they think about the price of a clip. In practice the price of one clip barely matters. The cost is the loop: you generate, it is not what you meant, you adjust the prompt, you generate again. A thirty second video with six or eight shots can easily turn into forty generations, because each shot can fail independently and the failures are not obvious until you have paid for them.
Three things I learned that apply beyond my own product:
- Cheap previews beat clever prompts. Generating still images first costs a small fraction of generating motion, and a still tells you most of what you need to know: does the character look right, is the colour world consistent, is the composition what you pictured. If the stills are wrong, no amount of motion will fix them. Move the failures to the cheap stage.
- The planning model matters more than the video model. We tried reducing how much the planning step "thinks" to make things faster. In our tests the plans got thinner and some failed outright. The step that decides what the shots are is the step that decides whether the video is any good, so we stopped trying to economise there.
- Consistency is a product problem, not just a model problem. The model has no memory of what a character looked like two shots ago unless you give it something to hold onto. Reference images and a written brief that travels with every shot do more than any prompt trick.
I also learned the unglamorous version of this last week: when the upstream model provider had a capacity incident, one of our steps took more than a minute against a gateway that gives up at 29 seconds. Users saw a timeout even though the work finished in the background. Nobody warns you that "the model is slow today" turns into a product outage unless you have designed for it.
The product I am building is called Melodious. It takes a song and turns it into a music video, and the first storyboard is free so you can see the plan before you spend anything. The link is melodious.ai if you want to look, but I would honestly be more interested in how others here handle the retry problem. Do you cap spend per shot, preview first, or just accept the redo rate?