Back
RCreddit.com
15
·17 hr ago·Dev community · RSS

"Vision" is the current bottleneck imo

View original
Video generation

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

This covers a coding tool or code-capability update — useful for developers assessing workflow changes and reusable value.

AI summary

The current bottleneck in AI development, particularly for tasks involving visual elements, is "Vision." While AI can perform well in some areas, it struggles with visual assessmen…

Blind people are obviously just as capable in some areas, but ask them to make a Minecraft clone (or an original video game with a focus on visual elements) and they'll really fucking struggle, not with the coding, but with the loop of assessing what they've done so far.

You just can't unit test everything in advance reliably, sometimes you have to load in the game and notice the sword is rotated incorrectly in the player's hand.

With one person with vision on their side they can do it, and that's basically the pairing a lot of vibe coders amount to, you are the set of eyes for the LLM, still a small amount of providing "common sense", taste, opinions, goal alignment, etc. but in terms of technical capability it's obviously very far along and imo not the bottleneck.

Memory is an issue too. "Dexterity" is ofc an issue with those trying to make robot irl workers. There are a whole host of bottlenecks but to me vision feels like a big one, to instantly notice issues, yeah if you ask ChatGPT what's wrong if you show it a picture of a player with a sword held blade first it'd probably notice, but can it notice it unprompted for a video? Could it notice if the sword was only rotated incorrectly in the Z axis such that the flat side was being "used" for swings?

Needs to be more efficient and more intuitive, rather than relying on the reasoning power they've cultivated imo.

Especially for video, norm for video is to give 1-5 images per second depending on the LLM, price, plan, company, etc. but that's not how people see videos. And it's not even how we see images, tokenisation of text probably isn't TOO far from how the brain does it, but for images it's probably WAY off.

"Vision" is the current bottleneck imo · BuzzRadr