Skip to content
RCreddit.com·

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

AI summary

A new method, CO2Jump, for concurrent image understanding and generation, utilizes Self-Correcting Coupled Markov Jump Processes. Evaluated on image editing, maze solving, and nonograms, it introduces datasets like JEdit-1M, JMaze-200K, and JNono-200K. CO2Jump demonstrated monotonic improvement in both editing quality and grounding across 8–512 sampling steps, outperforming other samplers in tasks requiring joint accuracy of textual answers and generated images.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 30, 2026, 07:28 UTC

IngestedOffset at this time: UTC+0Sep 30, 2026, 14:00 UTC

Published
Sep 30, 2026, 07:28
Ingested
Sep 30, 2026, 14:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University.

We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent.

Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allows low-confidence tokens to be masked again and regenerated, so earlier decisions can be revised as generation progresses.

CO₂Jump uses one model forward pass per denoising step. The sampler itself requires no additional training; our experiments compare sampling methods using the same task-specific fine-tuned model.

We evaluate image editing, maze solving and nonograms, and introduce three datasets: JEdit-1M, JMaze-200K and JNono-200K. On the puzzle benchmarks, joint accuracy requires both the textual answer and generated image to be correct. Across 8–512 sampling steps, CO₂Jump was the only sampler we compared that improved monotonically on both editing quality and grounding.

I’d be interested in suggestions for other tasks where text–image consistency and correctness can be evaluated together. Happy to discuss the method, evaluation or limitations.

Source·reddit.com