Can you use autoregressive diffusion to generate market data?
Kavish's initial diffusion head, designed for market data generation, used a 1,000-step cosine schedule and DDIM for inference. This setup proved unstable, leading to exploding denoising trajectories where 88-95% of values exceeded 8 standard deviations from the mean across 7 continuous targets. The model's predictions also degraded significantly further into the future, as seen in the widening spread in CME US synthetic rollouts, indicating areas for further research.
This report details the instability of an initial diffusion model configuration, with 88-95% of values exceeding 8 standard deviations, unlike typical stable models.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 14:56 UTC
收录当时偏移:UTC+02026年10月10日 02:00 UTC
- 发布
- 2026年10月9日 14:56
- 收录
- 2026年10月10日 02:00
- 来源类型
- 未分类
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition”
In quantitative finance we are used to models that take a stream of market data events for a given symbol (like resting orders being added to an order book, cancellations of those orders, executions) and predict that symbol’s future price. Generative models are less common. Imagine a model that could give not just a point estimate of a symbol’s price, but could actually synthesize book events—including the timing of their arrival on the exchange. You’d get price predictions, of course, but your rollouts would have hugely more texture than just that.
But what kind of data is market data? A generative model demands an answer to that question. Is it more like video, or text? That is, is it continuous or discrete? Market data seems to have features of both. An order book evolves via a series of turns as market participants put on, or take away, resting orders; there is no doubt a discrete action space. And yet, many of the most important parameters of a given order, like its price, have such a high cardinality that they basically appear continuous. Some simply are continuous: an easy one to overlook, but crucial for a generative model, is timing, as in “when does the new order arrive at the exchange.” Even that oversimplifies things. In practice the distributions are spiky: “pennying” (where you improve upon a price by one tick) is much more common than improving by two ticks, and orders tend to arrive in bursts, for instance around whole-number times.
One way to explore these complications is to treat your data as if it’s continuous and see how that breaks down. This summer, a research intern, Kavish, built a diffusion model for market data. Diffusion and flow-matching models are powerful tools for representing sequential, multimodal data in regimes such as image, video, robotics, and audio. Taking inspiration from Autoregressive Image Generation without Vector Quantization, Kavish created an event-level generative model of market data using autoregressive diffusion.
There were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better. But the headline result was that fully continuous diffusion did not lend itself to the jaggedness of real market data. Kavish’s various attempts to handle point masses via smoothing, on the one hand, and discretizing some of the model’s targets, on the other, have greatly clarified what an eventual generative model of market data might look like.
Modeling and setup
Kavish looked at four years of US equities data, with each row of data containing timestamp, price, kind (trade, order, cancel, etc.), and other information. The diffusion model would try to generate these features for the next event, which was then autoregressively appended to the stream.
Following the Li et al. paper cited above, he used an encoder–diffuser architecture, specifically a causally masked transformer encoder to produce a latent embedding at each position. During training, the latent at each position was fed into a small event-kind head, which outputs a 2-categorical probability distribution indicating whether the next event is a trade or a BBO update (“best bid and offer”). The continuous targets, like the price, elapsed time, etc., are generated via diffusion head that is conditioned on the corresponding latent and the true event-kind via standard AdaLN conditioning.
During inference, you pass only the encoded latent of the last event into the event-kind head and sample an event from the output distribution. The next event’s continuous features are generated via the diffusion head, conditioned on the last event’s latent and the sampled event.
DDPM vs. Flow matching
A DDPM or “denoising diffusion probabilistic model” predicts the noise that was added to a sample, and derives a clean estimate by subtracting the scaled noise prediction and dividing by the remaining signal level. DDPMs are highly successful in areas such as image generation, but at high noise levels, small errors in the noise prediction result in large errors in the clean target estimate. Flow matching instead interpolates linearly between noise and data, and trains the network to predict the velocity along that line (data minus noise). As a result, the sampled trajectories are nearly straight, and flow models can be integrated accurately in relatively few steps.
Fundamentally these are two different parameterizations of the same objective, but Kavish found that the differences mattered quite a bit. His initial diffusion head predicted on a 1,000 step cosine schedule, as recommended in Improved DDPM. At inference time, generation was run following DDIM. Unfortunately, this configuration was unstable, and led to exploding denoising trajectories, with 88-95% of values over 8 standard deviations away from the distribution mean on all 7 continuous targets. Note in the figure how in the leftmost box, the trajectories (those thin pink filaments, just barely visible against the background) quickly diverge to the edges.
This is caused by the scaling of when calculating. As the figure shows, using more sampling steps reduces the increment per step, which prevents the amount of overflow. Additionally, adding stochasticity to the sampling process () regularizes the denoising process to a standard gaussian, reducing deviations. But while it is possible to tune the noise schedule to account for this instability, or apply inference-time interventions to prevent denoising explosion, Kavish found that a rectified flows approach performed well out-of-the-box. So he switched to a flow-matching model.
Handling discontinuities
Most of the project was spent grappling with a fundamental problem: market data is neither fully continuous nor fully discrete. Video, for instance—for which diffusion models are well suited—is continuous both in snapshots (each frame is a series of continuous pixel intensities and colors) and in evolution (pixels can take on any continuous value from frame to frame). Market data, by contrast, has continuous snapshots but evolves in a very discrete way, governed by market microstructure.
This is evident when you look at the distributions of features, which are characterized by sharp boundaries:
The timing of orders isn’t normally distributed, but rather quite clumpy, with point masses around zero seconds (events that arrive simultaneously), the-earliest-possible-reaction-time, and whole-number seconds. And prices tend to cluster around the midpoint of the bid and ask, or the current bid, or the current ask, and thus you see sharp discontinuities there too.
Because there’s also an imbalance in categorical predictions—only 8% of events are trades, the rest being BBO changes; thus the categorical head overpredicts trades—Kavish experimented with fracturing the number of classes predicted by the categorical head. He broke BBO changes into a bunch of common classes, like “size only” (no change in price) or “ask up”, “bid down”, etc. This had the added benefit of allowing him to factor out “atoms” (or sharp discontinuities) in certain features, so that an elapsed time of zero could be predicted as its own class. He ended up with a 20-class categorical head, and simpler diffusion targets, like “how big was the non-zero time gap” (since zeroes are accounted for categorically) and “what’s the magnitude of the bid price change” (since direction is a category).
This approach led to materially better marginal distributions for both the event-kind and continuous targets, but obviously involved lots of hand-engineering. Using a categorical head to specially model every discontinuity in every feature becomes unscalable in larger feature sets. For instance, real data is likely to have far more variables, including high-cardinality ones like “which exchange was this on?”, potentially exploding the number of categorical classes.
So Kavish went looking for another way to handle discontinuities.
Atom smoothing
He used a procedure he ended up calling “atom smoothing.” The idea is to take the true, spiky distribution, smooth it out, and then re-sculpt in the spike, in a way that captures the probability mass there without creating a discontinuity. For example:
For this to work well, you have to be sure to pick the right representation of your variables: sharp spikes under some representation will be less sharp under another, and you need to expose and smooth any sharp spikes.
Results and rollouts