← All writing

PulseCheck: How We Validate Our Market Simulators

In previous posts, we introduced Pulse, our agent-based market simulator, and Portobello, its latest release. Every improvement we reported was measured against the same validation framework. In this post we introduce that framework: PulseCheck.

PulseCheck is our limit order book validation suite. It is designed to answer a simple question: does a simulated market behave like the real one? To answer that properly, looking at the price path alone is not enough. We also need to examine the structure of the order book, how liquidity is distributed across price levels, how orders arrive and are cancelled over time, and how the market responds to different types of events. Together, these features tell us whether a simulation reproduces the underlying market dynamics rather than simply generating a plausible price series.

A realistic price chart is not enough

Put a simulated price path next to a historical one and they may look remarkably similar. That tells us less than it might seem. A simulator can produce a convincing price series while getting the market underneath it wrong. Spreads might be too tight, liquidity might sit at the wrong depths, orders might cancel too slowly, or the book might recover unrealistically quickly after a trade. These differences matter when the simulation is used for execution analysis, stress testing or strategy research. The price is the visible output of a much richer market process. For Pulse to be useful in these settings, the behaviour of the underlying book has to be realistic too.

That is what PulseCheck is designed to measure.

Given historical limit order book data for a symbol and a set of Monte Carlo simulations of the same market, PulseCheck produces a quantitative view of where the simulated market matches the historical one and where it does not. PulseCheck combines distributional metrics, impact-response analysis and stylised facts into a single validation framework. The first two components draw on LOB-Bench (Nagy et al., 2025), while the stylised facts are adapted from Cont (Cont, 2001) and provide an additional check on the behaviour of the resulting price process.

What does PulseCheck measure?

The simulator is not expected to reproduce the exact historical sequence of events. Rather, it should generate a plausible alternative realisation of the same market, so validation focuses on matching distributions rather than individual events.

PulseCheck evaluates 21 distributions covering three parts of the market. Some of the metrics describe the shape of the book at a point in time, some describe how the book evolves, and others describe the behaviour of individual orders.

Book shape

The first group looks at the state of the order book itself.

MetricWhat it captures
spreadDistance between the best bid and best ask
touch_bid_volume / touch_ask_volumeVolume resting at the best bid / ask
total_bid_volume / total_ask_volumeVolume across the top 10 levels on each side
obiOrder book imbalance at the touch

These metrics tell us whether the simulator produces realistic spreads, liquidity and imbalance. A price series can look convincing while the book supporting it is systematically too liquid, too thin or too imbalanced.

Order flow

The second group looks at how the market evolves through time.

MetricWhat it captures
log_interarrivaltimeLog time between consecutive events
volume_per_minuteTraded volume per minute, inferred from book changes
ofi_rollingOrder flow imbalance averaged over 100 events
ofi_rolling_up / down / stayOFI conditioned on the direction of the next mid-price move

The conditional order flow imbalance metrics are particularly useful. A simulator may reproduce the overall amount of imbalance in the market while still getting its relationship with subsequent price movement wrong.

In that case the distribution looks right, but the structure underneath it is not. Conditioning OFI on the direction of the next mid-price move gives us a way to detect this.

Order behaviour

The final group looks at individual orders. These metrics require message-level data rather than book snapshots alone.

MetricWhat it captures
time_to_cancelLifetime of an order from placement to cancellation
ask_limit_depth / bid_limit_depthDistance from the mid-price at which new limit orders are placed
ask_cancel_depth / bid_cancel_depthDistance from the mid-price of cancelled orders
ask_limit_level / bid_limit_levelBook level (1 to 10) at which new limit orders arrive
ask_cancel_level / bid_cancel_levelBook level from which cancellations occur

Together, these metrics give us a fairly detailed view of the simulated market. We can ask whether the book has the right shape, whether it evolves at the right rate, and whether the orders inside it behave in a way that resembles the historical market.

How do we score a match?

For each metric, PulseCheck produces a distribution from the historical market and another from the simulation. The next problem is turning the difference between those distributions into something we can compare.

We use two distances: L1 and Wasserstein-1. They capture different kinds of error, so we keep both.

L1 distance

The L1 distance compares the historical and simulated distributions as histograms.

We pool the observations, construct a common set of bins using the Freedman-Diaconis algorithm, and calculate the proportion of each dataset falling into each bin. We then compute

L1(p,q)=12∑i=1B∣pi−qi∣L_1(p, q) = \frac{1}{2} \sum_{i=1}^{B} \left| p_i - q_i \right|,

Here, pip_i is the proportion of historical observations in bin ii, qiq_i is the proportion of simulated observations, and BB is the number of bins.

The factor of 12\tfrac{1}{2} means the score lies between 0 and 1. A score of 0 means the two histograms are identical, while a score of 1 means they contain no overlapping probability mass.

This bounded scale is convenient because every metric can be displayed on the same axis. It also makes the score relatively easy to interpret.

There is one important limitation. L1 tells us how much probability mass is in the wrong place, but not how far away that place is.

Suppose the historical spread is usually two ticks. A simulator that puts most of its mass at three ticks is wrong, but only slightly. Another simulator might put the same amount of probability at ten ticks. L1 can assign similar penalties to both because it only sees that the mass has moved to a different bin.

For that reason, we also look at Wasserstein distance.

Wasserstein distance

Wasserstein-1 is often described as the earth mover’s distance. The intuition is to imagine one distribution as a pile of earth and ask how much work is required to move it until it has the shape of the other distribution.

In one dimension,

W1(P,Q)=∫−∞∞∣FP(x)−FQ(x)∣ dxW_1(P, Q) = \int_{-\infty}^{\infty} \left| F_P(x) - F_Q(x) \right| \, dx,

where FPF_P and FQF_Q are the empirical cumulative distribution functions.

Unlike L1, Wasserstein distance takes the size of the error into account. Moving probability mass one tick costs less than moving it ten ticks. This makes it useful for detecting systematic shifts and errors in the tails of a distribution.

There is a practical complication. Wasserstein distance inherits the units of the underlying variable. A distance measured on volume might be expressed in thousands of shares, while a distance measured on spread might be expressed in ticks. The raw values therefore cannot be compared directly.

PulseCheck handles this by standardising both the historical and simulated values using the mean and standard deviation of the historical distribution:

W1z(P,Q)=W1 ⁣(P−μPσP,  Q−μPσP)W_1^{z}(P, Q) = W_1\!\left( \frac{P - \mu_P}{\sigma_P}, \; \frac{Q - \mu_P}{\sigma_P} \right),

The important point is that both distributions use the same historical normalisation. We do not independently standardise the simulation. If the simulated distribution is shifted or has the wrong scale, that error remains visible. It is simply expressed in units of historical standard deviations.

Why use both?

L1 and Wasserstein tell us slightly different things about the same comparison.

L1 is useful for understanding how much the two distributions overlap. Wasserstein adds information about how large the displacement is. A simulator can therefore score relatively well under one distance and poorly under the other.

For example, a high L1 score with a low Wasserstein score can indicate that the simulation misses some fine structure while remaining close to the right region of the distribution. A relatively low L1 score with a high Wasserstein score can point to a problem in the tails, where most of the distribution overlaps but a smaller amount of mass sits much too far away.

Looking at both helps us understand not just whether a simulation is wrong, but how it is wrong.

Visualising model performance

Twenty-one metrics and two distance measures produce a large number of scores, especially across many Monte Carlo runs, symbols, and trading days. We use spider plots to summarise these results and compare models directly.

Each axis represents one metric and each polygon one simulator. Distances are averaged across runs, and the radial axis is inverted so that zero lies on the outer rim. Better match with the historical market therefore produces a larger web, while weaker performance appears as a dent towards the centre.

The plots are particularly useful for model comparison. Because all simulators are evaluated against the same historical day using the same metrics and distances, their polygons can be overlaid directly. This makes it easy to see where models differ in their strengths and weaknesses, rather than reducing performance to a single score.

The example below compares Pulse ABM and LOBS5 for 700.HK on 22 December 2025. LOBS5 (Nagy et al., 2023) is a deep generative model that learns and autoregressively generates message-level order flow, whereas Pulse builds the market from the bottom up through interacting, instrument-calibrated agents. The two therefore provide contrasting approaches to generating synthetic markets.

PulseABM is shown as the mean across 10 independent Monte Carlo runs, whereas the LOBS5 result corresponds to a single inference run for one symbol and one trading day. Hovering over each point shows the corresponding distance and, where applicable, the standard deviation across runs.

L1 distance by metric for 700.HK on 22 December 2025, comparing Pulse and LOBS5 against the historical trading day. Pulse shows the mean across 2 independent Monte Carlo runs, while LOBS5 represents a single inference run for the same symbol and day. The outer rim corresponds to a distance of 0 and the centre to 1. Hover for exact values; standard deviations are shown for Pulse across runs.

Wasserstein distance by metric for the same comparison, normalised by the historical mean and standard deviation. The outer rim represents a distance of 0. Because the normalised Wasserstein distance is not bounded by 1, the scale is set by the largest distance in the comparison rather than a fixed limit, so no metric is clipped: the worst axis sits at the centre and every other axis is placed relative to it. Hover for exact values.

When reading these charts, distances of zero are not the target. The aim is not for a simulation to reproduce exactly a trading day, but to give a financially plausible alternative.

Impact response

The distributional comparison of different market statistics is a static analysis. However, the market is not static. When a new order comes through, it impacts the book. Thus, we also need a dynamic measure of how the book evolves throughout the day. This is where impact response comes in. The impact response analysis follows LOB-Bench (Nagy et al., 2025), which uses the event-response methodology of Eisler et al. (Eisler et al., 2012) as its basis.

We begin with events that affect the price or quantity at the best bid or ask. These are divided into market orders, limit orders and cancellations. Each category is then split according to whether the event immediately changes the mid-price or leaves it unchanged, giving six event types in total.

This distinction is useful because two events of the same broad type can have very different effects. A market order that consumes all remaining liquidity at the best quote immediately changes the mid-price. Another market order may be absorbed by the available liquidity and leave the price unchanged. Even when there is no immediate price movement, the event can still contain information about what happens next.

For each event type π\pi, we measure the average signed movement in the mid-price after the event:

Rπ(ℓ)=⟨(mt+ℓ−mt)εt⟩t∈πR_\pi(\ell) = \Big\langle \left( m_{t+\ell} - m_t \right) \varepsilon_t \Big\rangle_{t \in \pi}

Here, mtm_t is the mid-price at event time tt and ℓ\ell is the lag, measured in events rather than seconds. The sign εt\varepsilon_t is chosen so that each event’s expected impact on the mid-price is positive.

We calculate the response at log-spaced lags from 1 to roughly 200 subsequent events and bootstrap 99% confidence intervals for the historical data and each simulation. We use event time rather than wall-clock time because the simulated market may operate at a slightly different activity rate from the historical one. Measuring the response in events makes it easier to compare the dynamics of the two books directly.

The response curves can expose behaviour that looks completely normal in the distributional metrics. A simulator might have an accurate spread distribution while replenishing liquidity too quickly after a trade, causing impact to disappear too soon. Another might respond too strongly because liquidity does not react enough.

Stylised facts

The final part of PulseCheck steps away from the book and looks at the price series that emerges from it.

Financial returns have a set of empirical properties that recur across markets, instruments and sample periods. Rama Cont collected them into the list now usually called the stylised facts (Cont, 2001), and a recent study has revisited that list against modern US equity data and found that most of the facts still hold (Ratliff-Crain et al., 2025).

What makes them useful is that they are qualitative statements. They do not describe one particular trading day, they describe how financial price series behave in general. So unlike the distributional metrics and the impact curves, this part of the suite does not need a historical counterpart to compare against. We can ask whether a simulated price process has the properties a market’s price process should have, and answer that from the simulation alone.

PulseCheck implements eleven of Cont’s stylised facts. Returns are mid-price changes, sampled at 1s, 60s, 300s, 900s and 1800s, so each fact is checked at both microstructure and intraday scales. Three of the facts sweep a finer grid instead, from 1s to 500s in 50s steps, because what they test is how a property changes as returns are aggregated. Where a trading protocol is available, the series is first restricted to the phases in which orders can be submitted, so auctions and closed periods do not contaminate the returns.

Each fact is reduced to a pass or a fail against a criterion fixed in advance, and the statistics behind it are kept alongside the flag.

The eleven facts

Stylised factWhat we measurePass criterion
Absence of autocorrelationsLjung-Box test on returns, up to 10 lagsAll p-values above 0.05, at every frequency
Heavy tailsExcess kurtosis of returns, with Shapiro-Wilk and D’Agostino normality statistics alongsideExcess kurtosis above 0 at every frequency
Gain-loss asymmetrySkewness of returns, and the mean of positive against negative returnsSkewness below 0 at every frequency
Aggregational GaussianityShapiro-Wilk statistic as the sampling interval grows from 1s to 500sPositive slope, so returns become more Gaussian as they are aggregated
IntermittencyFano factor of the count of extreme returns, taken above the 99th percentile of absolute returnsFano factor above 1, so extreme events are over-dispersed rather than Poisson
Conditional heavy tailsKurtosis of returns normalised by local (60s) volatility, against the kurtosis of raw returnsTails survive the volatility correction: normalised kurtosis exceeds raw kurtosis at every frequency
Volatility clusteringLjung-Box test on absolute returns, up to 10 lagsAll p-values below 0.05 at the 1s sampling rate, the rate at which the criterion is defined
Slow decay of autocorrelationsPower-law exponent of the ACF of absolute returns, fitted on a log-log scaleAbsolute exponent between 0.2 and 0.4
Leverage effectCorrelation between returns and absolute returns across lags 1 to 100Mean correlation across lags below 0
Volume-volatility correlationTraded volume against absolute returns, with a Ljung-Box test on the demeaned product of the twop-value below 0.05 at every frequency
Asymmetry in time scalesCorrelation between fine-grained volatility (absolute returns) and coarse volatility (a 20-period rolling standard deviation) at symmetric leads and lagsPeak correlation at negative lags above the peak at positive lags

How we read them

The eleven fall into three groups, and it is the groups rather than the individual flags that tell us something.

Heavy tails, aggregational Gaussianity and conditional heavy tails describe the shape of the return distribution: fat tails at short horizons, thinning towards a Gaussian under aggregation, and surviving a correction for the local level of volatility.

Volatility clustering, slow decay of autocorrelations and intermittency describe memory. This group matters because the first fact in the table is easy to pass for the wrong reason: a simulator producing something close to noise will look comfortably uncorrelated. These are the checks that stop noise from being mistaken for a market.

Gain-loss asymmetry, the leverage effect and asymmetry in time scales describe asymmetry, and they are the hardest to reproduce. They require the direction of a move to matter and not only its size, and volatility at one horizon to lead volatility at another.

We read the profile across frequencies rather than the aggregate flag. Mid-price changes carry genuine microstructure effects at the finest sampling rates, so a fail at 1s with passes elsewhere means something quite different from a fail at every scale. We also run the suite on the historical day: a fact the real market does not satisfy for that symbol and date is not one we hold the simulator to.

Stylised facts therefore give us a third view of the simulation, from the price process that emerges from the order book rather than from the book itself.

Putting it together

There is no single number that tells us whether a simulated market is realistic.

The distributional metrics tell us whether individual properties of the book and order flow look right. Impact response tells us whether market events are followed by realistic dynamics. Stylised facts tell us whether the resulting price process behaves like a financial market.

A simulator can perform well on one of these checks and poorly on another. That is why we use them together.

This is the framework we have used to evaluate the improvements made in Portobello and the other Pulse results we have published. Rather than asking whether the simulated output simply looks realistic, PulseCheck lets us both quantitatively and qualitatively inspect where it matches the historical market and where it still needs work.

These metrics have limitations, as they all depend on statistical analysis of the market data. Our recent research has expanded to other, non-parametric and AI-based approaches to validating our synthetic market data, which we are excited to share soon.

Running it yourself

You can validate your own Pulse simulations against the same benchmark. PulseCheck is available today on the Pro tier: give it a symbol, a date and your simulation runs, and it returns the distances per metric and the impact response curves. If you have built your own generative model or ABM, you can upload its output and have it scored against the same historical day we hold Pulse to, and you can compare it to the Pulse simulations.

The API reference is at pulse.simudyne.com/docs/pulse-check to get started. If you would like validation scorecards for a particular symbol and date, get in touch at info@simudyne.com.

References

Cont, R. (2001). Empirical properties of asset returns: stylised facts and statistical issues. Quantitative Finance, 1(2), 223–236.
Eisler, Z., Bouchaud, J.-P., & Kockelkoren, J. (2012). The price impact of order book events: market orders, limit orders and cancellations. Quantitative Finance, 12(9), 1395–1419.
Nagy, P., Frey, S., Li, K., Sarkar, B., Vyetrenko, S., Zohren, S., Calinescu, A., & Foerster, J. (2025). LOB-Bench: Benchmarking Generative AI for Finance — an Application to Limit Order Book Data. https://arxiv.org/abs/2502.09172
Nagy, P., Frey, S., Sapora, S., Li, K., Calinescu, A., Zohren, S., & Foerster, J. N. (2023). Generative AI for End-to-End Limit Order Book Modelling: A Token-Level Autoregressive Generative Model of Message Flow Using a Deep State Space Network. Proceedings of the Fourth ACM International Conference on AI in Finance (ICAIF ’23).
Ratliff-Crain, E., Van Oort, C. M., Koehler, M. T. K., & Tivnan, B. F. (2025). Revisiting Cont’s stylized facts for modern stock markets. Quantitative Finance, 25(9), 1343–1373.