PulseCheck: How We Validate Our Market Simulators
In previous posts, we introduced Pulse, our agent-based market simulator, and Portobello, its latest release. Every improvement we reported was measured against the same validation framework. In this post we introduce that framework: PulseCheck.
PulseCheck is our limit order book validation suite. It is designed to answer a simple question: does a simulated market behave like the real one? To answer that properly, looking at the price path alone is not enough. We also need to examine the structure of the order book, how liquidity is distributed across price levels, how orders arrive and are cancelled over time, and how the market responds to different types of events. Together, these features tell us whether a simulation reproduces the underlying market dynamics rather than simply generating a plausible price series.
A realistic price chart is not enough
Put a simulated price path next to a historical one and they may look remarkably similar. That tells us less than it might seem. A simulator can produce a convincing price series while getting the market underneath it wrong. Spreads might be too tight, liquidity might sit at the wrong depths, orders might cancel too slowly, or the book might recover unrealistically quickly after a trade. These differences matter when the simulation is used for execution analysis, stress testing or strategy research. The price is the visible output of a much richer market process. For Pulse to be useful in these settings, the behaviour of the underlying book has to be realistic too.
That is what PulseCheck is designed to measure.
Given historical limit order book data for a symbol and a set of Monte Carlo simulations of the same market, PulseCheck produces a quantitative view of where the simulated market matches the historical one and where it does not. PulseCheck combines distributional metrics, impact-response analysis and stylised facts into a single validation framework. The first two components draw on LOB-Bench (Nagy et al., 2025), while the stylised facts are adapted from Cont (Cont, 2001) and provide an additional check on the behaviour of the resulting price process.
What does PulseCheck measure?
The simulator is not expected to reproduce the exact historical sequence of events. Rather, it should generate a plausible alternative realisation of the same market, so validation focuses on matching distributions rather than individual events.
PulseCheck evaluates 21 distributions covering three parts of the market. Some of the metrics describe the shape of the book at a point in time, some describe how the book evolves, and others describe the behaviour of individual orders.
Book shape
The first group looks at the state of the order book itself.
| Metric | What it captures |
|---|---|
spread | Distance between the best bid and best ask |
touch_bid_volume / touch_ask_volume | Volume resting at the best bid / ask |
total_bid_volume / total_ask_volume | Volume across the top 10 levels on each side |
obi | Order book imbalance at the touch |
These metrics tell us whether the simulator produces realistic spreads, liquidity and imbalance. A price series can look convincing while the book supporting it is systematically too liquid, too thin or too imbalanced.
Order flow
The second group looks at how the market evolves through time.
| Metric | What it captures |
|---|---|
log_interarrivaltime | Log time between consecutive events |
volume_per_minute | Traded volume per minute, inferred from book changes |
ofi_rolling | Order flow imbalance averaged over 100 events |
ofi_rolling_up / down / stay | OFI conditioned on the direction of the next mid-price move |
The conditional order flow imbalance metrics are particularly useful. A simulator may reproduce the overall amount of imbalance in the market while still getting its relationship with subsequent price movement wrong.
In that case the distribution looks right, but the structure underneath it is not. Conditioning OFI on the direction of the next mid-price move gives us a way to detect this.
Order behaviour
The final group looks at individual orders. These metrics require message-level data rather than book snapshots alone.
| Metric | What it captures |
|---|---|
time_to_cancel | Lifetime of an order from placement to cancellation |
ask_limit_depth / bid_limit_depth | Distance from the mid-price at which new limit orders are placed |
ask_cancel_depth / bid_cancel_depth | Distance from the mid-price of cancelled orders |
ask_limit_level / bid_limit_level | Book level (1 to 10) at which new limit orders arrive |
ask_cancel_level / bid_cancel_level | Book level from which cancellations occur |
Together, these metrics give us a fairly detailed view of the simulated market. We can ask whether the book has the right shape, whether it evolves at the right rate, and whether the orders inside it behave in a way that resembles the historical market.
How do we score a match?
For each metric, PulseCheck produces a distribution from the historical market and another from the simulation. The next problem is turning the difference between those distributions into something we can compare.
We use two distances: L1 and Wasserstein-1. They capture different kinds of error, so we keep both.
L1 distance
The L1 distance compares the historical and simulated distributions as histograms.
We pool the observations, construct a common set of bins using the Freedman-Diaconis algorithm, and calculate the proportion of each dataset falling into each bin. We then compute
,
Here, is the proportion of historical observations in bin , is the proportion of simulated observations, and is the number of bins.
The factor of means the score lies between 0 and 1. A score of 0 means the two histograms are identical, while a score of 1 means they contain no overlapping probability mass.
This bounded scale is convenient because every metric can be displayed on the same axis. It also makes the score relatively easy to interpret.
There is one important limitation. L1 tells us how much probability mass is in the wrong place, but not how far away that place is.
Suppose the historical spread is usually two ticks. A simulator that puts most of its mass at three ticks is wrong, but only slightly. Another simulator might put the same amount of probability at ten ticks. L1 can assign similar penalties to both because it only sees that the mass has moved to a different bin.
For that reason, we also look at Wasserstein distance.
Wasserstein distance
Wasserstein-1 is often described as the earth mover’s distance. The intuition is to imagine one distribution as a pile of earth and ask how much work is required to move it until it has the shape of the other distribution.
In one dimension,
,
where and are the empirical cumulative distribution functions.
Unlike L1, Wasserstein distance takes the size of the error into account. Moving probability mass one tick costs less than moving it ten ticks. This makes it useful for detecting systematic shifts and errors in the tails of a distribution.
There is a practical complication. Wasserstein distance inherits the units of the underlying variable. A distance measured on volume might be expressed in thousands of shares, while a distance measured on spread might be expressed in ticks. The raw values therefore cannot be compared directly.
PulseCheck handles this by standardising both the historical and simulated values using the mean and standard deviation of the historical distribution:
,
The important point is that both distributions use the same historical normalisation. We do not independently standardise the simulation. If the simulated distribution is shifted or has the wrong scale, that error remains visible. It is simply expressed in units of historical standard deviations.
Why use both?
L1 and Wasserstein tell us slightly different things about the same comparison.
L1 is useful for understanding how much the two distributions overlap. Wasserstein adds information about how large the displacement is. A simulator can therefore score relatively well under one distance and poorly under the other.
For example, a high L1 score with a low Wasserstein score can indicate that the simulation misses some fine structure while remaining close to the right region of the distribution. A relatively low L1 score with a high Wasserstein score can point to a problem in the tails, where most of the distribution overlaps but a smaller amount of mass sits much too far away.
Looking at both helps us understand not just whether a simulation is wrong, but how it is wrong.
Visualising model performance
Twenty-one metrics and two distance measures produce a large number of scores, especially across many Monte Carlo runs, symbols, and trading days. We use spider plots to summarise these results and compare models directly.
Each axis represents one metric and each polygon one simulator. Distances are averaged across runs, and the radial axis is inverted so that zero lies on the outer rim. Better match with the historical market therefore produces a larger web, while weaker performance appears as a dent towards the centre.
The plots are particularly useful for model comparison. Because all simulators are evaluated against the same historical day using the same metrics and distances, their polygons can be overlaid directly. This makes it easy to see where models differ in their strengths and weaknesses, rather than reducing performance to a single score.
The example below compares Pulse ABM and LOBS5 for 700.HK on 22 December 2025. LOBS5 (Nagy et al., 2023) is a deep generative model that learns and autoregressively generates message-level order flow, whereas Pulse builds the market from the bottom up through interacting, instrument-calibrated agents. The two therefore provide contrasting approaches to generating synthetic markets.
PulseABM is shown as the mean across 10 independent Monte Carlo runs, whereas the LOBS5 result corresponds to a single inference run for one symbol and one trading day. Hovering over each point shows the corresponding distance and, where applicable, the standard deviation across runs.
L1 distance by metric for 700.HK on 22 December 2025, comparing Pulse and LOBS5 against the historical trading day. Pulse shows the mean across 2 independent Monte Carlo runs, while LOBS5 represents a single inference run for the same symbol and day. The outer rim corresponds to a distance of 0 and the centre to 1. Hover for exact values; standard deviations are shown for Pulse across runs.
Wasserstein distance by metric for the same comparison, normalised by the historical mean and standard deviation. The outer rim represents a distance of 0. Because the normalised Wasserstein distance is not bounded by 1, the scale is set by the largest distance in the comparison rather than a fixed limit, so no metric is clipped: the worst axis sits at the centre and every other axis is placed relative to it. Hover for exact values.
When reading these charts, distances of zero are not the target. The aim is not for a simulation to reproduce exactly a trading day, but to give a financially plausible alternative.
Impact response
The distributional comparison of different market statistics is a static analysis. However, the market is not static. When a new order comes through, it impacts the book. Thus, we also need a dynamic measure of how the book evolves throughout the day. This is where impact response comes in. The impact response analysis follows LOB-Bench (Nagy et al., 2025), which uses the event-response methodology of Eisler et al. (Eisler et al., 2012) as its basis.
We begin with events that affect the price or quantity at the best bid or ask. These are divided into market orders, limit orders and cancellations. Each category is then split according to whether the event immediately changes the mid-price or leaves it unchanged, giving six event types in total.
This distinction is useful because two events of the same broad type can have very different effects. A market order that consumes all remaining liquidity at the best quote immediately changes the mid-price. Another market order may be absorbed by the available liquidity and leave the price unchanged. Even when there is no immediate price movement, the event can still contain information about what happens next.
For each event type , we measure the average signed movement in the mid-price after the event:
Here, is the mid-price at event time and is the lag, measured in events rather than seconds. The sign is chosen so that each event’s expected impact on the mid-price is positive.
We calculate the response at log-spaced lags from 1 to roughly 200 subsequent events and bootstrap 99% confidence intervals for the historical data and each simulation. We use event time rather than wall-clock time because the simulated market may operate at a slightly different activity rate from the historical one. Measuring the response in events makes it easier to compare the dynamics of the two books directly.
The response curves can expose behaviour that looks completely normal in the distributional metrics. A simulator might have an accurate spread distribution while replenishing liquidity too quickly after a trade, causing impact to disappear too soon. Another might respond too strongly because liquidity does not react enough.
Stylised facts
The final part of PulseCheck steps away from the book and looks at the price series that emerges from it.
Financial returns have a set of empirical properties that recur across markets, instruments and sample periods. Rama Cont collected them into the list now usually called the stylised facts (Cont, 2001), and a recent study has revisited that list against modern US equity data and found that most of the facts still hold (Ratliff-Crain et al., 2025).
What makes them useful is that they are qualitative statements. They do not describe one particular trading day, they describe how financial price series behave in general. So unlike the distributional metrics and the impact curves, this part of the suite does not need a historical counterpart to compare against. We can ask whether a simulated price process has the properties a market’s price process should have, and answer that from the simulation alone.
PulseCheck implements eleven of Cont’s stylised facts. Returns are mid-price changes, sampled at 1s, 60s, 300s, 900s and 1800s, so each fact is checked at both microstructure and intraday scales. Three of the facts sweep a finer grid instead, from 1s to 500s in 50s steps, because what they test is how a property changes as returns are aggregated. Where a trading protocol is available, the series is first restricted to the phases in which orders can be submitted, so auctions and closed periods do not contaminate the returns.
Each fact is reduced to a pass or a fail against a criterion fixed in advance, and the statistics behind it are kept alongside the flag.
The eleven facts
| Stylised fact | What we measure | Pass criterion |
|---|---|---|
| Absence of autocorrelations | Ljung-Box test on returns, up to 10 lags | All p-values above 0.05, at every frequency |
| Heavy tails | Excess kurtosis of returns, with Shapiro-Wilk and D’Agostino normality statistics alongside | Excess kurtosis above 0 at every frequency |
| Gain-loss asymmetry | Skewness of returns, and the mean of positive against negative returns | Skewness below 0 at every frequency |
| Aggregational Gaussianity | Shapiro-Wilk statistic as the sampling interval grows from 1s to 500s | Positive slope, so returns become more Gaussian as they are aggregated |
| Intermittency | Fano factor of the count of extreme returns, taken above the 99th percentile of absolute returns | Fano factor above 1, so extreme events are over-dispersed rather than Poisson |
| Conditional heavy tails | Kurtosis of returns normalised by local (60s) volatility, against the kurtosis of raw returns | Tails survive the volatility correction: normalised kurtosis exceeds raw kurtosis at every frequency |
| Volatility clustering | Ljung-Box test on absolute returns, up to 10 lags | All p-values below 0.05 at the 1s sampling rate, the rate at which the criterion is defined |
| Slow decay of autocorrelations | Power-law exponent of the ACF of absolute returns, fitted on a log-log scale | Absolute exponent between 0.2 and 0.4 |
| Leverage effect | Correlation between returns and absolute returns across lags 1 to 100 | Mean correlation across lags below 0 |
| Volume-volatility correlation | Traded volume against absolute returns, with a Ljung-Box test on the demeaned product of the two | p-value below 0.05 at every frequency |
| Asymmetry in time scales | Correlation between fine-grained volatility (absolute returns) and coarse volatility (a 20-period rolling standard deviation) at symmetric leads and lags | Peak correlation at negative lags above the peak at positive lags |
How we read them
The eleven fall into three groups, and it is the groups rather than the individual flags that tell us something.
Heavy tails, aggregational Gaussianity and conditional heavy tails describe the shape of the return distribution: fat tails at short horizons, thinning towards a Gaussian under aggregation, and surviving a correction for the local level of volatility.
Volatility clustering, slow decay of autocorrelations and intermittency describe memory. This group matters because the first fact in the table is easy to pass for the wrong reason: a simulator producing something close to noise will look comfortably uncorrelated. These are the checks that stop noise from being mistaken for a market.
Gain-loss asymmetry, the leverage effect and asymmetry in time scales describe asymmetry, and they are the hardest to reproduce. They require the direction of a move to matter and not only its size, and volatility at one horizon to lead volatility at another.
We read the profile across frequencies rather than the aggregate flag. Mid-price changes carry genuine microstructure effects at the finest sampling rates, so a fail at 1s with passes elsewhere means something quite different from a fail at every scale. We also run the suite on the historical day: a fact the real market does not satisfy for that symbol and date is not one we hold the simulator to.
Stylised facts therefore give us a third view of the simulation, from the price process that emerges from the order book rather than from the book itself.
Putting it together
There is no single number that tells us whether a simulated market is realistic.
The distributional metrics tell us whether individual properties of the book and order flow look right. Impact response tells us whether market events are followed by realistic dynamics. Stylised facts tell us whether the resulting price process behaves like a financial market.
A simulator can perform well on one of these checks and poorly on another. That is why we use them together.
This is the framework we have used to evaluate the improvements made in Portobello and the other Pulse results we have published. Rather than asking whether the simulated output simply looks realistic, PulseCheck lets us both quantitatively and qualitatively inspect where it matches the historical market and where it still needs work.
These metrics have limitations, as they all depend on statistical analysis of the market data. Our recent research has expanded to other, non-parametric and AI-based approaches to validating our synthetic market data, which we are excited to share soon.
Running it yourself
You can validate your own Pulse simulations against the same benchmark. PulseCheck is available today on the Pro tier: give it a symbol, a date and your simulation runs, and it returns the distances per metric and the impact response curves. If you have built your own generative model or ABM, you can upload its output and have it scored against the same historical day we hold Pulse to, and you can compare it to the Pulse simulations.
The API reference is at pulse.simudyne.com/docs/pulse-check to get started. If you would like validation scorecards for a particular symbol and date, get in touch at info@simudyne.com.