Methodology

How bars are built

One-minute candles are fetched from Dukascopy's public historical feed and aggregated locally to 15-minute bars on UTC-anchored boundaries. A bar is treated as complete when it is built from at least 12 of its 15 constituent minutes; incomplete bars are stored but excluded from every statistic. Where the market was closed, no bar exists — absent bars are never zero-filled, because an absent bar and a flat bar are different facts.

What the source is, and is not

Dukascopy is a broker CFD feed, not an exchange feed. Prices are Dukascopy's own quotes rather than exchange prints, and the volume field is a tick count rather than traded contracts. This is faithful for statistics of shape and timing — range, gaps, run lengths, time-of-day distributions — and it is not suitable for anything requiring true exchange volume or settlement prices. Read every figure here with that in mind.

Horizons

Everything on this site is reported at 4 hours, 1 day, 3 days and 5 days. Those come from the holding period actually traded — hours to days, usually within a week — and five days is one full trading week. Horizons are wall clock, not bar counts: a position held for four hours is held for four hours whether or not the venue was open, so an entry late on a Friday genuinely gets very few bars of price action. Each horizon publishes the median number of bars its windows contained, so that variation is visible.

Fifteen-minute bars remain the storage and entry-timing granularity, and the 15-minute range is still published as execution detail. It is no longer a headline: a multi-day stop sized to a 15-minute range is precisely the error that shakes a trader out of positions that would have worked.

The barrier grid, and why it needs a null

For every bar, taken as an entry at that bar's close, and for each (target, stop) pair, we record whether the target was reached before the stop within the horizon. There is no entry rule and no signal — entries are every bar, unconditionally — so this is the unconditional first-passage structure of the price path. It is not a backtest and no strategy is being evaluated, but it sits close enough to that line to be worth saying plainly.

A grid of hit rates on its own is arithmetic dressed as insight, because a driftless random walk already answers it. Over unbounded time, P(hit +R before −S) = S/(R+S) — the martingale reference — so a 1% target against a 0.5% stop works about a third of the time for a coin flip.

That figure is not printed on this page, for two reasons. It is wrong at a finite horizon (below), and once it is not the null it is a closed-form function of the row label: the row “+1% / −0.5%” reads 33.3% for every asset, at every horizon, in every year. It cannot take a different value, so no value it takes could change a decision, and beside a measured hit rate it invites precisely the comparison this section exists to warn against.

It is not, however, what significance is tested against, because it is wrong at a finite horizon. Every horizon here truncates the path, and it truncates asymmetrically: inside a four-hour window the tighter barrier is far more likely to be reached than the wider one, so conditioning on the trade having resolved over-weights it. Measured on synthetic driftless data, the gap between a cell's hit rate and its S/(R+S) reference runs to 20–25 percentage points — three times the size of the effect this page exists to detect. Publishing those as findings would be the exact failure this project names as its primary risk.

So the tested quantity is the long-minus-short asymmetry: the same (target, stop) pair traded in both directions, and the gap between them. Its null is exactly zero at any horizon — the short trade is the long trade on the mirrored series, and the horizon truncates both identically, so the truncation cancels instead of needing to be modelled.

Stated precisely, the null is sign symmetry: after the drift component is removed, a path and its mirror image are equally likely. That is a stronger statement than “zero drift”, and the difference matters. A series with no drift at all can still fail this null if it is skewed — if it grinds upward and falls abruptly, up-moves and down-moves reach a barrier by different routes even though neither direction is favoured on average. Volatility that depends on direction, serial dependence, and jumps break it the same way.

So a confirmed cell is evidence of directional path asymmetry, not of a tradeable directional edge. Skew is the first explanation to reach for, not the last. A cell sitting on zero is a real answer too: there is nothing there.

Drift, and what is actually tested

Gold rose substantially over this sample. A rising series reaches its upside barrier first more often for a reason that has nothing to do with path structure, so the raw long-minus-short gap would be significant almost everywhere and this page would report a “long-side edge” that is a restatement of the price chart.

So the tested quantity is the excess over the sample's own drift. The mean log return is subtracted from every bar, the path is rebuilt, and the asymmetry is measured again on it. Subtracting a constant removes exactly one thing — the average slope. Volatility, volatility clustering, skew, fat tails, autocorrelation and the intrabar range all survive untouched. Each cell publishes the raw gap, the part the drift accounts for, and the excess; they add up exactly, and only the excess carries a confidence interval and a significance mark.

Three limits. The mean is estimated, so the correction is itself slightly noisy — it errs toward finding nothing. It removes one whole-sample average, not a drift that changes across the sample. And momentum is invisible to this statistic by design: persistence amplifies up-moves and down-moves equally, so it changes how variance scales rather than who wins the race. That is the variance ratio's question, covered later on this page — so an absence of findings here is not evidence that the series has no structure of any kind.

Barrier placement

Barriers are placed as exact reciprocals: the upside at entry × (1+R), the downside at entry ÷ (1+R). The obvious alternative, entry × (1−R), is not the mirror image of the upside in log terms, and that small difference would put a systematic tilt into a statistic whose whole null rests on the two directions being exactly symmetric. The cost is that a “1%” downside rung sits 0.990% below entry rather than 1.000% — a hundredth of a percentage point of price.

Reciprocal barriers make the null exact rather than approximate: ln(1+R) and ln(1÷(1+R)) are the same distance from zero, by construction, for every rung on the ladder.

One bar can cross both barriers, and OHLC cannot order the two touches. Those resolve as the stop — the rule fails against the trader — and the rate is published per cell. It is bounded: measured across every long and short cell this site currently publishes, the highest ambiguous rate is 8.76% of resolutions, at the tightest rungs only, and a tie counts as a stop for the long and the short reading alike, so most of it cancels in the gap that is actually tested. Even the worst case moves the tested gap by well under a percentage point — smaller than every effect this site currently confirms, and more than an order of magnitude below the typical one.

Within a finite horizon some entries hit neither barrier. The primary table on the asset page shows the unconditional shares of all entries — target first, stop first, neither — which sum to 100%. The conditional rate — target first among only the entries that reached either barrier, labelled “of resolved” — and the resolved share itself are in the collapsed detail block underneath, not beside the headline number: a cell where only a tenth of entries resolved would read as a near-100% conditional rate on a handful of observations, which is why that figure is demoted rather than led with. A cell is withheld entirely unless it clears both publication floors: at least 100 resolutions in each direction, and a bootstrap effective sample size of at least 30 on both sides. The second floor is the one that usually bites — most of the cells this site currently withholds cleared the resolution count comfortably, some of them by more than a factor of ten, and were held back because their entries overlap too heavily to carry that much independent evidence. A resolution count on its own is not a sample size; see Overlap, and effective sample size below.

Excursions, and heat before the win

Not every window in the barrier grid resolves — especially at shorter horizons and wider rungs, many windows touch neither barrier before the horizon runs out. Excursions measure what those windows did anyway, resolved or not. For every entry, over its forward window, MFE (maximum favourable excursion) is the best price reached in the trader's favour and MAE (maximum adverse excursion) is the worst price reached against them — both taken from the window's intrabar high and low, never from closes: a stop is hit by the low of a bar, not its close, so measuring from closes would understate every excursion. The entry bar itself is excluded from its own window — an excursion is what happens after the position is opened, and the entry bar's own high/low sit partly behind the entry price rather than ahead of it.

MFE is signed positive and MAE signed negative, and neither is clamped at zero. A window that gapped away from the entry and never recovered has a genuinely negative MFE — price never once traded in the trader's favour before the horizon ran out. That is a fact about the window, not an error to correct, and it is published exactly as measured.

Heat before the win asks a narrower, more direct question: of entries whose MFE reached at least a target rung, what share drew down past a stop rung before the target was reached — i.e. would the stop already have closed the trade before it went on to work. The ordering is the entire statistic: a drawdown that happens after the target was already reached does not count, because the position was closed at the target and no longer exposed to the drawdown that followed. A version that instead asked “did MAE ever breach the stop AND did MFE ever reach the target” — without the ordering — answers a different, and materially larger, question than the one this statistic exists to answer.

Heat before the win is long-side only. There is no short-side version of it anywhere on this site.

Like the barrier grid, both of these are descriptive statistics about the shape of the price path, not a backtest: there is no entry rule, no exit rule and no position sizing here, only what price did after an unconditional entry at every bar's close.

Sigma equivalents

Each barrier rung is also shown in standard deviations of the move at that horizon: σ is the standard deviation of 15-minute close-to-close log returns across the window, scaled by the square root of the median number of bars in a holding window. So, on gold's current sample, “1% ≈ roughly 2σ at 4h” means a 1% barrier sits about two standard deviations away from a four-hour entry — an illustrative magnitude, not a fixed bound: the exact multiple is asset- and sample-specific and the live per-horizon figure is on the asset page.

Two caveats, and neither is small. The √t scaling assumes variance grows linearly with horizon — that is exactly the assumption the variance ratio tests, covered later on this page. And using σ as a unit does not make the distribution normal; these series have fat tails, so a 2σ barrier is reached far more often than a normal table would suggest.

Overlap, and effective sample size

Consecutive entries share almost their entire window: two adjacent entries in a 3-day study differ by one bar out of the roughly 220 such a window actually holds — the exact median is published beside every horizon. A row count is therefore not a sample size, and reporting n in the hundreds of thousands for such a study — on gold's current five-year sample, well over 100,000 — would be the largest false-confidence artefact this project could produce.

Every interval here is produced by a moving-block bootstrap over calendar weeks — whole weeks are resampled with replacement, so the correlation inside a week is carried into every resample and shows up as interval width. Each cell publishes an effective sample size: the number of independent observations that would have produced the interval shown. Expect it to be one to three orders of magnitude below the row count. That is the honest figure.

A horizon is suppressed entirely unless it holds at least 100 non-overlapping entries — the real count of eligible entries whose holding windows do not overlap one another. That is a deliberately different number from coverage-span equivalents (span ÷ horizon length, a calendar quotient), which is shown beside it on the asset page: coverage-span equivalents counts a weekend or a venue halt as tradeable time, so it overstates how much independent data actually exists. The gap between them is large enough to matter: on gold's current five-year sample, 4h runs roughly 11,000 coverage-span equivalents against 7,700 non-overlapping entries, and 5d runs roughly 370 against 260. These are illustrative magnitudes from one asset's current sample, not fixed bounds — the exact counts move as the nightly ingest grows the history, and the live per-horizon figures are always on the asset page. Twelve trailing months holds on the order of 70 five-day coverage-span equivalents but only around 50 non-overlapping entries — both below the floor of 100, so the 5-day grid does not report on the trailing-12-month view, by design.

The same non-overlapping count gates a second, stricter rule: a quantile at the qth percentile is suppressed unless it has at least 20÷min(q, 1−q) non-overlapping entries. The rule is symmetric because a tail is a tail regardless of which side of the median it sits on: p99 needs roughly 2,000, and p10 needs the same ~200 as p90 — MAE is signed negative, so its worst-case figure is p10, and a p10 built on too few independent observations is exactly as unsupported as a p90 built on the same count. Gating that rule on coverage-span equivalents instead used to let one figure through that should not have cleared it: the trailing-12-month 4-hour p99 move quantile cleared its ~2,000-observation requirement only by counting closed-market time as though it were tradeable, and it is withdrawn under the corrected gate.

Each window runs a hundred comparisons — four horizons by twenty-five stop/target pairs. At 95% confidence with no correction, roughly five of those look significant by chance alone, and a reader scanning the page naturally lands on the most striking one. A Benjamini–Hochberg false-discovery-rate correction at q = 0.05 is therefore applied across the whole window rather than per horizon, and both the corrected and uncorrected counts are shown — the gap between them is itself informative.

Surviving Benjamini–Hochberg is necessary but not sufficient. A cell is reported as a finding only if it also survives a second, independent check: both halves of the history agree on the direction of the effect (a split-half sign comparison). The split-half comparison is computed on the full-history window only, never on the trailing-12-month window — halving twelve months leaves each half too short for the check to detect real instability, and a check that cannot detect instability would still render as a reassuring green column that means nothing.

A cell can therefore be statistically significant yet unconfirmed. It is still shown — without emphasis, no bold, no ring — because the reader should see that the grid contains large, unreliable numbers and be able to see why they are unreliable, rather than have them quietly disappear. Its number, interval and split-half state stay on the page; only the visual claim of “this is a finding” is withheld. A large asymmetry on a cell where almost nothing resolved is arithmetic on a handful of observations, not an edge — the effective sample size beside every cell in the primary table, and the resolved share in the detail block beneath it, are how that becomes visible rather than asserted. Only cells that pass both guards — Benjamini–Hochberg and split-half sign agreement — are marked in bold and ringed on the charts.

Entry sampling is an open question rather than a settled one: sampling every bar maximises data and maximises overlap, while a strictly non-overlapping grid is honestly independent and far thinner. Both are computed, and the difference between them is published per horizon.

Costs

The spread is a configured assumption, not a measurement. This feed publishes bid candles only, so there is no bid–ask in the data; high–low estimators exist but add a second layer of uncertainty to a number whose entire job is to be a sanity check. One full spread is charged per round trip, against the median absolute move at each horizon.

Sample size

Every figure carries its n. Any cell computed on fewer than 100 observations is shown as “insufficient data” rather than a number, per cell rather than per statistic. Statistics built on overlapping windows carry a second and stricter floor — see Overlap, and effective sample size above.

How returns are defined

A bar's return is the log change from the previous close — except for the first bar after any break in trading, which uses that bar's own open instead. Without the exception, the first bar of a session would absorb the entire overnight or weekend move. Because that recurs every session, it would show up as a large and apparently very significant opening effect that is really just the gap. Gaps are real and are measured separately.

Trend, reversion, and how moves scale

If price moved by coin flips, the variance of a move would grow in exact proportion to the time held. The variance ratio measures whether it does: VR(q) = Var(q-bar move) ÷ (q × Var(1-bar move)), with a null of exactly 1. Above 1 moves extend, below 1 they retrace. This is the only statistic on this site that can see momentum — the barrier grid measures a race between two barriers, and persistence speeds both runners equally, so it is blind to trend by construction.

q is the median number of bars inside a wall-clock window of that horizon, taken from each asset's own bar calendar rather than from calendar arithmetic. That draws on the same wall-clock windows as the barrier grid's, though the two medians can differ by a bar or two: the barrier grid counts only entry-eligible bars, while this statistic counts every stored bar, per this project's rule that complete gates entries and never filters a path array. It means the labels are not naive either way: a five-calendar-day window from a random entry usually contains a weekend, so “5d” is materially fewer than five trading days of price action.

Only the heteroskedasticity-consistent standard error is tested. Lo–MacKinlay give two. The simpler one assumes the size of moves is constant over time; real markets cluster quiet periods and violent ones together, and under that assumption ordinary volatility clustering reads as mean reversion. Across every horizon this site currently publishes — all four assets, both windows — the simple version's z-scores run as far as −3.98 (gold's all-history 4-hour cell), while the robust version never exceeds ±1.8 in magnitude anywhere — its own largest, −1.79, is the S&P's trailing-12-month 3-day cell. Both are published, and the gap between them is the honest measure of how much a naive test would have overstated.

Variance scaling and distribution shape are separate facts, and both are consistent with this data. Nothing here rejects proportionality — the robust test fails to reject VR = 1 on every published cell, so variance grows in proportion to time as far as this (deliberately low-powered) test can tell — yet the typical move grows faster than the square root of time regardless. There is no contradiction: aggregation pulls the distribution toward normal, so median ÷ σ rises and kurtosis falls as the horizon lengthens. The consequence is practical — σ may be scaled by √t, but a median, a stop distance or a target may not. For the median the direction is consistent: on every asset and window published here, scaling one sets it too tight. The tails are not consistent — the p90 runs ahead of √t at 5 days on all four assets, but a few percent behind it at 1 and 3 days on some — so a stop or a target is read off the measured table, never scaled to it.

What this is not

These are descriptive statistics about what has already happened. Nothing here is a prediction, a signal, or trading advice. A pattern visible in historical data may reflect a persistent feature of the market, a temporary regime, or chance.