How to backtest the ICT Silver Bullet (2026)
The Silver Bullet is the most codeable setup in the ICT catalogue — a fixed one-hour window, a defined trigger, a defined invalidation — which makes it one of the few discretionary concepts you can test almost as taught. The part that decides the answer isn't the entry logic, though. It's the control: running the identical mechanics in an ordinary hour, to test whether the famous window is doing any work at all.
What is the setup, stated mechanically?
As commonly taught: during a specific one-hour session — most often 10:00–11:00 AM New York time, with 3:00–4:00 AM and 2:00–3:00 PM as secondary windows — price sweeps a nearby liquidity level, leaves a fair value gap on the reversal, and the trader enters on the gap targeting a liquidity level on the other side.
Three of those four elements are already numbers. The window is a clock. A liquidity sweep can be defined as trade beyond a prior swing extreme followed by a close back inside it. A fair value gap is a three-candle displacement where candle one's extreme and candle three's extreme leave a true gap. Only the target — "the next liquidity pool" — is genuinely soft, and a test has to replace it with something fixed.
A complete rule set you can copy
- Instrument and timeframe: micro index futures, 5-minute bars
- Window: entries permitted only between 10:00 and 11:00 AM ET (run the 2:00–3:00 PM window as a separate variant, never pooled)
- Precondition: a sweep of the prior 60-minute swing high or low inside the window
- Trigger: the first fair value gap forming in the direction opposite the sweep, minimum gap size of 1 tick
- Entry: limit at the gap midpoint
- Stop: 1 tick beyond the sweep extreme
- Target: 2R fixed (run prior session-range midpoint as a second variant)
- Maximum one trade per window per day; no re-entries
- Costs: commission and slippage inside every trade
- Control: the identical trigger and management run between 11:30 AM and 12:30 PM ET, with no sweep precondition
Record the windows separately. Pooling the morning and afternoon windows into one result hides the possibility that one of them carries everything.
The control is the whole test
The Silver Bullet's core claim is not about fair value gaps — those are traded in every session. The claim is that this particular hour delivers price with unusual reliability. That claim is testable in exactly one way: run the same mechanics in an unremarkable hour and compare.
If the neutral-window results land in the same range, the setup's edge — whatever its size — comes from the entry mechanics, not the clock, and the window is a scheduling convention rather than a source of advantage. If the anointed window separates clearly and repeatedly across variants and sub-periods, that's a finding worth having. Either way you learn something the chart examples can't tell you, because chart examples are only ever selected from the window they're promoting.
Run the control across the same date range, same instrument, same costs. Nothing else changes.
What breaks this test in practice
Time zones and daylight saving. The window is specified in New York time, and futures data is frequently stamped in exchange time or UTC. A one-hour offset silently tests a different hour for half the year — this is the single most common way a Silver Bullet backtest ends up measuring nothing. Verify the window on a handful of individual days by eye before trusting any aggregate.
Sample size hidden by the calendar. One window a day means at most about 250 opportunities a year on one instrument, and preconditions cut that further. A "two-year backtest" can easily contain fewer than 200 trades, which is a small sample wearing a long date range. If you need more, add instruments rather than loosening the rules.
Fair value gap definition drift. Minimum gap size, whether wicks or bodies define the gap, and whether a partially filled gap still counts — each choice changes the trade count materially. Fix the definition before the first run and keep it fixed across variants.
Intrabar resolution. A stop one tick beyond the sweep extreme and a 2R target frequently sit inside the same 5-minute bar. Without lower-timeframe fill modelling the tester resolves that by assumption, and the assumption usually flatters. Run it both ways.
Reading the result, and one structural finding you get before any number
Read trade count first, then profit factor after costs versus the neutral-window control, then the shape of the returns. But there's a conclusion available before any number arrives, straight from the setup's design: one window per day on one instrument is a low-frequency profile.
That matters if the destination is a funded account, because evaluation clocks read frequency directly — minimum trading day requirements and 30-day evaluation windows exclude selective systems regardless of their edge. A setup producing a couple of qualifying signals a week needs either several instruments or a program without a time limit. The full mapping of which rules read which strategy properties is in which prop firm rules can a systematic strategy actually pass.
Once you have a trade list, the single historical sequence is one draw from a distribution — push it through the Monte Carlo simulator and read percentiles, using this guide to reading Monte Carlo results. Sanity-check the list first with the backtest quality checker, and if the sample looks thin, the math on how much history you need is in how long you should backtest.
The same treatment applied to related concepts: order blocks, CRT / candle range theory, inverted fair value gaps and liquidity sweeps.
FAQ
What time is the Silver Bullet window?
Most commonly 10:00–11:00 AM New York time, with 3:00–4:00 AM and 2:00–3:00 PM ET taught as secondary windows. For testing, treat each window as a separate variant rather than pooling them — they are separate claims.
Can the ICT Silver Bullet be automated?
The core can: window, sweep, fair value gap, entry and stop are all definable numerically. What resists automation is the discretionary layer most teaching adds on top — daily bias, higher-timeframe narrative, skipping unfavourable days — and that layer is where tested and taught versions diverge.
Why test a neutral window?
Because the setup's claim is about the hour. Without a control run in an ordinary hour you cannot distinguish an edge from the window from an edge in the entry mechanics, which would show up in any session.
Is one setup per day enough for a prop firm evaluation?
Rarely on a single instrument. Roughly 250 opportunities a year before filters means long quiet stretches, and evaluation clocks punish exactly that profile — frequency is a rule-set constraint, not just a patience question.
---
This article describes a testing procedure, not a performance claim. Any figures you generate are your own backtest results, not live performance. Past performance does not guarantee future results. Not financial advice.