A rollout plan that says start small and scale up as confidence grows is not a plan, because confidence is not an observable and nobody has ever failed to have it. In practice it describes a profile that goes live at a token size, behaves for six weeks, and gets larger every time nothing bad happens, until it is carrying real risk on the strength of a quiet run. No criterion was ever stated that the strategy could have failed, so no promotion was a decision. It was an absence of alarm, recorded as a judgement.
The alternative is dull and it holds up in a review. Four capital stages. Each one has a written entry gate, a written exit gate and a minimum dwell time, all three fixed before the first live order. The unit of advancement is evidence, and the evidence is named in advance so a stage can be failed rather than merely survived.
Four stages, and what each one actually buys
The stages are not four sizes on a ladder. Each answers a question the previous one structurally could not.
Stage zero, alert only. The profile runs and is not permitted near capital. Blockcircle runs paper alongside live and reports the two separately, so this is a state the engine supports rather than something you improvise with a disabled profile. What stage zero buys is infrastructure proof. Signals arrive, they arrive once rather than three times, the timestamps are coherent, the instrument list matches the mandate, and the profile trades at the frequency you were promised. What it cannot buy is anything about fills, because paper fills flatter you by construction. Keep any profit figure from this stage out of the deck.
Stage one, minimum viable size. Small enough that losing the whole stage allocation is an operational irritation rather than a reportable event, and large enough that the fills are genuinely yours. The purpose is one measurement: implementation shortfall against the fill assumption used in the backtest, per instrument, in basis points, with a distribution rather than a mean. You are buying one number per instrument and there is no other way to get it.
Stage two, attributable size. Roughly a third of target. The P and L is now large enough to survive a monthly review as a real line rather than a rounding artefact, and the guards start binding occasionally instead of theoretically. This is where you discover what the caps do to the strategy, because a cap that has never bound has told you nothing.
Stage three, target size. At this point the binding constraint stops being your process and becomes capacity and crowding, which are different problems with other owners.

Writing a gate that a strategy can actually fail
A gate is only a gate if it is countable, sourced from a named artefact, and judged by someone who does not own the strategy. Anything phrased as satisfied with performance is a mood. The numbers below are the ones I work with for a profile trading a daily or four hour signal. They are mine, not the product's, and the dwell figures need recalibrating to your own trade frequency.
| Promotion | Minimum dwell | Entry evidence required |
|---|---|---|
| Zero to one | 20 trading days and 15 signals, whichever is later | Delivery log with zero missed signals. Instrument list reconciled to mandate. Paper trade list matching the signal feed one for one. |
| One to two | 40 trading days and 30 fills | Measured shortfall per instrument with a distribution, not an average. No unexplained reconciliation break older than two days. At least one guard hit that behaved as documented. |
| Two to three | 60 trading days and 60 fills | Shortfall still inside the band measured at stage one. Capacity note refreshed against current volumes. Supervisor sign-offs complete for every period, including the quiet ones. |
Notice what is absent. No gate mentions the strategy making money. That is deliberate and it is the part people argue with. Over twenty or forty trades the P and L is dominated by noise, so a profit gate promotes on luck and a loss gate demotes on luck, and either way a random number makes the decision. What the gates test instead is whether the operation is sound and whether live behaviour matches the documented behaviour. Return evidence needs a sample you will not have for a year.
Dwell time is the control that gets deleted first
Dwell is where the pressure lands, because it is the only gate that cannot be satisfied by working harder. Someone will argue that the profile has clearly settled and the remaining three weeks are ceremony. Hold the line, in two units at once: calendar time and signal count, whichever is longer.
Both units are needed because they fail in opposite directions. A four hour profile can produce thirty fills in a fortnight, which satisfies a count gate while having seen only one market condition. A daily profile can produce three signals in a month, which satisfies a calendar gate on a sample too small to describe. The two profiles listed on my overview at capture were Follow: NDX MTE 1day and Follow: XAG/USD MTE 4h, which is this pair of clocks under one engine, and one calendar dwell applied to both would mean two very different amounts of evidence.
Add one qualitative requirement to the final promotion. The dwell period should contain at least one genuine change in market conditions, and if it did not, the promotion memo has to say so. That sentence is what stops a strategy reaching full size having only ever traded one regime, which is the most common way a well governed rollout still ends badly.
The ladder has to run downward as well
A staged rollout with no exit gates is a ratchet, and a ratchet is not a control. Write the demotion criteria at the same time as the promotion criteria, in the same document, and give the supervisor standing authority to apply them rather than a recommendation to escalate.
I split demotions into two classes. Straight to stage zero, no discussion, for anything that means you no longer know what the system is doing: an unexplained reconciliation break older than two days, a signal delivery failure, a position outside the mandate, or the profile acting on a price you cannot verify was live. Demotion by one stage, reviewed at the next scheduled meeting, for evidence that the strategy is worse than documented but still understood: shortfall persistently outside the stage one band, guard hits at a rate that means the cap has become the position sizer, or a drawdown beyond the level written into the stage.
The distinction matters because the two failures need different responses. The first class is an integrity failure and size is irrelevant to it. The second is an economics failure, where the strategy may well be fine at a smaller number.
What the engine enforces, and what only your process does
Be precise about the boundary, because this is where staged rollouts are quietly abandoned while appearing to continue. Autopilot enforces per profile controls, describing paper mode, risk guards, trailing stops and partial take profits, with per source rules and an audit trail. The overview separates eleven profiles from ten enabled, so enablement is per profile, and the actions row offers engine wide pause, resume and an emergency kill switch. Those are real controls, and the ones that will save you at three in the morning.
None of them knows your stage ladder exists. There is no reading on that page that distinguishes a profile deliberately at stage one from a profile someone quietly moved to target size on a Friday. Stage discipline lives entirely in your change record and in the arithmetic of who is allowed to alter a size parameter. If the same person can write the gate, judge the gate and change the size, you have a document rather than a control.
And the last thing worth saying plainly: reaching stage three does not make an unattended profile safe. It makes it measured, which is a different property. A profile at full size with working guards and a named supervisor will still open positions on a stale price that has not moved for six hours, still fire into a venue outage, and still keep trading a strategy whose edge stopped existing three weeks ago, because none of those conditions announce themselves as errors. The staging buys you the right to be surprised at a size you can absorb. It does not buy you the absence of surprises.