Do Candlestick Patterns Work? What Two Studies Found
Two studies read for this post ran candlestick rules against a bootstrap significance test: on Dow component stocks nothing came back significant, and on three CME currency futures contracts only a handful of forms did.

Two studies located for this post put candlestick patterns through a formal statistical test, and neither found much. Marshall's 2005 Massey University thesis tested candlestick rules on the stocks in the Dow Jones Industrial Average from 1992 to 2002 and concluded that the method "does not have value". Röthig, Röthig and Chiarella tested the same family of rules on British pound, Swiss franc and Japanese yen futures at the CME from 1978 to 2013, and reported that only very few of them carried statistically significant predictive value.
That is the short answer. The more useful answer is why a pattern that looks obvious on a chart is so hard to test: how it gets defined, how rarely it occurs, and what counts as the trend before it are the same problems facing anyone trading one.
Do candlestick patterns work?
In the two formal tests read for this post, very little survived. Marshall found that on Dow component stocks the strategies were unprofitable even before trading costs and risk adjustment were subtracted. Röthig and co-authors found that across all three currency futures contracts, only the Long White/Black Candle and the White/Black Paper Umbrella showed statistically significant predictive value.
Marshall's t-tests did produce statistically significant results, and the thesis reports that every one of them pointed the wrong way. The forms labelled bullish, one-bar and multi-bar alike, were followed by returns below the sample average, and the bearish ones by returns above it — the reverse of what candlestick theory predicts. Under the sturdier bootstrap test, which does not rely on the assumptions a t-test makes about financial data, that significance disappeared.
The futures paper is slightly less bleak. Most of its rules were either missing from the bootstrap samples or came back with p-values above 0.05. Two did reject the null of no predictive value — Three Outside Up and Bullish Engulfing — in the Japanese yen contract alone. Both scope limits matter: Marshall's evidence is about US blue-chip stocks, and Röthig's is about three currency futures contracts. Neither tested an equity index future such as the E-mini S&P 500.
What is a candlestick, and what turns one into a pattern?
A candlestick encodes four prices from one period: the open, the high, the low and the close. Marshall's thesis puts the rule for colour plainly: when the close is above the open the body is white, when it is below the body is black. Röthig and co-authors name the parts: open-to-close is the body, and the thin extensions from it up to the high and down to the low are the upper and lower shadows. A single day's candle is called a single line in the candlestick literature.
Consecutive single lines make up the patterns. Marshall separates them into continuation patterns, which are read as the existing trend persisting, and reversal patterns, which are read as the trend changing. The distinction carries a practical consequence the thesis is explicit about: single lines are claimed to have forecasting power on their own, while a reversal pattern cannot be identified at all until you have decided what the prior trend was.
A candle is also a function of where you cut the tape. Quantower documents ten chart types, switched from an Aggregation Type menu, and six chart styles with Candle among them. Its tick chart starts a fresh bar once a fixed count of trades has printed, 500 in the documentation's example, where a time chart starts one at the end of each period such as five minutes; volume bars are cut by tick or exchange volume with no time element in the aggregation at all. The same trades produce different bodies, different shadows and therefore different patterns depending on which you picked.
Who defined candlestick patterns, and when?
Marshall's thesis traces the principles the West now calls candlestick analysis to rice trading in Japan in the 1700s or earlier, and credits today's methodology to the trading principles of Munehisa Homma, a merchant who began trading at the Sakata rice exchange in 1750. The thesis names Shimizu (1986) as the seminal Japanese-language text on the subject.
The Western arrival has a date. The thesis attributes it to Steve Nison's Japanese Candlestick Charting Techniques, published in 1991, and makes that date load-bearing. The thesis starts its sample on 1 January 1992 for two stated reasons: no seminal English-language candlestick book existed before 1991, and major data providers only began supplying open, high, low and close prices from the middle of that year.
For the definitions themselves, Marshall works from the practitioner literature — Bigalow, Fischer and Fischer, Morris, Nison, Pring, and Wagner and Matheny — cross-checked against a translation of Shimizu to make sure the English books had not quietly dropped or altered anything. That matters for the next section, because the thing being tested is the practitioner definition, not a researcher's reinterpretation of it.
How did the two tests differ?
The two tests share a basic design — a mechanical recogniser for the forms, a ten-day holding period, a bootstrapped null model — and then part company on how the prior trend gets defined, which is itself a sign of how loosely these rules are written down.
| Marshall (2005) | Röthig, Röthig & Chiarella (2015) | |
|---|---|---|
| Market | Stocks in the Dow Jones Industrial Average | British pound, Swiss franc and Japanese yen futures at the CME |
| Period | 1 Jan 1992 – 31 Dec 2002 | 3 Jan 1978 – 29 Nov 2013 |
| Data | Individual stock prices | Datastream futures prices, 8,937 observations for the yen contract |
| Prior trend | Price above or below a 10-day exponential moving average, following Morris (1995) | A 65-day exponential moving average of closes, attributed to Nison (1991) |
| Holding period | Ten days, with sensitivity runs at five and two | Ten days, entering at the next day's open |
| Significance test | Bootstrap: 500 resampled series from a fitted null model | Parametric bootstrap on a multivariate Pair-Copula null model |
| Finding | Strategies did not have value on this sample | Few rules had statistically significant predictive value |
How do you test a candlestick signal against chance?
By generating artificial price histories that preserve the market's statistical behaviour but carry no pattern information, running the same rule on those, and asking how often the fake data beats the real data. Both papers do a version of this, and it is the step that dissolved most of the apparent results.
Marshall's bootstrap works by fitting a null model — a random walk, an AR(1), a GARCH-M or an EGARCH — to the close series, resampling its residuals 500 times, and building 500 artificial price series that are random by construction but keep the statistical fingerprint of the original. A rule counts as significant at 5% if fewer than 25 of those 500 simulated series beat the real one. Because candlestick rules need highs and lows as well as closes, the thesis extends the method by sampling from the distribution of each bar's (high − close)/close and (close − low)/close distances to manufacture plausible highs and lows around each simulated close.
Röthig and co-authors measure a rule's profit as the share of signals after which the ten-day return finished on the right side of zero, which puts the bar for "profitable" at better than 50%. Their own footnote concedes the obvious limitation: a rule can be right less than half the time and still make money if the winners are larger. They also had to patch the data: on some days the settlement price printed outside the day's range, struck after a thin close, so the high and low were widened to take it in.
How often do these patterns actually appear?
Rarely — the quiet problem underneath every test of them. In thirty-five years of daily data on three currency futures contracts, Röthig and co-authors identified the hammer between two and seven times per contract. Here are their identification counts for a selection of the bullish forms they classified:
| Form | British pound | Swiss franc | Japanese yen |
|---|---|---|---|
| Hammer | 2 | 6 | 7 |
| Tweezer bottom | 7 | 12 | 9 |
| Dragonfly doji | 1 | 9 | 2 |
| Three outside up | 33 | 43 | 30 |
| Bullish engulfing | 78 | 97 | 82 |
| Bullish harami | 174 | 183 | 156 |
| Long white candle | 2,195 | 2,093 | 2,016 |
A rule with two realisations in thirty-five years cannot be evaluated, and the paper's results reflect it: the hammer's ten-day outcome finished above zero on neither of the two pound signals, on two of the six franc signals and on two of the seven yen signals. Marshall ran into the same wall and handled it by exclusion — any single line or pattern appearing fewer than ten times across the whole sample was dropped before testing, on the grounds that no test on it would be robust.
The contrast with the bottom row of that table is the thing to take away. Long white candles are everywhere, and they are also one of the forms that survived the futures paper's significance test. Rarity and testability move together. In this sample the single lines — the simplest forms, one bar each — are the ones with enough observations to say anything about, and the multi-bar patterns are not.
Why does the definition change the answer?
Because the practitioner literature fixes some thresholds and leaves others open, which means the person running the test has to choose the rest, and the choices move the results. Marshall is candid about this: candlestick books are precise about some conditions, quoting Morris (1995) that where a white candle has to open near its low and finish near its high, the gap "should be less than 10% of the open-close range". They are flexible about others, such as how far apart open and close must sit before a candle counts as long. The thesis is careful to say where the problem bites: the books are clear on which groupings of single lines add up to a pattern, so the latitude is all in defining the individual lines.
The catalogue shrinks dramatically once those criteria are applied. Of the forms listed in Morris (1995) — 18 single lines, 44 reversal patterns and 14 continuation patterns — Marshall's screens left 14 single lines and 14 reversal patterns, and not one continuation pattern qualified. Twenty-two of the reversal patterns were excluded as having no explanatory power and eight as too infrequent to test.
Röthig and co-authors hit the same imprecision and solved it differently, building a Mamdani-type fuzzy system to convert the verbal rules into something a computer can score, on the explicit grounds that the written rules are imprecise and rest on traders' causal knowledge. Their prior-trend filter is a 65-day exponential moving average; Marshall's is ten days. Both cite the practitioner literature for the choice. A pattern that requires a prior downtrend is a different pattern depending on which of those you use, and so is its measured performance.
This is the same hazard that curve fitting in backtests creates from the other direction: when a rule has free parameters, someone has to set them, and the setting that looks best on the sample is rarely the one that holds up outside it. Marshall's response was sensitivity analysis — re-running everything across different single-line definitions, trend definitions and holding periods — and the thesis reports the conclusion held.
How does a charting platform decide a pattern is there?
By arithmetic on the open, high, low and close of the current and preceding bars, with every condition written out as a comparison. Sierra Chart's CandleStick Patterns Finder offers 63 numbered patterns, from 01 Hammer to 63 Doji, each with a three-letter chart label such as HMM, UEN or MST, and lets you search for up to six at a time per copy of the study.
The documentation spells out the conditions. For Bullish (Up) Engulfing, code 03, every one of these has to hold: the prior bar's last price sits below its open, the current bar's last price sits above its open, the current high exceeds the prior high, the current low undercuts the prior low, the current open is below the prior bar's last price, and the current last price is above the prior open. Bearish Engulfing is the mirror image. There is no interpretation anywhere in that list: the study either finds those six comparisons true across a pair of bars or it finds nothing.
Trend is the part the study has to infer for itself. Sierra Chart marks 60 of its 63 patterns as requiring trend detection — the exceptions are Kicker, Kicking and Doji — and determines it with a linear regression over a window of prior bars, four by default and excluding the bar being labelled, treating a slope of 1 or more as an uptrend and −1 or less as a downtrend, after scaling the slope by a value derived from the recent price range. The page also notes that switching trend detection off increases how often a pattern code appears. Detected pattern numbers go into a study subgraph, readable from alert formulas, spreadsheet studies or custom code. The page's own stamp dates its last revision to 1 February 2023.
What does this mean on a funded evaluation account?
The tested horizon is not available to you. Both papers measure results over ten days after a signal, and on a simulated evaluation account like ours everything closes by late afternoon and every day starts flat, so a candlestick signal cannot be held for the period in which it was tested. Whatever these studies measured, it is not what an intraday trader is doing with a hammer on a five-minute chart.
Sizing is the part no pattern can help with. A candle tells you nothing about how much to risk, while the evaluation rules settle it: the gap between your balance and your drawdown line is the whole loss budget you have, and stop distance in ticks times tick value is what spends it. That arithmetic is the subject of futures position sizing, and it holds no matter which pattern put you in the trade.
If candlestick patterns are part of how you read a chart, the honest way to carry the research is as a bound on expectations. The forms are a compact way to see a session's open, high, low and close at a glance. The evidence located here does not support treating them as predictions, and our risk disclosure says what happens to most people who try. Our write-up of the head and shoulders pattern goes through the testing on that one.
FAQ
Do candlestick patterns work on intraday charts?
Neither study tested that. Marshall's thesis uses daily stock prices and Röthig's paper uses daily futures prices, so both measure what happened over the ten days following a daily-bar signal. Neither result is evidence about a five-minute chart, and Quantower's documentation is a reminder of why: bars can also be cut by trade count or by traded volume, and each of those aggregations draws different candles from identical trades.
Which candlestick signals held up in the futures test?
Across all three currency futures contracts, the Long White Candle, Long Black Candle and White and Black Paper Umbrella showed statistically significant predictive value. Two patterns, Three Outside Up and Bullish Engulfing, rejected the null of no predictive value in the Japanese yen contract only. Most of the rest either produced p-values above 0.05 or appeared too rarely to measure. Statistical significance in that test is a statement about predictive value, not a profit figure.
Do these results apply to the E-mini S&P 500?
No. Marshall's sample is the stocks in the Dow Jones Industrial Average and the futures paper covers British pound, Swiss franc and Japanese yen contracts. Neither includes an equity index future, so neither says anything about how a candlestick rule would have performed on one. Applying a currency futures result to an index future would be an assumption, not a finding.
Why do platforms ship so many candlestick patterns if the evidence is thin?
They are cheap to compute and unambiguous once defined. Sierra Chart's study carries 63 of them, and its Bullish Engulfing definition reduces to six comparisons between the open, high, low and last price of two bars — the kind of test that runs on every bar of every chart and exports to alerts. Availability in software is not evidence about profitability, and the two studies cited here answer that question only for the markets they covered.
Sources
- Massey University Research Online — Marshall (2005), Candlestick Technical Trading Strategies: Can They Create Value for Investors? (thesis PDF)
- Massey University Research Online — thesis record page
- UTS Quantitative Finance Research Centre — Röthig, Röthig & Chiarella (2015), Research Paper 362, On Candlestick-based Trading Rules Profitability Analysis via Parametric Bootstraps and Multivariate Pair-Copula based Models
- Sierra Chart — Technical Studies Reference, CandleStick Patterns Finder
- Quantower — Chart Types
- Quantower — Tick Bars
- Quantower — Volume Bars
Educational content about futures markets and simulated trading. Not investment advice, and not a solicitation to trade. Trading futures involves substantial risk of loss. Read the full risk disclosure.