How to Backtest Your Own Trading Strategy (Without Coding)

Backtesting means taking your written buy and exit rules, running them over past daily price data, and measuring what they would have made, what they would have lost, and how deep they would have drawn down. It does not answer "will this stock go up" — it answers "what is the temperament of my rules." You do not need to code: as long as every rule is a condition a machine can evaluate, a tool can run the whole history for you.

Key takeaways

  • A backtest should give you at least four numbers: win rate, profit factor, max drawdown, and Sharpe ratio. Reading only total return is the most common trap.
  • Expectancy = win rate × average win − loss rate × average loss − cost per round trip. Subtract costs and many "profitable-looking" strategies turn negative.
  • One win out of six signals tells you nothing. With tiny samples, the numbers are noise.
  • Lookahead bias and survivorship bias both inflate results; you have to actively rule them out.
  • Whatever you tune on one slice of data, verify on another slice you never touched (out-of-sample).

What a backtest measures

A meaningful backtest gives you at least four numbers:

  • Win rate: the share of trades that closed with a profit. It is the most intuitive number and the least useful on its own.
  • Profit factor: total money made on winning trades ÷ total money lost on losing trades. Above 1 means net profit. A low win rate with a high profit factor still makes money; a 70% win rate where each win is one-third the size of each loss still loses.
  • Max drawdown (MDD): the largest peak-to-trough fall in the equity curve. It decides whether you can actually hold on — a strategy with a great annual return and a 60% drawdown is one you would most likely abandon at the bottom, never seeing the recovery.
  • Sharpe ratio: excess return per unit of volatility. Two strategies with the same total return can feel completely different to hold; the higher-Sharpe one is the smoother ride.

Read all four together. Any one of them alone can be contradicted by the other three.

Expectancy: what one trade is worth on average

Fold win rate and trade size together and you get expectancy — how much you make, on average, every time you take a trade. The formula is short:

Expectancy = win rate × average win − loss rate × average loss − round-trip cost

Work through a concrete case. Say your backtest reports a 40% win rate, an average win of +10%, and an average loss of −5%. Each trade is 100 shares at $100, so $10,000 per position.

  • A win: sell at $110, make $1,000.
  • A loss: stopped out at $95, lose $500.
  • Over 10 trades: 4 wins = +$4,000, 6 losses = −$3,000, gross profit +$1,000, or +$100 per trade (+1%).

Positive so far. Now add costs. Slippage plus commission plus any stamp duty comes to roughly 1.5% per round trip in this example (Hong Kong runs higher — see Hong Kong trading fees), which is $150 per trade and $1,500 over ten.

  • Net result: $1,000 − $1,500 = −$500, or −$50 per trade (−0.5%).

A strategy with +1% gross expectancy is a losing strategy after costs. That is why "are costs included" matters more than "what is the total return" in any backtest report. You can also read off the profit factor: gross, $4,000 ÷ $3,000 ≈ 1.33; after costs, winners net $850 each and losers net $650 each, so $3,400 ÷ $3,900 ≈ 0.87 — below 1.

Three steps to a clean backtest

  1. Write the rules as explicit conditions. Not "buy when it's low," but something a machine can evaluate, like "RSI < 35 AND today's change > 1.5%" (see how to turn logic into buy rules). Exit rules need the same precision — how much loss, how much gain, how many days at most. A backtest without an exit rule is not a backtest.
  2. Use data that is long and wide enough. Cover at least one full cycle, ideally including a real bear market. Test only last year's bull run and every "buy and hold on" rule looks brilliant.
  3. Don't overfit. Parameters tuned to a perfect past usually just memorized noise. A rule with eight conditions, each set to two decimal places, is almost certainly fitted. Fewer, blunter conditions tend to survive the future better.

The two most common biases

Lookahead bias: the rule uses information that was not available at the time. The classic case is confirming a signal on today's close but assuming you bought at today's open — in reality, by the time you see the close, the open is long gone. A subtler version is using restated financials to judge what the valuation looked like back then. The clean approach: a signal confirmed at the close of day T can only be filled on T+1 or later.

Survivorship bias: testing only stocks that still trade today. The ones that were delisted, acquired, or suspended after collapsing are not in your list — and those are exactly where a strategy steps on landmines. Your universe is "the survivors," so the results come out flattering. Eliminating this completely is hard, but at minimum know that your backtest is biased in that direction, and do not treat it as a conservative estimate.

Small samples are noise

Suppose your rule fired only six times in the past two years and won once — a 17% win rate. Does that prove it is a bad strategy? No. One extra win and the rate is 33%; one extra loss and it is 0%. With a sample that small, a single random event flips the whole picture.

By the same logic, five wins out of six does not prove it is good either. As a rough guide, a few dozen trades show you the general direction; comparing parameter settings meaningfully takes hundreds. If your rule is so strict it fires a handful of times in several years, either loosen the conditions, widen the universe, or accept that this rule cannot currently be validated.

Out-of-sample checking: leave a slice untouched

This is the most practical defense against overfitting. Split the history by time: tune your parameters on the earlier part (say 2019–2023) and do not look at the later part (2024 to now) until you are done. Once parameters are locked, run the rule on the later slice once.

  • If out-of-sample win rate and profit factor land in the same neighborhood as in-sample, the rule probably caught something real.
  • If in-sample looks beautiful and out-of-sample falls apart, you fitted noise — go back and remove conditions, do not add more.

Another split is by market: take parameters tuned on US stocks and run them unchanged on Hong Kong or China A-shares. A rule that holds up across markets is usually more robust.

How to set this up in Stock Compass

Backtesting in Stock Compass is a Pro feature. It runs your own buy and exit rules over daily bars, on up to 500 tickers at a time, and reports win rate, profit factor, max drawdown, and a trade-by-trade list. It does not connect to a broker, does not place orders, and does not recommend stocks — it only runs your rules.

Here is an example strategy you can backtest as-is:

{
  "name": "20-day breakout + trend filter",
  "market": "us",
  "rules": {
    "buy": {
      "v": 2,
      "outerOp": "OR",
      "groups": [
        {
          "innerOp": "AND",
          "conditions": [
            { "indicator": "is_20d_breakout", "operator": "==", "value": 1 },
            { "indicator": "adx_14", "operator": ">", "value": 25 },
            { "indicator": "vol_vs_avg20d", "operator": ">", "value": 1.5 }
          ]
        }
      ]
    },
    "add": { "v": 2, "outerOp": "OR", "groups": [] },
    "trim": { "v": 2, "outerOp": "OR", "groups": [] },
    "exit": {
      "v": 2,
      "outerOp": "OR",
      "groups": [
        { "innerOp": "AND", "conditions": [ { "indicator": "floating_loss_pct", "operator": ">", "value": 5 } ] },
        { "innerOp": "AND", "conditions": [ { "indicator": "floating_gain_pct", "operator": ">", "value": 15 } ] },
        { "innerOp": "AND", "conditions": [ { "indicator": "holding_days", "operator": ">", "value": 40 } ] }
      ]
    }
  }
}

Each condition in one line:

  • is_20d_breakout == 1: today's close is a new 20-day high.
  • adx_14 > 25: ADX above 25 means a trend exists, so this is not a fake breakout inside a range (see filtering false breakouts with ADX).
  • vol_vs_avg20d > 1.5: volume ratio above 1.5x — the breakout has volume behind it.
  • floating_loss_pct > 5: stop out once the position is down more than 5%.
  • floating_gain_pct > 15: take profit once it is up more than 15%.
  • holding_days > 40: if neither of the above has triggered after 40 trading days, time is up — exit.

The three exit groups are joined by OR: any one of them closes the trade. When the results come back, check the trade count first, then whether expectancy after costs is positive, and finally whether the max drawdown is one you could sit through.

Separately, every live signal you receive is logged in signal history with its actual T+1, T+5, and T+10 move (gross of costs). That works as a continuously running out-of-sample check — once a rule goes live, its record on the real future accumulates on its own.

Common mistakes

  • Reading only total return. 40% a year with a 60% drawdown is a strategy you will not hold to recovery. Read all four numbers.
  • Forgetting costs. The example above shows a 1.5% round trip turning positive expectancy negative.
  • Filling on the signal day. Confirming on the close and filling at that day's price is the most common lookahead bias.
  • Concluding from tiny samples. Six trades tell you nothing, whether it is 1 of 6 or 5 of 6.
  • Re-tuning until it looks good. Once you tune on the out-of-sample slice too, it is not out-of-sample anymore.

Summary

Backtesting is not about finding a "perfect strategy." It is about knowing a rule's temperament before you bet on it: how often it fires, how often it wins, how deep it bleeds, and whether anything is left after costs. Write the rules precisely, price in costs, leave a slice of data untouched, and do not conclude until the trade count is there. Validate first, execute second.

FAQ

What win rate counts as good?

Win rate alone is meaningless. A 40% win rate with wins twice the size of losses is profitable; a 70% win rate with losses three times the size of wins is not. Read win rate together with profit factor, expectancy after costs, and max drawdown.

If it profits in a backtest, will it profit live?

Not necessarily. Backtests carry traps — slippage, fees, survivorship bias, lookahead bias, overfitting — and every one of them pushes results in the flattering direction. A backtest helps you rule out clearly bad rules and understand how a rule behaves; it is not a forecast of future returns.

How much historical data does a backtest need?

Enough to include at least one full up-and-down cycle, ideally a real bear market, and enough trades to matter — a few dozen at minimum, hundreds if you want to compare parameter settings. If your rule only fires a handful of times over the whole period, the result is noise regardless of how many years it covers.

How do I know if I have overfitted?

Split your data by time, tune on the earlier part only, then run once on the later part you never looked at. If the out-of-sample numbers collapse compared to in-sample, you fitted noise. Rules with many precisely tuned conditions are the usual culprits; fix them by removing conditions, not adding more.

What is the difference between backtesting and signal history in Stock Compass?

Backtesting (a Pro feature) runs your buy and exit rules over past daily bars on up to 500 tickers and reports win rate, profit factor, and max drawdown. Signal history records every live signal your rules produce going forward and scores it at T+1, T+5, and T+10, gross of costs — a forward-running check rather than a look back. Neither connects to a broker or places trades.