Ask a trader whether they overtrade and you'll get one of two answers. Either "no, I'm selective," or "probably, yeah" — said the way people admit they should drink more water. Neither answer is worth anything, because neither is a measurement.
The problem is the word. Overtrading sounds like a volume complaint, so traders defend themselves with volume: eight trades is a lot, three is disciplined. But the number by itself is meaningless. A trader taking eight A-grade setups on a day the market handed out eight is not overtrading. A trader taking a third trade at 2pm because the first two were flat and the screen is right there almost certainly is.
There is a version of the question that can actually be answered, and your journal already holds the data. Not "how many trades is too many," but: at which trade of the session does your expectancy stop being positive? That has a number attached, the number is different for every trader, and finding it takes about half an hour.
Key takeaways
- Sort your closed trades by their position in the session — 1st of the day, 2nd, 3rd — and compute expectancy per bucket. In the worked 400-trade journal below, the first two trades of the day returned +72.4R and everything after them gave back −24.9R, or 34.5% of the gross.
- Costs move the cutoff a trade earlier than the raw numbers suggest. The third trade is +0.13R gross and +0.02R net — fees, funding, and slippage take essentially all of it before direction is involved.
- The late trade doesn't underperform because it's late. It underperforms because 55% of fourth trades were C-grade setups, against 10% of first trades. Frequency is the symptom; selection is the disease.
- The finding is suggestive, not established. That −24.9R sits 1.61 standard errors from zero — short of the conventional bar. The decision is still easy, because the asymmetry isn't: capping costs about 2.8R if the effect is noise and saves about 24.9R if it's real.
- A hard count cap is the wrong fix. Roughly 25 of those 142 late trades were graded A or better, and a "stop after two" rule throws them out with the rest. Gate on grade, not on count.
Overtrading is a selection problem wearing a frequency costume
Every real definition of overtrading reduces to the same thing: taking trades that don't meet the bar you set when you weren't in the market. Frequency is how it shows up in the data, not what it is. That distinction matters because it decides what you do about it.
If overtrading were genuinely about volume, the fix would be a cap — three trades a day, done. But traders who impose caps usually find one of two things. Either the cap holds and they feel disciplined while quietly skipping their best afternoon setup, or the cap breaks the first time a clean opportunity shows up at trade four, and once it has broken once it never binds again. Both outcomes come from the same mistake: capping the count when the thing that decayed was the standard.
Three mechanisms do the decaying, and they compound:
- The bar drops as the session ages. A setup you'd have skipped at 9am starts to look tradeable at 2pm, because the alternative is a day with nothing in it. Nobody experiences this as lowering their standard. It is experienced as the setup improving.
- Position in the session correlates with emotional state. The fourth trade is disproportionately taken by someone who is down on the day, up on the day and pressing, or bored. Each of those is a different distortion, and none of them is present on the opening trade. When the trigger is specifically a loss, the mechanism has its own name and its own fix — see revenge trading.
- Costs stay fixed while edge shrinks. Fees, funding, and slippage don't know it's your fifth trade. They take the same toll off a marginal setup as off your best one, which means the marginal trade is the one where costs are most likely to exceed the edge entirely.
A worked journal: 400 trades, 140 sessions
Here is a six-month journal, 400 closed trades across 140 sessions, sorted by each trade's position within its own session. Expectancy is net of fees, funding, and slippage, stated in R — multiples of the amount risked — so trades of different sizes compare directly.
| Trade of session | Trades | Net expectancy | Total R |
|---|---|---|---|
| 1st | 140 | +0.34R | +47.6R |
| 2nd | 118 | +0.21R | +24.8R |
| 3rd | 78 | +0.02R | +1.6R |
| 4th | 38 | −0.28R | −10.6R |
| 5th or later | 26 | −0.61R | −15.9R |
| All | 400 | +0.12R | +47.4R |
The account made money. The trader is profitable, the equity curve goes up and to the right, and nothing on a monthly P&L statement would prompt a second look. Underneath, two businesses are running in the same account. The first two trades of each session earned +72.4R. Everything from the third trade on returned −24.9R — handing back 34.5% of what the good half produced.
Per trade, the gap is starker than the totals: +0.28R across the first two, +0.12R across all 400. The extra 142 trades didn't add a smaller edge to the pile. They diluted a good one by more than half.
Costs move the cutoff one trade earlier
The third-trade bucket is the interesting one, because on a gross basis it still looks fine. Split each bucket into what the market gave and what execution took back:
| Trade of session | Gross expectancy | Costs | Net expectancy |
|---|---|---|---|
| 1st | +0.43R | −0.09R | +0.34R |
| 2nd | +0.31R | −0.10R | +0.21R |
| 3rd | +0.13R | −0.11R | +0.02R |
| 4th | −0.15R | −0.13R | −0.28R |
| 5th or later | −0.46R | −0.15R | −0.61R |
Gross, the edge survives through the third trade and turns negative at the fourth. Net, the third trade is already a rounding error — 78 trades of screen time, risk, and attention for +1.6R total. Costs pull the honest cutoff forward by a full trade, and a trader looking only at gross P&L will defend a bucket that is, after execution, doing nothing.
Notice too that the cost column climbs as the session goes on: 0.09R on the opener, 0.15R on the fifth. That isn't a coincidence and it isn't fees changing. Late trades tend to be tighter-stopped and more impatient — and cost measured in R rises as your stop distance shrinks. A five-basis-point fill is trivial against a 1% stop and expensive against a 0.2% one. The marginal trade is both the least profitable and the most expensive to execute, which is a bad combination in one place.
The mechanism: grade, not fatigue
It's tempting to read the decay as tiredness, and tiredness is real. But the journal has a field that tests it directly. If each trade carries the setup grade it was assigned at entry — the A+/A/B/C scale popularized by Lance Breitstein — then the grade mix by trade number tells you whether the trader is executing the same setups worse, or executing worse setups.
| Trade of session | A+/A | B | C |
|---|---|---|---|
| 1st | 61% | 29% | 10% |
| 2nd | 44% | 37% | 19% |
| 3rd | 23% | 40% | 37% |
| 4th | 12% | 33% | 55% |
| 5th or later | 8% | 25% | 67% |
The answer is not subtle. The first trade of the day is an A-grade setup 61% of the time; the fourth is a C-grade setup 55% of the time. The trader hasn't gotten worse at trading over the course of the afternoon. They have gotten worse at declining, and the expectancy table is just the grade table with money attached.
This reframes the whole exercise. If the fourth trade underperformed while holding grade constant, the honest conclusion would be fatigue and the fix would be a stopping rule. Because grade collapses instead, the fix is a standard — and that requires that your grades mean something in the first place, which is a playbook problem, not a discipline one. Grades assigned after the fact, or assigned to describe how a trade felt, will produce a flat table and tell you nothing.
"What's my expectancy by trade number of the session — and does my setup grade drop as the day goes on?"
No article can answer that, this one included. The cutoff is a fact about your trades, and the honest range across traders is wide: some journals bend at trade three, some run clean to six, and a few show the opposite pattern because the trader's real setup only appears once a session is underway. A coach grounded in your own journal can bucket the trades, subtract the costs attached to each one, cross the result against the grade you assigned at entry, and tell you which of those three you are. A general chatbot can only restate the concept, because your fills and your grades are the inputs it doesn't have.
What the sample size will and won't support
The temptation at this point is to treat −0.28R on the fourth trade as a finding. It isn't one. That bucket holds 38 trades, and with a typical R-multiple spread the 95% interval around it runs roughly ±0.41R — wide enough to contain zero comfortably. Read alone, the fourth-trade number is barely evidence at all.
Grouping helps but doesn't rescue it. Pooled, the 142 trades from the third onward average −0.176R with a standard error near 0.109R — about 1.61 standard errors from zero, short of the conventional 1.96 bar. The right sentence is "suggestive, and consistent with a mechanism I can see in the grade data," not "proven." Splitting a journal into buckets is exactly the operation that destroys the sample you need, and five buckets is a lot of splitting for 400 trades.
The decision survives the uncertainty anyway, because the payoff is lopsided. Suppose the late trades are genuinely neutral — the +0.02R the third bucket showed. Capping then costs about 2.8R over 142 trades, roughly one good trade's worth, spread across six months. If instead the pooled −0.176R is real, capping saves about 24.9R. You are risking a rounding error against half a year of leaked edge. You don't need statistical significance to take a bet priced like that — you need it before you claim the effect as a fact, which is a different question.
Five ways this will mislead you
Trade-number analysis is cheap to run and easy to over-read. Each of these has cost someone money.
- Trade number is a proxy for four things at once. Position in the session is entangled with time of day, market state, your P&L at the moment of entry, and how long you'd already been at the screen. A clean-looking decay curve may be a time-of-day effect wearing a different label. The grade column is what separates them; without it you have a correlation and a story.
- Opportunity isn't evenly distributed across days. A trending session genuinely offers more valid setups than a chopping one, so the five-trade days are partly selected for being good days. That biases the late buckets upward, not down — which makes a negative late bucket more believable, and makes a mildly positive one much weaker evidence than it looks.
- A hard count cap discards your late winners. About 25 of those 142 post-second trades carried an A grade. A "stop after two" rule deletes them along with the C-grade filler, and if your best setup happens to be a late-session one, a count cap is the most expensive rule you could write.
- The cap changes the behavior it's measuring. The 52.6% uplift a "first two trades only" counterfactual produces on this data is arithmetic, not a forecast. Knowing you have two shots changes which two you take — sometimes for the better, often by making you fire the first one early so you don't waste the allowance. Treat the counterfactual as a size estimate of the problem, never as a projected return.
- Buckets get thin fast. The fifth-plus bucket in this journal holds 26 trades, several of which are probably the same bad Tuesday. One catastrophic session can create an entire "finding." Check how many distinct days feed each bucket, not just how many trades.
Where to start
Pull your last 200 to 400 closed trades and add one column: the trade's ordinal position within its session, counting from the first entry of that day. It's a sort and a counter — twenty minutes in a spreadsheet, and most journals can export the timestamps you need.
Then compute two things per bucket, and only two: net expectancy in R, and the share of trades graded A or better. If both decline together, you have a selection problem and the mechanism is visible. If expectancy declines while grade holds flat, you have an execution or fatigue problem and the fix is different — shorter sessions, or a break rather than a bar. If neither declines, you are not overtrading and you can stop worrying about a number that was never yours.
When the curve does bend, write the rule as a standard rather than a ceiling: after the second trade of the session, A-grade setups only. That keeps the late winners, removes the filler, and has the useful property of being falsifiable — in three months you can check whether the trades it let through actually earned their place. A count cap can't be checked, because it never lets you see what it excluded. The point of the exercise isn't to trade less. It's to stop paying full price for the trades that were never worth taking — which, like expectancy itself, is a question of arithmetic long before it becomes a question of willpower.