Do Gamma Walls Work? Testing the Gamma Wall Without Fooling Yourself
Scope: measured on the futures contracts listed (ES, NQ, RTY, YM, GC, SI). Findings apply to the contracts measured and do not automatically transfer to ETFs or other instruments.
A call wall sitting a few points above spot gets touched on almost any ordinary day, so a touch rate measures distance, not the strike. A real test needs two things a published hit rate almost never carries: a control level the same distance away, and a prediction the gamma mechanism actually makes about behaviour on arrival. The short answer: a raw hit rate, however high, is not evidence either way, and you can check any published one yourself in an afternoon with the control described below. If a level on your screen tomorrow comes with a hit rate, this is how to know whether that number told you anything before you trade around it.
All 26 research notes
01Claims retail traders inherit — tested
Do gaps get filled? What a volume spike means Do volatile days offer more? Calendar effects, tested Reversal stories need controls What moved Bitcoin in August Stop hunts, tested The VWAP magnet, tested02Before you read any indicator
When futures actually trade Best time of day to buy an ETF How much does SPY move? Premarket and after hours The gap before you see it ETFs vs futures SPX vs SPY vs ES Do SPY and QQQ move together?03The account is a variable too
Why accounts blow up The account is a variable04How the levels are computed
GEX: open interest vs volume Why platforms disagree on GEX05How much the levels move
The gamma flip moves all day Call and put walls, explained06Whether the levels carry information
Testing the gamma wall Point of control, tested07What the executed trades add
One market, many tapes Order flow + gamma confluenceWhy does "how often is the wall touched" answer the wrong question?#
Because touch rate is mostly a function of distance. Any level inside a typical session's range is reached by ordinary movement, whatever is written on it. A number sitting near the top of its scale cannot discriminate between a strike that mattered and a strike that merely sat close to price.
The common test takes the wall at the close, asks whether the next session's high-low range contains it, and publishes the fraction. Three things go wrong. The level comes from that session's own chain, so a same-day touch is partly guaranteed by construction and must be excluded. The metric saturates: when nearly every candidate is touched, nothing separates one from another. And nothing is compared, so the reader cannot tell whether the wall beat a nearby number.
What does a mirror control look like?#
For every wall, place a phantom level at the same absolute distance on the opposite side of spot, or pick an unremarkable strike the same distance away on the same side, and score it with the identical rule on the identical sessions. The wall's own rate is meaningless on its own. Only the gap between the two carries information. If all you have is a screenshot, take the wall's distance from that day's close, pick the nearest round strike the same distance away, and count both over the same sessions.
Build the control carefully. Score both levels as a matched pair within each session, so the comparison is paired rather than two samples drawn from different days. Apply exclusions symmetrically: drop a session's control whenever the wall is unscoreable. The mirror inherits that day's volatility, which is the point, since touch odds depend more on range than on the label.
The mirror is not perfect. Index price distributions are not symmetric around spot, so levels equidistant above and below do not face identical odds. The cleaner alternative is a same-side control: an arbitrary strike at matched distance, chosen without reference to open interest. Report both differences rather than defending one.
What does the gamma mechanism actually predict?#
Not arrival. The theory says hedging near a strike where dealers (the market makers on the other side of most option trades) are net long options runs counter to the move and damps it, while hedging near a strike where they are net short runs with the move and amplifies it. That is a claim about the character of movement, not about whether price shows up.
So the second test measures behaviour, conditioned on the assumed sign of dealer positioning. Three quantities matter, each scored against the same control level: realized range in the window before first touch, reversal frequency within a fixed interval after touch, and the signed slope of price over that interval. A long-gamma assumption predicts compression and more reversals than the control; a short-gamma assumption predicts continuation. Counting touches cannot distinguish the two, because both permit arrival.
Why is the positioning sign the fragile input?#
Open interest says how many contracts are open at a strike, never who is long and who is short. Every wall statistic rests on an assumed sign, usually that investors sell upside calls and buy downside puts. When that assumption is wrong, predicted damping becomes acceleration, and a correct mechanism scores as a failure.
The proxies used to infer that sign are indirect: whether trading at the strike has been predominantly buyer- or seller-initiated, and what built the pile. Those are estimates stacked on the estimate being tested. A result reported without stating the sign rule is a joint test of mechanism and proxy, and a joint test cannot say which half failed. The honest form reports the behavioural result separately under each sign assumption. Background on why the sign is inferred rather than observed is in call wall and put wall.
| Test 1: touch rate | Test 2: behaviour on approach | |
|---|---|---|
| Question asked | Did price arrive? | Did movement change character near the strike? |
| Control needed | Phantom level, matched distance, same sessions | The same, plus grouping by assumed positioning sign |
| Fails when | The level sits near price, so almost everything is touched | The sign assumption is wrong, or the sample is too thin to split |
| Evidence value | The difference from control, never the raw rate | The conditioned difference in behaviour |
What would settle this, and what would not?#
A single month cannot. Splitting sessions by positioning sign, then by instrument, then by distance bucket, divides an already short record into groups too small to separate a real effect from noise. Settling it takes years of sessions per instrument, rules fixed before the data is scored, and a result published whichever way it lands.
Two further constraints are structural. The expected effect is small, because hedging is one flow among many trading the same underlying, and small effects need large samples. Results also do not transfer: regimes imply different hedging density, expiration cycles change strike concentration, and one index future says little about another. Pooling across regimes for a bigger sample answers a question nobody asked.
What should you demand from anyone quoting a wall hit rate?#
Ask for the control level and the rule that generated it. Ask for the window, the session count, and whether same-day touches were excluded. Ask how the wall was defined, since different weightings name different strikes on the same afternoon (why the numbers differ, open interest versus volume). Ask which sign assumption was used, and whether the test was specified before the results were seen. A hit rate without those is a number missing its denominator.
Limits#
- A touch test ignores path. One session grinding into a level and another gapping through it score identically and mean different things.
- The behavioural test needs intraday data aligned to the chain timestamp; daily bars cannot measure compression on approach.
- The wall moves during the session, so "the wall" must be timestamped before it can be tested at all (intraday relocation).
- None of this establishes causation. A surviving difference from control is association, consistent with the mechanism and with whatever else clusters at round and popular strikes.
- A null result from a short sample usually means the design lacked power. It is not, by itself, evidence that the mechanism is absent.
Related reading#
- Call wall and put wall — where the sign assumption comes from
- Gamma flip moves — the flip level needs the same timestamp discipline
- Why GEX numbers differ — how the definition changes which strike is called the wall
- GEX from open interest versus volume — two inputs, two different maps to test
- Volume profile POC hit rate — the same control design on a level where it discriminates
- Stop hunt pattern, tested — another popular claim put through a control
See these levels on a live chart
Option-derived levels and futures tape on one timeline.