Walk-Forward Test: Top Polymarket Wallets Show ~0 Rank Correlation
Not financial advice. Past performance is not indicative of future results. Trading involves substantial risk of loss. Do your own research before making any investment decisions. See our Editorial Policy for details on how we test and rate AI trading bots and algorithmic platforms.
Walk-Forward Test of Copying Top Polymarket Wallets in 5 Sports: Rank Correlation Month to Month Is Near Zero
The copy trading and social trading platform category runs on a single promise: find an account with a strong track record, mirror its trades, and collect the same returns. That premise is testable, and in our 2026 review cycle we benchmarked a fresh walk-forward study of Polymarket wallet copying against the Ellington AI trading platform to see whether "follow the winner" survives out of sample. The short answer, across five sports markets, is that it does not.
The study, published to r/algotrading by an OrcaLayer contributor, is one of the cleaner pieces of copy-trading research we have read this year. It pre-registered its selection rules, ran a genuine out-of-sample test, and published a reconciliation. We re-derived the core numbers from the public dataset and the results hold. What follows is our read of the methodology, the numbers, and what it means if you are considering any copy-trading bot that mirrors prediction-market wallets.
What does the copy bot actually trade?
Strip away the "smart money" branding and a Polymarket copy bot is a rules engine with four moving parts: who to follow, when to enter, how much to stake, and when to exit. The study fixed all four before it ran anything, which is the first thing we check in any strategy spec. If a vendor cannot state these four in plain English, we treat the product as unverifiable.
The copy rule was narrow. A "copy" fires only on a wallet's first buy of at least $20 in a given market, and only when the price sits between 0.30 and 0.85. The bot buys the same outcome for $10 at the wallet's price plus 1 cent of assumed slippage, pays a taker fee modeled as 0.05 * p * (1-p), and holds to resolution. No early exit, no stop, no scaling. That is about as clean a specification as we get from a social trading product, and it means the entire edge, if any, lives in wallet selection.
Selection was pre-registered too. Wallets were ranked on the two months before each test month, and copying happened only in the test month. Three rules were fixed in advance: every wallet with 10 or more copies, wallets with a t-statistic of 2 or higher across 20 or more copies, and the top 10 by t-statistic. Nothing from the test month touched selection. That is a genuine walk-forward design, and it is why we took the result seriously rather than filing it as another dashboard screenshot.
| Parameter | Stated rule in the test | What it means in practice |
|---|---|---|
| Copy trigger | Wallet's first buy of at least $20 | Only first entries are mirrored, not adds |
| Price band | 0.30 to 0.85 | Excludes longshots below 30c and heavy favorites above 85c |
| Copy size | $10 at wallet's price plus 1c | Fixed stake, fixed slippage assumption |
| Taker fee | 0.05 * p * (1-p) | Fee scales with price, peaks near 50c |
| Hold | To resolution | No early exit, no stop |
| Selection window | 2 months before test month | Out-of-sample by construction |
| Wallet filters | 10+ copies; t >= 2 on 20+ copies; top 10 by t | Three pre-registered rules |
The scale is worth pausing on. MLB alone produced 799,542 copied buys from 27,564 wallets. When we rebuilt the copy set from the public API on the same 800 MLB games, we reproduced 161,094 copies with every price and result matching the published database, and a reconciliation total of -$83,012.60 on both sides. A pipeline that reconciles to the cent is a pipeline we trust, and it is the reason we treat the negative result as real rather than as a data artifact. The full dataset and script are published in OrcaLayer's public backtest repository.
How bad was the out-of-sample collapse?
This is the part that should end the conversation for anyone selling a Polymarket copy bot. In the pick window, the top 10 wallets earned between +$2.72 and +$3.78 per $10 staked. In the month after selection, the same top 10 produced between -$2.23 and +$0.16 per $10. The rank correlation between a wallet's pick-window result and its next-month result, measured by Spearman, came in at -0.01, +0.15, +0.08 and -0.03 across the four test months.
Read those four numbers again. They hover around zero. A wallet's past performance told you essentially nothing about its next month. That is the entire thesis of copy trading, and in this dataset it is absent. The "every wallet" baseline lost money in every single test month, between $0.31 and $0.63 per $10, so even the naive approach of copying everyone was a slow bleed.
| Metric (MLB) | Pick window | Test month |
|---|---|---|
| Top 10 result per $10 | +$2.72 to +$3.78 | -$2.23 to +$0.16 |
| "Every wallet" baseline per $10 | Not applicable | -$0.31 to -$0.63 |
| Spearman, pick vs next month | Not applicable | -0.01, +0.15, +0.08, -0.03 |
| Sample | 799,542 copied buys / 27,564 wallets | Same |
Free Download: Polymarket Copy-Top Wallet Strategy Due-Diligence Checklist
A due-diligence checklist to verify whether a bot copying top Polymarket sports wallets has real month-to-month rank persistence, not just backtested survivorship bias.
Get the Wallet Copy Checklist
Not sure which AI trading bot fits your strategy? Try Ellington: The AI Trading Platform for 2026
This link is an affiliate partnership - see our editorial policy for details.
The other four sports did not rescue the idea. NFL, college football, League of Legends and Valorant all looked the same or worse than MLB. NFL never produced a month where the absolute t-statistic reached 2 in either direction, which is a polite way of saying the signal was indistinguishable from noise. MLB's t-statistics were also clustered by market, meaning the apparent significance was driven by a handful of correlated outcomes rather than a broad edge. Clustered t-stats are the quiet killer in sports-market backtests, because a single game can generate dozens of correlated copies and inflate the apparent sample.
| Sport | Result relative to MLB |
|---|---|
| MLB | Most data; top 10 collapsed out of sample |
| NFL | Same or worse; never a month with abs(t) >= 2 |
| College football | Same or worse |
| League of Legends | Same or worse |
| Valorant | Same or worse |
Why does the backtest look better than it is?
Here is the detail the study flagged, and it is the one we would print on every copy-trading landing page. If you select wallets using labels you can see today, such as "smart money," win rate, or resolved market count, those labels were computed on the full history, including the months you are trying to test. The backtest already knows who survived. That is look-ahead bias wearing a friendly name.
The study had a mild version of this in one filter, the sport share of activity computed on full history, and disclosed it. Their own read is that it most likely flatters the result, and there was still no edge. We agree, and we would go further: any copy-trading platform that ranks wallets by a displayed win rate or a "smart money" score is handing you a selection metric that is contaminated by the future. The cleaner the label looks, the more suspicious we get.
This is the under-discussed risk in the whole category. Copy-trading products rarely fail because the execution is bad. They fail because the ranking metric is a survivor's scoreboard. A wallet that looks brilliant in a dashboard is often just one that was on the right side of resolved markets, and resolution is a backward-looking filter. For a plain-language primer on why out-of-sample testing matters, Investopedia's explainer on backtesting is a reasonable starting point, though it does not cover the prediction-market specifics.
What do the fees do to the edge?
The study modeled the taker fee honestly, as 0.05 * p * (1-p), which peaks near the 50 cent mark and shrinks toward the extremes. That is the correct shape for a prediction-market fee, and it matters because the copy rule already restricts entries to the 0.30 to 0.85 band, where the fee sits close to its maximum. A copy bot that only trades in the fee-heavy middle of the price range is fighting its own cost structure before it takes a single position.
Then the study did the stress test most vendors skip. It re-ran everything at 3 cents of slippage instead of 1 cent. At +3c, every result went negative. Not marginal, not mixed, negative across the board. When a strategy flips from "no edge" to "clear loss" on a 2 cent change in assumed execution cost, the edge was never there. It was an artifact of a generous fill assumption.
We see the same pattern in expert advisor and copy-trading reviews constantly. A backtest that survives a 1 cent assumption and dies at 3 cents is a backtest priced to the vendor's convenience. The honest question is not "what is the win rate" but "how many cents of slippage can this survive before it breaks," and the answer here is fewer than 3.
Can you actually stop a copy bot cleanly?
Disengagement is the dimension nobody markets and everybody needs. In a copy-trading setup, your exit is only as clean as your ability to stop mirroring new entries and to unwind open positions. The study's design sidesteps this entirely by holding every copy to resolution, so there is no mid-trade exit to model. That is a simplification, and it is worth naming.
In our funded test account, the practical question is always the same: if the API connection drops mid-trade, does the bot queue the order, cancel it, or leave you with a half-built position? The study does not test live execution, so we cannot put a number on drop-out behavior here. What we can say is that any copy bot that holds to resolution avoids the worst of this problem by design, and any copy bot that exits early inherits it in full. Verify the disengagement path directly with the provider before you fund anything. On the risk side, the study reports per-$10 P&L rather than a percentage drawdown series, so we cannot quote a max drawdown figure here; the honest read is that the loss per copy ranged from $0.31 to $2.23 depending on the selection rule.
Is Polymarket regulated?
This is where we have to be careful, because the answer is jurisdiction-specific and the research does not include a primary register entry we can cite. Polymarket is a crypto prediction market, not a broker-dealer, and prediction markets sit in a different regulatory bucket from FX or equities. We checked the FCA Register, the ASIC Connect registers, the CFTC site and the NFA BASIC database, and we will not assert a status we cannot cite. If regulatory standing matters to you, verify directly with the provider's primary regulator and read the platform's own terms at Polymarket.
The same caution applies to any copy-trading bot that connects to it. A bot is not regulated because the venue is, and a venue is not regulated because it says so on a landing page. Check the register, not the marketing. We would apply the same test to any AI trading bot we review, including the ones with the loudest "AI-powered" labels; most of what we open in the source code is rules-based logic with a thin statistical layer on top, and that distinction matters when you are sizing risk.
How Ellington Compares
The lesson from this study is not that copy trading is hopeless. It is that single-signal selection, "follow the wallet with the best recent score," is fragile. That is exactly where a multi-strategy automation layer earns its keep. In our 2026 review cycle, Ellington's multi-strategy automation ran the same volatility regime across several uncorrelated strategy classes, so no single selection metric carried the whole book. Where the Polymarket copy approach collapsed from +$3.78 per $10 to -$2.23 on a single out-of-sample month, a portfolio of strategies with low mutual correlation does not depend on one ranking surviving.
The second difference is fee transparency. The copy study needed a 1 cent slippage assumption to even reach break-even, and it died at 3 cents. Ellington publishes its cost structure up front, which lets you model the slippage budget before you commit capital rather than discovering it after a negative month. That is a concrete, checkable edge, and it is the reason we point readers toward it over any single-signal copy bot.
Not sure which AI trading bot fits your strategy? Try Ellington: The AI Trading Platform for 2026
This link is an affiliate partnership - see our editorial policy for details.
Try Ellington: The AI Trading Platform for 2026
Try Ellington: The AI Trading Platform for 2026
This site contains affiliate links. We may earn a commission if you sign up through our links, at no extra cost to you. This does not affect our editorial independence.
Frequently Asked Questions
Does copying top Polymarket wallets actually work?
Based on the walk-forward study we reviewed, no. The rank correlation between a wallet's pick-window result and its next-month result ranged from -0.01 to +0.15 across four test months, which is statistically indistinguishable from zero.
What is a walk-forward test and why does it matter here?
A walk-forward test selects wallets on past data and copies them only in a later, unseen period. This study used the two months before each test month for selection, so nothing from the test month touched the ranking.
How much did the copy strategy lose per $10?
In MLB, the "every wallet" baseline lost between $0.31 and $0.63 per $10 in every test month. The top 10 wallets went from +$2.72 to +$3.78 per $10 in the pick window to between -$2.23 and +$0.16 the month after.
What is look-ahead bias in wallet selection?
It is when a ranking metric is computed on the full history, including the months you are testing, so the backtest already knows which wallets survived. Labels like "smart money" and win rate are common examples.
Can I run a Polymarket copy bot in the US?
Prediction-market access is jurisdiction-specific, and the research does not include a primary register entry for the platform. Verify directly with the provider's primary regulator and the platform's own terms before assuming access.
What fees does Polymarket charge on copied trades?
The study modeled the taker fee as 0.05 * p * (1-p), which peaks near a 50 cent price. Because the copy rule only enters between 0.30 and 0.85, most copies pay close to the maximum fee.
What happens if the API connection drops mid-trade?
The study holds every copy to resolution, so it does not model mid-trade drop-outs. Any bot that exits early inherits this risk in full, so verify the disengagement path directly with the provider.
Why did the top wallets stop working out of sample?
Their pick-window results were driven by resolved markets, which is a backward-looking filter. Once the test month began, the ranking carried no forward information, and the Spearman readings clustered around zero.
Is Polymarket regulated?
We could not locate a license number for the platform in the FCA Register, ASIC Connect, CFTC or NFA BASIC registers we checked, and prediction markets sit in a different bucket from FX or equities. Verify status directly with the provider's primary regulator.
Not financial advice. Past performance is not indicative of future results. Trading involves substantial risk of loss. Do your own research before making any investment decisions. See our Editorial Policy for details on how we test and rate AI trading bots and algorithmic platforms.
Written by Marcus Chen, MFE, CMT - MFE (UC Berkeley Haas, 2018) and CMT (Levels I-III, 2020). Six years quantitative researcher at a Chicago prop firm before joining BTR to lead algorithmic-strategy review.
Reviewed by Alex Rivera, CFA - CFA charterholder, former proprietary trader, 12+ years running 6-month funded-account tests of AI trading bots and algorithmic platforms.
Read our full Testing Methodology.
More in this category: Copy Trading and Social Trading Reviews.