Data Findings

Ten months of watching Polymarket

I have been polling every tradeable market on Polymarket since October 2025 and keeping what I saw. Here is what the data says about the exchange.

Gent Blaku·x.com/_GentB·linkedin.com/in/gentblaku

Price moves detected     144,785   23 Oct 2025 onward, 155 days observed
Distinct markets moved    14,301
Daily catalogue census    37,419   every listed market, snapshotted daily
Raw quote samples        394,723   bid, ask and spread, not just the mid
Outcomes at settlement     3,497   captured as markets close

A threshold detector records volatility events, not trades. Absolute rates from it are unreliable. Comparisons made within it, where both sides passed the same instrument, largely cancel that bias, and every finding below is the second kind unless it says otherwise.

Finding 01

Volume concentrates. Volatility does not.

Held byShare of volumeShare of price moves
Top 10 markets11.1%2.1%
Top 100 markets58.0%11.1%

Money pools into a handful of markets. Price discovery does not. The 100 busiest markets carry 58% of all volume on the platform but only 11% of the repricings, and the same roughly fivefold gap holds at the top ten. Meanwhile 39.1% of the 14,135 markets that ever moved, moved exactly once, and 83.9% of the catalogue has traded under $10,000 in its life.

These are two very different pictures of the same exchange, and only one of them is visible from a volume leaderboard. The long tail is where the platform is doing its actual work of turning news into a number.

Volume comes from the daily census and moves from the threshold detector, so this is a comparison across two instruments rather than within one. A gap of that size is far larger than any plausible instrument bias, but it is not a like-for-like measurement and should not be quoted as one.

Finding 02

There are about 7,130 events, not 33,710 markets

Options in the eventEventsMarkets
1 (standalone binary)2,7612,761
2 to 52,3766,868
6 to 201,77517,712
21 to 502095,850
50 or more9519

Only 8% of listed markets are standalone binaries. Everything else is one option inside a larger event, and more than half the catalogue sits in events with six to twenty options. Counting markets rather than events overstates the real decision surface by roughly 4.7x.

Why it matters beyond the headline number. It explains why the largest markets are the hardest to price. A big event is a wide option field, and most legs of a wide field never develop a two-sided book, so "biggest market" and "most liquid market" are not the same thing and often are not even close.

Finding 03

The cheapest markets have the tightest books

Price bandAvg absolute spreadRejected by a 10%-of-mid filter
Under 10%1.1c82.6%
10 to 25%4.3c53.6%
25 to 50%6.4c36.3%
50 to 75%6.4c24.2%
75 to 100%2.8c6.9%

Measured across 150,000 quote samples. Markets trading under 10% carry the tightest absolute books on the exchange, averaging just over a cent wide. They are also the ones a conventional spread filter throws away, because dividing a fixed spread by a small mid inflates the ratio: an identical 2-cent book reads as 33% at a mid of 6% and as 4% at a mid of 50%.

This is a finding about the API's consumers, not its data. Any client filtering on spread as a percentage of mid is silently blind to more than four fifths of the low-probability range, which is precisely the range where a market moving is most informative. I know because my own bot did it for months.

Finding 04

Sports and news run on opposite clocks

Hour (UTC)Share of sports movesShare of news moves
01:0012.7%4.0%
07:000.4%6.1%
16:002.6%5.9%
23:005.3%1.9%

Sports volatility concentrates between 18:00 and 02:00 UTC, peaking around 01:00 with 12.7% of all sports moves, and all but disappears between 06:00 and 11:00, bottoming at 0.4%. That is a 29x swing between its busiest and quietest hour. News does the reverse: spread across European and US business hours, quietest overnight.

The two classes are close to evenly split by move count, so neither is riding the other. They are genuinely two different exchanges sharing one venue, and any capacity, alerting or staffing decision made on a blended hourly average is being made on a curve that describes neither.

Finding 05

18.3% of markets listed as active have already ended

Of 33,710 listed activeCount
End date already passed6,164
Passed more than 90 days ago334
No end date at all1,370

Taken from Gamma's own fields on a full catalogue walk. "Active" does not mean "has not finished". A July baseball game was still listed as open in late August. Anyone building a "what can I trade right now" view inherits roughly one part in five of noise unless they know to re-filter on the end date themselves.

Finding 06

A quarter of the catalogue is sport, and the API will not tell you which quarter

ClassMarketsShare
News and everything else21,60864.1%
Sports8,62325.6%
Crypto2,4387.2%
Entertainment1,0413.1%

About 36% of what Polymarket lists is not news. That classification is mine, not theirs, and that is the point: Gamma returns an empty category for every sports market, so the split above had to be reconstructed from slug structure. The load-bearing rule is a fixture date in the slug, validated against the full catalogue with two false positives out of 8,675, both of them drought markets shaped like a fixture.

The single most repriced market of the ten months was "US forces in Venezuela by October 31?", with 288 separate repricings. The rest of the top eight are mechanical: Bitcoin and Ethereum price-bracket markets, and same-day football. Volatility and newsworthiness are different axes, which is why a threshold alone cannot decide what is worth reading.

Finding 07

An unquotable market is almost never a small market

Options in the eventMarketsReturns a two-sided quote
1 (standalone)2,07598.2%
2 to 45,69796.9%
5 to 1215,06191.0%
13 to 4010,28181.4%
Over 401,02682.7%

About 10.4% of live listed markets return no tradeable two-sided quote. The CLOB has no bid, no ask, or neither. Measured on seven consecutive daily censuses of the full catalogue, the share sits between 10.1% and 10.7% every day, so this is a property of the listings rather than of any one walk.

The intuition is that these are the dead long tail, and that is wrong. Quotability barely tracks volume at all, and what movement there is runs the wrong way:

Volume bandMarketsReturns a two-sided quote
Under $1k21,56888.1%
$1k to $10k7,18890.6%
$10k to $100k3,57894.2%
$100k to $1M1,25589.7%
Over $1M55184.8%

Markets that have traded under $1,000 are more likely to be quotable than markets that have traded over a million. Volume is not what decides it. Field size is, and it decides it monotonically. A standalone market is quotable 98.2% of the time. One option inside a 13-to-40-way field is quotable 81.4% of the time. The million-dollar events are disproportionately large fields, which is the whole of the dip in the second table.

This is a listing-quality observation, not a pricing one. Inert options sit beside live ones in the same field with nothing in the response distinguishing them. A client rendering a 30-way election field gets a full grid of rows back and has to discover on its own that a fifth of them cannot be traded. The failure is silent: an option with no book is not flagged, it simply has no price.

Markets whose end date has already passed but which are still listed as active, Finding 05's roughly one in five, are 46.2% unquotable, against 10.7% for genuinely live ones. That is expected, and it is why every figure above excludes them. Pooling the two produces a headline of 14.4%, which is mostly a statement about expired listings wearing an active label.

Seven daily snapshots, one per day, so this measures the catalogue as it stands rather than a trend. About 1.7% of the unquotable set is contamination rather than signal: sampling 60 of them against Gamma, one carried no CLOB token ID at all, which is a different defect that lands in the same bucket. The predecessor of this number was wrong. See below.

Finding 08

Finding 05's expired listings are a queue, and it drains at about 43% a day

DayListedLeft by next dayArrivedTurnover
25 Aug34,0233,8755,08511.4%
26 Aug35,2334,1664,68111.8%
27 Aug35,7483,7915,57910.6%
28 Aug37,5364,2056,00811.2%
29 Aug39,3395,9235,75415.1%
30 Aug39,1706,6704,91917.0%
31 Aug37,4195,6983,94115.2%
1 Sep35,6624,2614,62711.9%

Between a tenth and a sixth of the catalogue is replaced every day. Roughly five thousand markets leave and five thousand arrive, and the arrivals are new in every sense: their median traded volume is $41. The headline count of about 36,000 markets is not 36,000 durable things. It is a stable core of roughly 23,500 that appear on every single day, plus a daily tide of short-lived listings, most of them dated sports fixtures that exist for a day or two and settle.

What leaves is not random. Splitting each day's listings by whether the end date had already passed, and excluding same-day expiries from both sides so the boundary case cannot flatter either number:

State on the dayPopulationGone the next day
End date already passed2,503 to 3,76134.5% to 55.4%
Still running27,229 to 30,3901.4% to 4.2%

Averaged over the eight pairs that is 43.1% against 2.2%, a factor of twenty, and the gap holds on every single pair. Delisting is almost entirely explained by the market having finished. That is the mechanism behind Finding 05: expired listings are not an inert pile, they are a queue being worked off, and the reason the pile never empties is that a fresh cohort expires every day.

The queue is slow enough to be visible. Of the 2,448 markets listed as active on 2 September whose end date had passed before that day began, the median had been sitting there 14.3 days. One in ten had been there over 109 days, and 752 of them ended more than a month ago.

Half of the expired queue still has a live order book. 1,213 of those 2,448 markets return a two-sided quote, which means you can still trade a market whose event is over and whose outcome is known. 232 of them have taken more than $10,000, and 39 more than $100,000. Finding 07 says an unquotable market is usually a leg of a large field. This is the other half of that: a quotable market is not necessarily a live one, and nothing in the response distinguishes the two.

Eight consecutive day pairs, 25 August to 2 September 2026. Earlier days are excluded because they were recorded on a different pricing basis. The census is one snapshot per day, so any market listed and delisted inside the same day is invisible and the turnover figures are lower bounds. "Left" means Gamma stopped returning the market as active, which is what a client sees; it does not mean the record was deleted.

How I know this is Polymarket and not my own instrument. A daily count that swings by five thousand is exactly what an unreliable walk looks like, and I assumed that was the answer for a week. Three checks say otherwise. First, if the walk were dropping markets at random they would vanish and come back: across 74,599 markets and 330,224 market-days, a market is absent on a day between two days it was present 18 times. Second, a walk that truncates cannot know when a football match finished, yet departures track the end date at four to six times the base rate on all eight pairs. Third, I took 30 markets that disappeared between 30 and 31 August and asked Gamma directly: all 30 came back closed and resolved, none still active, none missing. The first version of this finding also put the drain at 72% against 1.9%, which was two different expiry cutoffs being compared with each other rather than a real effect; excluding the same-day boundary from both sides gives the 43% against 2.2% above.

What is not on this page matters as much as what is. One finding was cut while writing it, and one had to be measured twice. A calibration result relating price to eventual outcome no longer reproduces, because the overlap between markets I alerted on and markets whose settlement I captured is 17. And the first version of Finding 07 was measuring my own 8,000-token pricing cap rather than Polymarket's order books. It put the unquotable share near 80%, and I noticed only because the number came out as exactly 0.0% in one volume band. Pricing the full catalogue instead of the first 8,000 moved it to 10.4%, where it has since held for seven straight days.

That is the second time a cap in my own code has produced a confident, wrong, publishable number. The first one I published before catching. Measuring an exchange from the outside is mostly the work of telling its behaviour apart from your instrument's, and a number that looks clean is the one to check hardest.

Finding 08 is the first one here that went the other way. The daily count swinging by five thousand markets looked exactly like a broken walk, and I wrote it off as one for a week before testing it. It was real. The same habit that catches a bad number also throws away a good one, so the check has to run in both directions.

Findings 01 to 06 recomputed from the live database on 24 August 2026; the dataset summary above and Finding 07 on 31 August 2026; Finding 08 on 2 September 2026.
Gent Blaku · x.com/_GentB · linkedin.com/in/gentblaku
The bot is @p0lywhales. Not affiliated with or endorsed by Polymarket.