Quantitative note · VCT 2023–2026 · 2026-08-31
Naive. Across 1,732 matches and 91,393 pre-match comments the thread lands on a coin flip, and adds nothing once you know the price.
vlr.gg match pages carry a comment thread, and much of it is people calling the score before the game. This asks whether those calls contain anything the betting market has not already priced.
They do not. Across 1,732 VCT matches from 2023 to 2026, a signed measure of which way each thread leans has no relationship to the result once the bookmaker's implied probability is in the model (β₂ = −0.272, p = 0.085). On its own the crowd is right 50.9% of the time, which is a coin flip with rounding. The market, by contrast, is genuinely informative: AUC 0.697, and the favourite wins 63.7% of the time.
The measurement itself went through six rounds of correction before it settled, and how a comment becomes a number is set out in full in section 02. Section 07 is candid about what that process cannot do: 40% of comments stay unresolved, the weights are chosen rather than fitted, and a noisy regressor biases toward exactly the null reported here. Against that, the coefficient has been re-derived eleven times across four seasons and six rounds of extraction fixes, and has never left one standard error of zero. It is also null in each of the four seasons taken on its own.
logit(a_won) = β₀ + β₁·implied_prob_a + β₂·crowd_lean
The market works. AUC 0.697, Brier 0.2202 against 0.2500 for a coin flip. Favourites win 63.7% of VCT matches.
The crowd does not. Directional accuracy 50.9%, AUC 0.510. As a solo predictor its coefficient is +0.097 with p = 0.51, which is to say nothing at all.
Together, the crowd adds nothing. β₂ = −0.272, p = 0.085. A likelihood-ratio test against the market-only model agrees. Out of sample, over five expanding-window folds, adding the crowd never helps:
Walk-forward AUC · market + crowd, versus market alone
The grey tip marks the one fold where the market-only baseline scored higher. It wins fold 1 by 0.001; the combined model is ahead by thousandths in the other four, for a mean difference of +0.0022. Nothing separates them, a 0.002 AUC gap sits far inside fold-to-fold variation, since fold 5 alone swings 0.05. This is what a coefficient of zero looks like out of sample.
A single coefficient over four seasons could be an average of eras that disagree. It is not, but the seasons are far from alike, and the way they differ is the best argument in the study for asking this question rather than a simpler one.
| Season | Matches | Market AUC | Favourite wins | Crowd accuracy | β₂ | p |
|---|---|---|---|---|---|---|
| 2023 | 254 | 0.782 | 71.3% | 59.1% | −0.160 | 0.752 |
| 2024 | 427 | 0.662 | 60.9% | 50.8% | −0.271 | 0.358 |
| 2025 | 496 | 0.697 | 64.1% | 50.6% | −0.140 | 0.665 |
| 2026 | 555 | 0.678 | 62.0% | 47.4% | −0.451 | 0.090 |
2023 is the instructive one. The first franchised season had the widest talent gaps between organisations, and it shows everywhere: the market reaches an AUC of 0.782 and favourites win 71.3% of matches, both far above any later season. The crowd looks good there too, at 59.1% directional accuracy, its best showing anywhere in the data, and comfortably better than a coin.
And it still adds nothing: β₂ = −0.160, p = 0.752. The crowd was accurate precisely when matches were easy to call, which is exactly when the market was accurate as well. Raw crowd accuracy swings twelve points across the four seasons, from 59.1% down to 47.4% as the regions closed up, while the incremental contribution stays null in every one of them.
Read as an accuracy question, the crowd looks skilled in 2023 and useless in 2026. Read as the question actually asked, does it know anything the price does not, the answer is the same in both, and the apparent skill turns out to be a property of the fixtures rather than of the crowd.
A null is only worth reading if the study had the power to detect an effect. With a standard error of 0.158, anything at |β₂| ≥ 0.443 would have reached significance. That is a 21% shift in win odds per standard deviation of crowd lean. The observed value is 0.272. So: no usable edge, rather than proof of exactly zero.
β₂ was re-estimated as the sample grew and as the extraction improved:
−0.212 −0.059 −0.271 −0.098 −0.163 −0.178 −0.145 −0.170 −0.166 −0.319 −0.272
Every value sits inside one standard error (≈0.158) of the others and of zero. Three of those rounds were substantial. One removed roughly a thousand comments that were predictions about other matches, and β₂ moved by 0.004. One corrected six extraction rules (negation binding to the wrong verb, transitive verbs read backwards, rhetorical stakes counted as forecasts) and moved it from −0.166 to −0.319.
The last is the strongest test available: the corpus went from two seasons to four, adding 675 matches and 43,644 comments from a period the rules were never tuned on. β₂ moved to −0.272 and the standard error fell from 0.204 to 0.158, tightening the estimate by a quarter. More data made the answer more precise and left it in the same place. If a real signal were hiding under measurement noise, errors and samples of that size should have surfaced it.
Every pre-match comment is reduced to a signed number on one axis:
+1 means fully behind team A, −1 fully behind team B,
0 no view. Two things determine what a comment contributes: which
rule read it, and how strongly it was written.
Rules run in order of how much they can be trusted. The first one that resolves wins.
| Basis | Example | Notes |
|---|---|---|
| score/series | LOUD 2-0 | Strongest signal on the site, and lexically neutral |
| score/blowout | G2 26-0 NRG | Bravado for a sweep; scaled by how lopsided |
| score/percent | DRX 51% - 49% TL | A stated coin flip, so it barely counts |
| counterfactual | sen choked, shldve been 2-0 | Defends the named team; the literal score says the opposite |
| preference | ...but not over TL | Ranks one team above another |
| outcome | TL easy win | An explicit claim, not a passing reference |
| mvp / player | CINEMA mvp tkzin | Resolved through a dated roster |
| trash | Loud fans crying in the corner | Abuse aimed at one side backs the other |
| cheer | COME ON SENTINELS | Allegiance rather than analysis |
| mention | VARREL LFGGGG | Weakest. Naming a team is not backing it |
Direction says who. Conviction says how loudly, and it is the weight in the aggregate:
crowd_lean = Σ(direction × conviction) / Σ(conviction)
A comment at 0.85 moves the match figure about four times as far as one at 0.21. It starts at 0.55 and moves with the language:
| Signal | Effect |
|---|---|
Certainty: ez, free money, lock, no diff, fácil | +0.30 |
Hedging: ig, probably, might, 50/50, acho | −0.28 |
| Exclamation mark | +0.05 |
That figure is then scaled by how far the rule can be trusted. A sweep
scoreline keeps all of it; a bare mention keeps 60%; cheering keeps 45%. The
percentage rule is the interesting one, because it scales on the stated margin:
51% - 49% collapses to roughly 0.03, since the
author is telling you they cannot separate the teams. Reading that as a confident
call would invert what they meant.
Conviction is not a probability, and it is not a claim the comment is right. It measures how firmly a view was put and how much the rule that read it deserves believing.
| Flag | Meaning | Effect |
|---|---|---|
| S | Explicit /s, author-marked sarcasm | direction inverted |
| F | Fan-conflict wording: cope, copium, delusional | conviction × 0.6 |
| $ | Mentions odds, units, a payout, a price | feeds odds_talk_lean |
| r | A reply, not a top-level post | none (yet) |
F exists because of a mistake. Those words were originally treated as sarcasm and used to invert a comment. Measuring showed the detector fired on nine comments out of 91,393, and reading all nine showed the words are almost always aimed at the opposing fanbase while the author asserts their own pick:
Delusional SEN fans lmao, keep coping. DRX wins easily 2-0
Inverting that turns a correct read into a wrong one. So the flag now only signals heat, and pulls the weight down. Unmarked sarcasm remains undetected, and the note does not pretend otherwise.
$ changes nothing on its own. It supports a separate figure,
the lean among people talking about price. On the 2025 grand final that subgroup
read +1.000 in a thread that was otherwise dead even.
r is currently descriptive. Replies argue with their parent, so a thread is not a set of independent votes, and this flag is how you would notice that biting.
Polarity is the positive-to-negative axis of a piece of text. It is not a side
concern here, it is the quantity the whole study rests on: crowd_lean
is nothing but an average of per-comment polarities, and if that measure is wrong
the regression that follows is meaningless.
So polarity is necessary. What is not sufficient is polarity as a sentiment model computes it, because that version has no target. It answers "is this text angry or pleased?" when the question is "which of these two teams does this text favour?" Every comment in a match thread has two candidate subjects. Mood alone cannot say which one is meant, and abuse aimed at either side scores identically negative.
What the project needs is polarity conditioned on an entity,
aspect-based sentiment. Formally, (target, polarity) pairs where the
target is a team or a player, projected onto one signed axis. The
direction field is exactly that: the polarity of a comment toward
team A. It is not a mood score, and building it is closer to entity resolution
than to sentiment analysis.
Running VADER over 3,466 comments that the extractor resolved:
| Measure | Value | Reading |
|---|---|---|
| corr(polarity, direction) | +0.018 | no relationship |
| sign agreement, non-neutral | 51.2% | a coin flip |
| scored neutral overall | 46% | invisible |
| scorelines scored neutral | 57% | the best signal, unseen |
| corr(polarity, outcome) | +0.002 | predicts nothing |
| corr(direction, outcome) | +0.052 | weak, but real |
The three comments below show why. Scores are VADER's compound output.
| Comment | Polarity | Actual signal |
|---|---|---|
Furia 2-0 Loud fans crying in the corner | −0.477 | max pro-Furia |
LOUD 2-0 | 0.000 | max pro-LOUD |
Bro free money here, how tf is LOUD 2x odds | +0.511 | pro-LOUD |
The first is read backwards: the hostility belongs to Furia's supporters and the sentence is a confident Furia pick. The second is the strongest signal on the site and carries no sentiment vocabulary at all. The third gets the right sign by accident, because "free money" is cheerful language; rewrite it as "terrible line, LOUD at 2x is robbery" and the score flips while the meaning holds.
Measurement error pulls coefficients toward zero. A polarity-based
crowd_lean, 46% zeros, 51% coin-flip signs, would have returned
β₂ ≈ 0 and a comfortable p-value whether or not the crowd knew anything. The
headline would have looked identical and meant nothing, because there would be
no way to separate the crowd is uninformative from the measure is
uninformative.
The null is worth reading only because the measure demonstrably captures something: it resolves 60.2% of comments, every resolution traces to a named rule you can inspect line by line, and the one convention that mattered most was settled by hand-reading 26 comments.
Deterministic rules handle 60.2%. Most of what remains is banter, scheduling and reply fragments that should stay unresolved.
Every franchised VCT season, 2023 through 2026. The shape changed over that span: 2023 ran three regions with one continuous league each and opened with LOCK//IN São Paulo; China was franchised in 2024, which is also when the Kickoff / Stage 1 / Stage 2 structure appears. Event IDs came from vlr's own season index pages rather than guesswork, and all-star fixtures filed under VCT events are excluded.
| Season | Tier | Matches | Comments | Median thread | Coverage |
|---|---|---|---|---|---|
| 2023 | League | 165 | 8,737 | 45 | 59.2% |
| 2023 | Masters | 55 | 7,746 | 120 | 55.4% |
| 2023 | Champions | 34 | 3,842 | 94.5 | 54.7% |
| 2024 | League | 353 | 14,712 | 35 | 58.5% |
| 2024 | Masters | 40 | 4,486 | 107 | 53.7% |
| 2024 | Champions | 34 | 4,306 | 97.5 | 54.4% |
| 2025 | League | 422 | 17,458 | 35 | 61.3% |
| 2025 | Masters | 40 | 4,527 | 99 | 55.1% |
| 2025 | Champions | 34 | 3,824 | 98.5 | 57.0% |
| 2026 | League | 507 | 16,489 | 29 | 64.0% |
| 2026 | Masters | 48 | 5,266 | 96 | 51.3% |
International events draw roughly three times the thread depth of a league match, which ought to be the best possible conditions for a crowd signal. Nothing appears there either.
They also have the lowest coverage, and that is not a defect. Big-event threads are full of bracket talk, and a prediction about a different fixture is correctly ignored.
The comment cutoff is strict: posted before the match started, compared against vlr's own pre-match flag on 60 sampled matches, agreeing on all 60.
vlr archives only the winning team's pre-match price on a finished match. The loser's is gone. I first concluded that made historical odds unusable, since assigning a reconstructed price to a team seemed to need the outcome.
With a book of total size K, one price determines the other. For a true p(A) = 0.65 at K = 1.068:
A wins → published A @ 1.44 → p(A) = 0.65 A loses → published B @ 2.23 → p(A) = 1 − 0.419 = 0.65
Identical either way. The outcome picks which side you subtract from, not the answer, so the recovered value is exactly what a pre-match observer would have computed. It is asserted at five probability levels in the test suite.
That unlocked 1,051 of 1,732 matches.
Two things were checked rather than assumed. A favourite-longshot skew would make the reconstruction error correlate with the outcome; measured slope is −0.0038. And K was fitted from observed quotes rather than taken from a textbook.
The three books vlr lists price an even match at 1.87 / 1.87. Two implied probabilities of 53.5% sum to 106.95%, and that surplus is the bookmaker's margin. Every strategy in this study has to clear 6.95% before it breaks even, which is why a small edge is not enough: the crowd would need to be right 53.5% of the time simply to stand still. It manages 50.9%.
These are crypto books, and the level matters. A traditional bookmaker quotes nearer 1.83 / 1.83, a 9.09% overround, raising the same hurdle by two points. Neither figure moves the coefficient: β₂ barely changes across K ∈ [1.05, 1.12].
The genuine trap, still avoided, is reading the displayed price as "team X's odds". Which team is displayed is the label.
The pipeline runs on a single link. It fetches the page, applies the cutoff, resolves each comment, and puts the crowd's position next to the market's. Upcoming matches work by definition, since every comment on them is pre-match.
$ python src/predict.py https://www.vlr.gg/734304/nrg-vs-loud-... NRG vs LOUD VCT 2026: Americas Stage 2, Upper Semifinals status: upcoming pre-match comments : 30 resolved to a lean : 25 (83% coverage) CROWD LEAN +0.038 toward NRG MARKET P(NRG) = 75.0% [two_sided:ggbet] crowd, naively as a probability: 51.9% disagreement: crowd is 23.1% higher on LOUD
The disagreement line is description, not a tip. Given everything above, a gap between crowd and market is a fact about the thread.
Three flags exist for checking the extractor rather than using it:
--all prints every comment in thread order, --unresolved
prints only the misses, and --csv dumps 18 columns per comment.
--unresolved is the one that finds bugs. Three sign inversions were
caught that way, none of them visible in a thousand matches of aggregate
statistics.
Three separate things could be wrong here: how a comment becomes a number, how those numbers are combined, and what the combined figure is then asked to prove. They fail in different ways and deserve separating.
The extractor is a cascade of hand-written rules, not a trained model. That buys auditability: every number traces to a named rule and you can read the line that produced it. It costs generality. Each rule was added after watching it fail on real threads, which means the ruleset is shaped by the threads I happened to read. A league or a language I never looked at may break patterns I never wrote.
Coverage is 60.2%. The remaining 40% is mostly banter, scheduling, reply fragments and ASCII art, and should stay unresolved. But not all of it. Non-English predictions are the largest recoverable group, and Turkish rather than Chinese is the notable one, which was the opposite of what I first assumed.
The weights are chosen, not fitted. A sweep scoreline counts full, a bare mention 60%, cheering 45%. Those figures are judgements about how far each rule can be trusted, not estimates from data. The same is true of the 48-hour recency half-life. Fitting them against outcomes is possible in principle and was not done, partly because a signal this weak gives very little to fit against.
Precision was preferred to recall throughout. Roughly a thousand cross-fixture predictions, 570 statements of preference rather than forecast, and every map-veto recap are now discarded rather than counted. Each removal lowers coverage. In a signed mean a comment pointing the wrong way costs more than one that is simply absent, so that trade is deliberate.
Comments are treated as independent votes. They are not. Replies argue with their parent, and a thread of forty posts is not forty opinions. Depth is recorded on every row and currently does nothing with it.
Every commenter counts the same. A first-time poster and someone with a long accurate record carry identical weight. Per-user skill, with shrinkage for low volume, is the most interesting thing left undone.
Fanbase size is not normalised. A team with a larger following generates more comments regardless of how good they are, and the raw share of voices inherits that.
Unmarked sarcasm is invisible. An explicit /s is
handled. Nothing catches "great, another flawless performance incoming", and no
lexicon will.
Validating an interpretation needs the right ground truth. One scoreline convention was settled the wrong way for a long time because I tested it against match results. That test could never have worked: outcomes only separate two readings if the text predicts outcomes, and this study had already shown it does not, so both readings scored about 50%. Reading 26 comments by hand settled it immediately. Outcome labels answer what words predict; only a hand-labelled sample answers what they mean.
Measurement noise biases toward the result found here. A null is the easy thing to produce with a noisy regressor, and at 60.2% coverage the crowd measure is noisy. Against that: the estimate has been re-derived eleven times across four seasons of data and six rounds of extraction fixes, and has never moved outside one standard error of zero. One round removed a systematic contamination of about a thousand comments and shifted β₂ by 0.004; the most recent corrected six rules and moved it 0.153, still inside that band.
The study was powered for a moderate effect, not a small one. Anything at |β₂| ≥ 0.443 would have reached significance. Below that it would not have been seen. This is evidence of no usable edge, not proof of exactly zero.
The sign is more persistent than the null language suggests. Rebuilding the signal 25 ways (four weighting schemes, six bases, two recency windows, three thread depths, five leagues, four seasons, three tiers) leaves 24 of 25 cells negative, and one of them crosses p < 0.05: the unweighted signal at β₂ = −0.350, p = 0.042. One significant cell in 25 is exactly the chance rate at that threshold, and it is not independent of the headline: same comments, same matches, differing only in whether conviction weights them. The conviction weighting was fixed long before this run, so adopting the variant that clears the bar after seeing which one does would be picking the specification from its p-value.
Still, a direction that survives that much slicing is not nothing. The honest statement is not "there is no effect" but "there is something faintly negative that four seasons and 1,732 matches cannot separate from zero once the price is in the model", which is an argument for another season of data rather than for a claim.
Nothing here is a betting claim. There is no cost model: overround, stake sizing and line availability are all unmodelled, and a significant coefficient still would not have been a profitable one.
The price carries no timestamp. vlr stores one figure per book with no time attached, so whether it is an opening or a closing line is unknown. A strict claim about market efficiency needs the closing line, and this is not quite that.
The archived prices are slightly generous, and every P&L here
inherits that. Backing both sides of every match should lose the
overround, about 6.50%. On the prices vlr archives it loses 2.98%:
the mean archived winner price is 1.940 where a 1.87 / 1.87 book would pay 1.870.
Part of that is the median-of-three-books stored in the cache, part is the missing
timestamp above. Charging the full margin via
backtest.py --enforce-overround moves following from −7.38% to
−10.74% and turns the fade from +1.32% into −2.35%. The direction of
the headline is unaffected, following the crowd loses either way,
but any strategy whose margin is a few points should be read at the stricter
number.
Four seasons, top tier only. Franchised VCT 2023–2026. The game, the rosters, the leagues and the betting markets all change; 2023 looks materially unlike 2026. Nothing here is claimed beyond them.
A coefficient of −0.272 with p = 0.085 is hard to feel. Here is the same finding as a betting account: $100,000 opened on 13 February 2023, the first match of the first franchised season, and run to 27 August 2026, 1,731 settled bets across every VCT match in the corpus, backing whichever team the thread favoured, settled at the prices vlr archived.
The cleanest version. Equal money on every match, so nothing but the crowd's direction is being tested.
Flat $100 a bet · bankroll over 1,731 settlements
Following loses 7.38% of everything staked. The hit rate is 50.9%, and at a 6.95% overround you need 53.5% to stand still. That gap is the entire result, expressed in money.
The dashed line is the mirror bet, backing whichever team the thread does not. It is the control, not a proposal, and it is dealt with below.
The same bets, staked by how one-sided the thread was. If the crowd's confidence carried information, this should beat flat staking.
Binned stakes · $100 at 50-55%, $150 55-60%, $200 60-65%, $250 65-70%, $300 70-75%, $350 75-80%, $400 80-85%, $500 85-100%
It does not. ROI moves from −7.38% to −6.76%, six tenths of a point and well inside noise, while the cash loss nearly doubles to −$24,238, because sizing by confidence puts twice as much money through the book. More exposure, the same rate of loss.
The bins carry no gradient either. The 80-85% band returned −13.5% across 397 bets, while the 70-75% band, less confident, not more, is the only positive one at +2.7%. The most confident band of all, 85-100%, returns −4.7%. Confidence and return are unrelated across the whole range, which is what you see when the thing being sized against is noise.
Over the full sample it looks like it: the binned fade finishes +$11,827, profitable in 82% of bootstrap resamples. Four seasons make that checkable, and it fails three ways.
| Test | Result |
|---|---|
| Spread across seasons? | No. 2024 alone gives +$12,143 of the +$11,827. The other three seasons sum to −$317. |
| Holds out of sample? | No. Fitted on 2023-24 it returns +7.94%, profitable in 88% of resamples. Traded forward over the next 1,050 bets: +0.16%. |
| Survives the real margin? | No. The archived prices are ~3.6% generous (§07). Charged the full 6.95% it goes to −0.45%. |
What survives is the other side. Following the crowd reliably loses, 0% of 10,000 resamples profitable, an interval of [−$21,898, −$3,685] that never touches zero, in every season, region and confidence band. But "following loses" is not "fading wins", because both sides pay the margin. The crowd's anti-predictiveness is worth about four points of ROI against a neutral bet; the hurdle is 6.95%.
Note also how the fade earns: it wins less often than following, 49.1% against 50.9%, and profits only because its winners arrive at longer prices. That is a more fragile thing to own than a hit-rate advantage, and it depends on getting those prices, which an untimed archived line cannot promise.
| Region | Bets | Hit | ROI |
|---|---|---|---|
| EMEA | 379 | 53.3% | −2.63% |
| International | 285 | 54.0% | −4.34% |
| Americas | 372 | 56.2% | −5.10% |
| China | 320 | 48.1% | −9.11% |
| Pacific | 375 | 43.2% | −15.29% |
Over four seasons no region escapes. The spread runs from −2.63% in EMEA to −15.29% in Pacific, 12.7 points across 285-379 bets each, and every cell is negative. On two seasons alone International looked profitable at +3.68%; with four it sits at −4.34%, which is a useful reminder of how little a 122-bet sample was saying.
The losses also deepen by season, −3.55% in 2023 to −10.42% in 2026, tracking the fall in crowd accuracy shown in section 01. No season clears the hurdle.
Per team, two quantities: how far its crowd's belief sat above the team's real win rate, and what backing that team actually returned.
49 teams · crowd bias against return · bubble = bets placed
The relationship is real and moderate. DetonatioN FocusMe is the clearest case: across 59 matches its thread backed them at 53.8% while the market priced them at 32.8% and they actually won 30.5%. The market had them almost exactly right; the crowd was 23 points out, and following it cost 35.8% of everything staked on them. The biggest dollar loss is Global Esports at −$1,369, less overrated than DFM but appearing more often. Bias sets the rate; volume sets the bill.
The left-hand side is the mirror image. FNATIC were underrated by 18.6 points, the crowd gave them 51.9% across 112 matches while they won 70.5%, and backing them returned +22.4%. The crowd was wrong about FNATIC for four straight seasons in the same direction. That is the single most persistent bias in the data, and it is still not tradeable: the bookmaker had them at 65.0%, so almost all of the gap was already in the price.
Knowing a fanbase is deluded does not help, because the bookmaker knows it too and has already priced it; that is what both examples above show. The tradeable quantity was never "is the crowd wrong" but "is the crowd wrong in a way the price has missed", and that is β₂ = −0.272, p = 0.085, over four seasons and 1,732 matches.
The best-returning side, Wuxi Titan Esports Club at +36.6%, came off 12 bets at an average price of 2.88 with a 41.7% hit rate. Twelve bets at near-3.00 is variance wearing a nice hat.