Gambling methodology
How IndexFair rates gambling brands in the United Kingdom.
Every brand receives a score from 0 to 10, computed from public reviews, regulator filings, and operational signals. The score is the deterministic output of a versioned formula — given the same inputs and the same methodology version, the same number reproduces.
What our score means
v2.1 · United Kingdom · compositeThe IndexFair score is a single value between 0 and 10, attached to one brand, one vertical, and one country. It is the weighted sum of four signal families — Trust & Compliance, User Experience, Operational, and Measured Product — that we call the four-block composite.
Within each block, aspects are scored independently from their underlying signals, then combined according to a public weight table (§3). The composite is computed per the formula below; weights are pinned to the active methodology version, and each score is traceable to retained structured evidence and calculation snapshots for that version.
overall = 0.245 · A_trust + 0.28 · B_ux + 0.175 · C_ops + 0.30 · M_measured
When a block can’t be measured for a brand, the default is re-weighting: the block is left out and the remaining blocks share its weight at their original ratio. For ranked gambling cells the Measured Product block is instead filled with the market cohort median for display — disclosed on the brand page — while ranking uses the more conservative 40th percentile of measured peers, so a missing measurement alone cannot lift a brand’s rank. One honest limit: a brand whose product measures below that conservative baseline can still rank below an unmeasured peer; closing those measurement gaps is a standing collection priority.
Not every brand has enough qualifying signal coverage to publish every aspect projection. We label each brand's evidence basis explicitly:
- CompleteAt least 95% of methodology dimensions have ≥ 5 reviews. Highest-confidence display — every aspect carries a measured number.
- Broad80–94% of dimensions covered. Confident signal on most aspects; one or two show "insufficient signal" placeholders.
- Partial55–79% of dimensions covered. Composite score is still computed but reflects a narrower evidence base. Several dimensions show "insufficient signal" placeholders on the brand page.
- LimitedUnder 55% of dimensions covered. We do not publish a composite score for these brands; per-aspect details remain visible where signal exists, but the brand is not directly comparable to better-covered operators.
When an aspect shows — (“Insufficient signal”) instead of a number, it has not cleared the 30-signal model floor and any active evidence-source gate. Once those gates clear, the aspect switches from “—” to an overall-aligned projection.
Each block earns an absolute score from its own evidence — there is no grading on a curve. The combined composite reaches the published 0–10 scale in one of two modes, and every score artifact says which one it used. When a (country, vertical) has at least 12 ranked brands we apply a per-market ruler, so the same number means the same rank-position across markets: the median ranked brand sits near 5.5 and the strongest near 8.5, with genuine leaders above. The ruler is pinned to the ranked-brand distribution and never lets our own coverage gaps move a brand's level; a market where every ruler-calibrated brand is similar maps everyone to the midpoint — we do not manufacture spread.
Below 12 ranked brands — and in any market with no ranked brands yet — we drop the ruler and publish the brand's own honest composite on absolute bands clamped to the published range. That number is not calibrated against peers, and we label it as “small market — absolute bands” next to the numeral so a small-market score is never read as if it were peer-ranked. The ruler itself is also more stable: anchors come from interpolated percentiles, a small change in cohort statistics will not re-publish, and the per-market calibration anchors are now published as a dated, append-only time series.
A brand is ranked in a vertical only when it is a genuine operator of it — enough vertical-specific review evidence and a real product in that vertical.Otherwise it is listed as “Not yet ranked in {vertical}” rather than ranked on a reputation earned elsewhere (a casino-led brand does not top the sportsbook list). The caption is process framing, not a verdict. This is suppression, never a markdown: its score in the vertical it actually operates is unchanged.
Sources
v2.1 · United Kingdom · signalsEvery IndexFair score is built from genuine user reviews and observed facts about each brand. We continuously re-evaluate which sources carry real signal and which have decayed — so the mix below is the backbone of our coverage, not an exhaustive list. We deliberately don't publish every source we read: a score whose full input list is public is a score that can be flooded to order.
Main sources · genuine reviewsVerified-install reviews at scale — the closest thing to a representative cross-section of a brand’s actual users.
Real-time, unsolicited reports — including the unresolved complaints that never reach a formal review platform.
Real-time, unsolicited reports — including the unresolved complaints that never reach a formal review platform.
High-volume written feedback, read for the substance of each report rather than the headline star rating. We extract the real user reviews underneath these pages — never the marketing copy, and never the ranking the site is paid to show.
This list changes. Sources are added, down-weighted, or dropped as their signal quality shifts — a feed that starts carrying incentivised or templated reviews loses weight automatically (§4). Scores are computed per country; coverage for United Kingdom reflects the sources live there at the last recompute.
We independently estimate how much real-world traffic each brand draws and how its users arrive (direct, search, referral). This is a corroborating signal, not a score on its own: a brand with steady organic traffic and consistent review sentiment is treated as more reliable than one whose reviews spike with no matching audience behind them. Thin or anomalous patterns pull our confidence — and the interval around the score — down.
Beyond what users say, we measure the product directly. For casinos we count the live game library and assess game-provider quality; for sportsbooks we measure the market margin (the built-in overround a bettor pays) and how deep the markets run per event. These are first-party measurements — what the product actually is, not what the operator claims. The one disclosed exception: where an app-first sportsbook publishes no odds on the web, we instead compute its margin from a licensed third-party odds feed and label that provenance on the brand. These feed the composite as the Measured Product block, worth 30% of the overall when a measurement is present.
Aspects and weights
v2.1 · United Kingdom · casino weights| Subcomponent | Weight | What it measures | Math |
|---|---|---|---|
LICLicence status | 40% | Active licence verified against regulator register. | binary signal |
REGRegulator history | 20% | Five-year window of UKGC sanctions or warnings. | penalty decay |
AFFAffiliate integrity | 15% | AffiliateGuardDog + GPWA mapping. | category mean if missing |
YRSYears in operation | 10% | Years the operator has been in the market, from the brand founding year on record. | log scale |
LOYLoyalty composite | 15% | Traffic loyalty mix from Similarweb (direct, branded search, paid, affiliate). Contributes when traffic metrics exist; Block A renormalises over the other subcomponents when absent. | live at 0.15; omitted and renormalised when inputs are absent |
Technical detail
Block A includes Loyalty (LOY) in the active composite; the displayed weights are the active methodology schedule.
| Aspect | Weight | What it measures |
|---|---|---|
UXUX & onboarding | 32% | Overall experience of the site and app — registration friction, native experience, navigation, performance, and how the games feel to play. |
PAYPayouts | 31% | Time-to-cash, reliability, complaint rate per unit paid out. |
BONBonuses & fairness | 14% | Wagering reasonableness, T&C clarity, post-claim disputes. |
SUPCustomer support | 10% | Response latency, resolution rate, language coverage. |
KYCKYC & verification | 8% | Account safety, KYC proportionality, breach history. |
RGResponsible gambling | 5% | Self-exclusion, deposit and time limits, exclusion-scheme integration. |
| Subcomponent | Weight | What it measures | Math |
|---|---|---|---|
RESResponse engagement | 100% | Whether operators reply to negative reviews on covered review platforms and app stores: the share of a brand’s negative reviews over the last 90 days that drew an operator reply within 14 days. Silence on a covered platform is scored as a measured zero, not an exemption — non-engagement no longer dodges measurement. Because the complaint-response signal below has no ingested data yet, this is currently the only subcomponent producing a value, so it decides the whole of this block on its own — which is why the block itself now carries a proportionally smaller share of the overall score (see the note below the table). Auto-pasted template replies do not count, whether an operator repeats one template or rotates several. A partial-coverage signal: currently around four in ten rated gambling brands carry it (as of June 2026). Where a brand has too few negative reviews to measure, the signal is simply absent for it. | reply rate × 10 (measured zero on silence) |
CMPComplaint response rate | 0% | Whether the operator replied to a logged complaint within 14 days — a responsiveness signal, not a verdict on how the complaint was resolved. Defined and wired, but the complaint feeds have ingested no rows to date, so its effective weight is zero: it has no effect on any published score yet. When those feeds start delivering data it will take its share of this block back, and the block will carry a larger share of the overall score again. | reply-within-window rate |
Technical detail
One thing to know about this block as a whole: with the complaint-response signal not yet carrying data, this block is currently decided entirely by whether an operator replies to negative reviews — a signal directed at us and at the public record rather than at the customer whose complaint it was. Because it rests on one of its two scorable signals rather than both, the block carries a proportionally reduced share of the overall score: around 14% where product measurement is absent and around 10% where it is present, instead of the 25% and 17.5% a fully-populated block would carry. If the complaint feeds start delivering data, that share returns to full on its own.
Where a brand carries neither signal, this block is absent and its weight redistributes proportionally over the blocks that are present (renormalisation convention, §5). Cross-source consistency (XSV) is shown above as a confidence signal — computed and reported, never weighted into the score.
| Grade | Measured value | Score |
|---|---|---|
| Elite | raw ≥ 6.0 | 9.0 |
| Strong | 3.5 – 5.9 | 6.5 |
| Moderate | 1.5 – 3.4 | 4.5 |
| Thin | > 0 | 2.0 |
| Grade | Measured value | Score |
|---|---|---|
| Huge | ≥ 3,000 games | 9.0 |
| Large | 1,000 – 2,999 | 6.5 |
| Medium | 400 – 999 | 4.5 |
| Small | < 400 | 2.0 |
Technical detail
Graded on absolute published bands, not against the cohort. A measure we could not collect for a brand is left out and the rest re-weight — never scored as zero.
Provenance: most measurements are first-party — we read the operator's own product. The exception is app-first sportsbooks that publish no odds on the web; there we compute the margin ourselves from the operator's real prices read via a licensed third-party odds feed, and label the source on the brand (“odds sourced via [vendor]; margin computed by IndexFair”). A first-party reading always takes precedence where one exists (ADR 0211).
Coverage: live for a subset of brands (GB casino and sportsbook in rollout; US not yet measured, as of June 2026). For an unmeasured brand, this block is absent and its 30% redistributes proportionally over the remaining three blocks — the brand is neither rewarded nor penalised for the gap (renormalisation convention, §5).
Fake review filtering
v2.1 · United Kingdom · reviews · United KingdomA two-layer filter runs before any review enters the corpus. Heuristic checks catch the obvious patterns; an LLM judgment pass catches the remainder. Only the aggregate filter rate is published — reviewers are never named.
- 01Duplicate text fingerprint across operator corpus
- 02Account age < 14 days at review time
- 03Posting cadence > 8 reviews per day
- 04Sentiment ratio incongruent with star rating
- 05Reposted content from outside the source
- 01Coherence with stated factual claims
- 02Concrete event reference vs vague praise
- 03Author intent — personal vs commercial
- 04Cross-check against complaint corpus
- 05Per-language style anomaly detection
How we calculate
v2.1 · United Kingdom · formulaSeven deterministic stages explain the certified Block-B evidence, base block recomposition, served headline mapping and overall-aligned aspect projection. Persisted parent rows can also carry certified upstream adjustments, which the projection preserves. Internal review evidence and the public aspect numeral are deliberately different quantities.
- 01Weighted aspect evidenceFor each aspect a, eligible review signals are aggregated with the active source, confidence, content-quality and time-decay controls. Excluded or likely-fake rows do not enter the numerator or denominator.
raw_a = Σ (w_i · s_i) / Σ w_i over eligible signals i in aspect a
- 02Bayesian smoothingOn current signed gambling cells, a small effective sample is pulled toward the frozen empirical prior for the same leaf vertical. With the empirical-prior flag active, the mean uses k_mean = 10; the separate κ = 50 parameter controls uncertainty width, not the public level.¹ If that flag is inactive, the scorer falls back to prior = 5.0 and k_mean = 50.
E_a = (n_effective · raw_a + k_mean · prior) / (n_effective + k_mean); empirical flag ON: prior = prior_leaf, k_mean = 10; OFF: prior = 5.0, k_mean = 50
- 03Internal User Experience blockThe persisted per-aspect value E_a is model evidence on 0–10, not yet the public aspect numeral. Active aspect-definition weights combine the evidence into Block B; methodology v4+ then applies the fixed, monotone Block-B spread.
B = spread( Σ (v_a · E_a) / Σ v_a )
- 04Certified base recompositionThe certified composer combines Trust & Compliance, User Experience, Operational Signals and Measured Product over the blocks actually present. Missing measurements are null, never zero. C_base isolates the Block-B contribution; it is not assumed to equal persisted raw_parent, which can already contain upstream evidence, completeness or cap adjustments.
C_base = compose(A, B, C, M)
- 05Published cell scoreThe country × vertical cell’s frozen terminal mapping converts its persisted raw_parent to the served headline score. In a small cohort, absolute mode uses raw_parent directly within the published band; ruler mode uses the saved market parameters.
P_parent = R_cell(raw_parent)
- 06Overall-aligned aspect projectionFor public aspect a, its internal evidence replaces Block B while A, C, M and the same frozen cell ruler stay fixed. The result answers a counterfactual about the whole published model; it is not the average mood of people who chose to post.
q_a = methodology_version ≥ 4.0.0 ? spread(E_a) : E_a; raw_a = clamp(raw_parent + compose(A,q_a,C,M) − C_base, 0, 10); P_a = R_cell(raw_a)
- 07Integrity and evidence gateA numeral appears only when the facet and parent share methodology, processing-build and flag-set stamps, the facet is not stale, at least 30 signals support it, and replaying the parent ruler reproduces the served score. Failure shows an insufficient-signal state; raw sentiment is never a fallback.
show P_a iff stamps match ∧ computed_facet ≥ computed_parent ∧ n_signals ≥ 30 ∧ replay(parent) ≈ served
Time decay
v2.1 · United Kingdom · 0.5^(d / h_source)Every eligible review carries an exponential time-decay factor using its source’s configured half-life. The scorer has no hard age cutoff: older evidence remains in the corpus with continuously diminishing weight. The chart uses an illustrative 540-day half-life and stops at 48 months only as a display horizon.
What we do NOT do
Defensive commitments. If a behaviour you'd expect from a rating site is not listed here, assume we do not do it.
- ✗We do not accept brand payment to alter, reorder, or suppress scores.
- ✗We do not surface review text verbatim — aggregate metrics only.
- ✗We do not promote brands by ranking position. Order is computed; placement is not for sale.
- ✗We do not use marketing language. Words like "the best", "leading", "trusted" do not appear.
- ✗We do not recommend a brand by name in editorial copy.
Current version
Full methodology & changelog
- v2.12026-06-24minorcurrent
VPN profiles go live as audit-anchored. A VPN’s no-logs standing is reported only when an independent third party has assessed it — a published independent no-logs audit, or a court or seizure finding — and is described as assessed by independent audit, naming the auditor and date, never as proven or guaranteed. No overall VPN safety rating is published; service quality is shown separately from this factual no-logs standing.
- v2.02026-06-09major
The betting score now differentiates on the absolute market margin measured from each operator’s own live prices. Brands with a genuinely sharper product reach the top of the scale on measured merit; brands we cannot yet measure are placed at a neutral position rather than penalised. Vertical eligibility keeps casino-led brands out of the betting ranking.
- v1.02026-06-02major
First public methodology. The headline score combines regulatory and compliance standing, aggregated genuine user-review signal, and operational signals — computed per country and per vertical, and displayed on a fixed band that keeps licensed and offshore brands on separate, non-overlapping scales.