WellbeingEffect

Analysis diagnostics

Before reading any effect off the Results page, the machinery has to earn trust. We build it up from the simple to the demanding: which search terms carry the signal, whether each host's digital twin actually tracks it out-of-sample, whether the comparison is contaminated by seasonal structure, and finally how every single event day stands against chance. We show each, openly β€” including where it's weak.

1 Β· The WISTs at a glance

Start with the raw vocabulary. Which wellbeing-indicative search terms carry the signal? Frequency = mean calibrated search level across treated markets; variability = coefficient of variation; zero-share flags flooring. High-frequency, low-flooring terms anchor the index; high-CV terms move the most.

WISTCategory Frequency β‡… Variability (CV) β‡… Zero-share β‡…

2 Β· How strong is each digital twin?

Next, the comparison engine. For every host market and term we fit donor weights on the first 80% of the pre-period and report error on the held-out 20% (validation MAPE) before refitting on the full pre-period (fit MAPE). A twin whose validation error balloons relative to fit is overfit; we flag twins above 10% validation MAPE as weak.

What MAPE means. Mean Absolute Percentage Error: for each pre-period week, take the gap between the host's real index and its twin's prediction, express it as a percent of the real value, drop the sign, and average across weeks. A MAPE of 2% means the twin is off by about 2% in a typical week β€” so lower is a tighter twin. We use it because it's scale-free: a big market and a small one can be held to the same bar, and the 10% line above is a single, comparable threshold.

Lower = better. Paired bars per host: in-sample fit vs out-of-sample validation. Dashed line = 10% (weak-twin threshold).

3 Β· Seasonality of the host-vs-twin gap

The twins are fit to minimize error over the whole pre-period, so any seasonal mismatch between a host and its donor blend shows up as a recurring calendar pattern in the gap (observed βˆ’ twin). The chart below averages that gap by calendar month over the entire pre-period. A flat profile means clean comparisons in every month; a bump in some months means windows that land there need a same-season baseline rather than an all-weeks one.

What we found (and what we did about it). In-time placebo tests on the World Cup window of 2023, 2024 and 2025 all showed spurious "effects" against an all-weeks baseline β€” the gap runs systematically higher in summer. Against a matched-season baseline (same Jun 11 – Jul 19 window of other years) all three placebos are null. The Results page therefore defaults to the matched-season baseline; the all-weeks baseline remains available as a sensitivity view.
Get updates as results come in
We'll email when new results land.