Analysis diagnostics
Before reading any effect off the Results page, the machinery has to earn trust. We build it up from the simple to the demanding: which search terms carry the signal, whether each host's digital twin actually tracks it out-of-sample, whether the comparison is contaminated by seasonal structure, and finally how every single event day stands against chance. We show each, openly β including where it's weak.
1 Β· The WISTs at a glance
Start with the raw vocabulary. Which wellbeing-indicative search terms carry the signal? Frequency = mean calibrated search level across treated markets; variability = coefficient of variation; zero-share flags flooring. High-frequency, low-flooring terms anchor the index; high-CV terms move the most.
| WIST | Category | Frequency β | Variability (CV) β | Zero-share β |
|---|
2 Β· How strong is each digital twin?
Next, the comparison engine. For every host market and term we fit donor weights on the first 80% of the pre-period and report error on the held-out 20% (validation MAPE) before refitting on the full pre-period (fit MAPE). A twin whose validation error balloons relative to fit is overfit; we flag twins above 10% validation MAPE as weak.
What MAPE means. Mean Absolute Percentage Error: for each pre-period week, take the gap between the host's real index and its twin's prediction, express it as a percent of the real value, drop the sign, and average across weeks. A MAPE of 2% means the twin is off by about 2% in a typical week β so lower is a tighter twin. We use it because it's scale-free: a big market and a small one can be held to the same bar, and the 10% line above is a single, comparable threshold.
3 Β· Seasonality of the host-vs-twin gap
The twins are fit to minimize error over the whole pre-period, so any seasonal mismatch between a host and its donor blend shows up as a recurring calendar pattern in the gap (observed β twin). The chart below averages that gap by calendar month over the entire pre-period. A flat profile means clean comparisons in every month; a bump in some months means windows that land there need a same-season baseline rather than an all-weeks one.
4 Β· Day-level placebo standing β every city, every event day
This is the test that carries the conclusion. What gets "scored"? For each event day we take that city's wellbeing-index gap β its observed index minus its twin's prediction that day β and ask where that gap ranks among the same host-minus-twin gaps on the city's own no-event days. A day at the 95th percentile means the gap that day beat the gap on 95% of the city's quiet days. We pool every such day into one picture. Under the null β no hosting effect β each day's percentile is uniformly spread on 0β100, so the dots would scatter evenly and ~5% would land past the 95th-percentile line. Mass piling toward the right is the effect. Each dot is one event day; cities are ordered by their median standing.
Pooled effect vs. 10,000 placebo re-randomizations
The paper-grade test. We pool every event day into one number β the mean gap in each city's own standard-deviation units β then ask how often a fake assignment does as well: for each city we slide its event block to a random stretch of its quiet days and recompute, 10,000 times. The histogram is that null; the line is what really happened. No distributional assumptions; each city judged against its own volatility; contiguous blocks preserve day-to-day autocorrelation.