top of page

Cast members can lose you the day. They cannot win it.

2 days ago
6 min read

I labeled 21,656 paragraphs of Reddit discussion about seven Orlando theme parks, rated each one on five things visitors care about, and tested which of them actually move a person's verdict on the whole visit.


Four of the five move it both ways. Great rides lift the verdict, bad rides sink it. Same for queues, same for food and theming, same for value.


Staff service does not. Bad service is one of the strongest negative signals in the data, at -1.59 in log-odds. Good service comes in at +0.42 and cannot be distinguished from zero.


Service is a floor, not a ceiling. You can fall through it. You cannot stand on it.

Each attribute enters the model twice, against a baseline of not being mentioned. The pale band behind each bar is the 95% interval. Staff service is the only row where the two bars are not roughly mirror images.
Each attribute enters the model twice, against a baseline of not being mentioned. The pale band behind each bar is the 95% interval. Staff service is the only row where the two bars are not roughly mirror images.

That's the finding. Most of the work was making sure it wasn't an artifact.


Paragraphs, not posts

My first version attributed each post to a park, then required the post to name exactly one park before using it.


That threw away 2,879 posts. Which 2,879? The trip reports. The posts where someone spends 1,200 words on four parks across five days, which is exactly where the densest evaluative writing lives. I had built a filter that discarded my best data.


The fix was to move down a level. Attribute sentiment lives in paragraphs: a paragraph about a 90-minute queue tells you about queues, not about Magic Kingdom overall. Verdicts live in whole posts, usually stated once near the end.


So labeling runs per paragraph and analysis runs per post-park pair, aggregating that pair's paragraphs. Usable observations went from 3,333 to 21,656, and measured accuracy went up at the same time, which is not the usual trade.


The model extracts, the code decides

I took this rule from my last project and it earned its place again.


The labeling model is never asked to compute a score, rank an attribute, or judge what counts as important. It answers one question per paragraph per attribute: does this say something good, neutral, bad, or nothing at all. Everything downstream is Python, where it can be tested.


That matters most for the outcome variable. If I had asked a model "was this a good visit" while also asking it to rate five attributes, the answers would contaminate each other and the whole test would be measuring itself. Instead, overall_verdict is filled only when an author judges the park as a whole, and would_return is captured separately as a behavioral signal. Neither is derived from the attribute ratings.


I checked the analysis code before trusting it, by running it against 15,000 simulated labels with a known structure planted in them. It recovered all five correctly.


Making the attribution trustworthy

Paragraphs get assigned to parks by pattern matching, which is fallible in ways you cannot see from the summary statistics. So I hand-judged 140 segments against the source text and measured it.

version

precision

v1, post level

0.60

v2, paragraph level

0.69

v2.1, ride rosters added

0.74

v2.2, final

0.82

Two sources got cut after measurement rather than before.


SeaWorld Orlando scored 0.12. Five correct out of forty. Its coaster names are used as universal benchmarks by enthusiasts, so "not as good as Mako" is a sentence about a park in Germany. No pattern fixes that.


r/rollercoasters scored 0.38 for the same reason one level up. Enthusiast trip reports name an Orlando ride in passing while reviewing a park elsewhere. "Hyperspace Mountain and Avengers Assemble" pulled Magic Kingdom into a Disneyland Paris report.


Both were things I wanted in the dataset. Measuring them is what got them out.


The failure that nearly shipped

To cut cost I packed eight paragraphs into each API request instead of two. Overall label agreement against the two-paragraph run came back at 93%, which passed every check I had.


Then I looked at the disagreements by direction instead of in aggregate, and found that packing eight had lost 41% of queue detections and 29% of price detections. Not scattered errors. One-sided losses, concentrated in attributes that show up as a passing mention rather than as a paragraph's headline. Cram eight paragraphs into one prompt and the model reports what each one is mainly about, quietly dropping the secondary mentions.


93% agreement and a systematically broken dataset, at the same time. The aggregate metric was not just insufficient, it was actively reassuring.


The final run used two paragraphs per request.


I only caught it because I had written a comparison that reported direction. I had written that tool, then not looked at its directional output for two runs.


Why the obvious analysis misses this

Fit one coefficient per attribute, the way you normally would, and every one comes back significant, spread across a range of 1.07 to 1.44, with confidence intervals lying across each other. Four of the five clear p < 0.0001. The model says all five matter about equally and there is nothing further to find, and the apparent ranking between them is noise.


The problem is not precision, it is shape. One coefficient per attribute assumes the effect is symmetric, that being good helps as much as being bad hurts. Collapsing both directions into one number turns "bad service is devastating" into "service matters a lot", which points at the opposite management decision.


So each attribute enters twice instead, as rated-good and rated-bad, both against a baseline of not being mentioned. Same 816 pairs, same scale, different shape:

Both panels share an axis, which is the whole comparison. The left model's five estimates fit inside a quarter of the range the right one needs.
Both panels share an axis, which is the whole comparison. The left model's five estimates fit inside a quarter of the range the right one needs.

attribute

good

bad

ratio

attractions

+1.18

-1.76

0.67

queues and crowding

+1.40

-1.15

1.22

food and environment

+1.26

-1.25

1.01

value and price

+1.23

-1.06

1.16

staff service

+0.42

-1.59

0.26

Bold is p < 0.05. Standard errors are clustered on post, because a trip report covering four parks contributes four correlated rows and treating them as independent would overstate everything.


Four attributes sit near 1.0 on the ratio. Staff service sits at 0.26, and tightens to 0.13 when you drop planning posts and speculation to keep only firsthand accounts.


What I'd tell an operator

Cast members are already carrying enormous weight. That is precisely why more of them, or better ones, is not where the marginal dollar goes.


The asymmetry says service is a hygiene factor in the classic sense. Getting it wrong is expensive and getting it right buys nothing measurable. The operational implication is not "invest in service", it is "eliminate the bottom tail". The bad interaction is what shows up in the data, and the goal is fewer of them rather than more good ones.


Attractions, queues and food are where upside lives. They move the verdict in both directions, so investment there is visible in a way service investment is not.


The uncomfortable version: if you are choosing between a service excellence initiative and fixing the ride that breaks down twice a day, this data says fix the ride.


Limitations

Results are conditional on someone stating a verdict. 816 of 8,981 post-park pairs made the analysis sample, because most authors never judge the park as a whole. People do that more readily after strong experiences, and the surviving balance is 615 good to 201 bad. This describes pairs where someone rendered a judgment, not all visits.


The headline rests on 65 observations in the staff-service-bad direction, the thinnest cell in the model. The claim is that good service has no detectable effect: its interval runs from about -0.30 to +1.14, so a modest positive effect is not excluded.


The park comparison is complete for three parks. At a 30-pair floor, 23 of 35 park-by-attribute cells are usable. Magic Kingdom, EPCOT and Epic Universe clear it everywhere. All seven clear it on attractions and queues. Value, service and food are thin for the smaller four, so Animal Kingdom's eye-catching 69% on service rests on 16 pairs.

Uncoloured cells fall below 30 rated pairs. They keep their number and their count, so you can see what is being withheld rather than wonder what got dropped.
Uncoloured cells fall below 30 rated pairs. They keep their number and their count, so you can see what is being withheld rather than wonder what got dropped.

Magic Kingdom and Islands of Adventure carry 0.70 attribution precision, with intervals reaching 0.48. Magic Kingdom's residual errors are resort-hotel paragraphs, because r/WaltDisneyWorld is substantially a hotel-planning forum. Islands of Adventure's are multi-park paragraphs, because visitors genuinely treat the two older Universal parks as one destination.


Epic Universe is front-loaded. It opened inside the window, so 33% of its paragraphs fall in its first three months, when novelty and opening-period operations were both unusual.


Reddit is not a representative sample. It skews toward enthusiasts, repeat visitors and passholders.


Interactive version on Tableau Public. Labeled dataset on Kaggle. Code, validation suite and audit tooling on GitHub. Source posts via the Arctic Shift Reddit archive; no post text, titles, usernames or links are published.




































Comments


bottom of page