World-as-Lab Field Experiment
Test operating assumptions in the real setting where results must work
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 5
- Confidence
- 97%
List's research agenda treats the world as the laboratory. Begin with a consequential operating claim that practitioners currently support with intuition, then test alternative treatments in the real market rather than relying only on an artificial task. The field setting preserves features that can change decisions: the actual population, experience, anonymity, observation, market selection, and competition. Measure behavior, compare conditions, and investigate mechanisms and side effects so the result can improve the next decision. When deliberately assigning a harmful condition would be unethical, use naturally occurring variation and construct a credible comparison, as List did with good and bad Uber trips. The final discipline is replication: determine where the result persists and where a different audience or circumstance changes it before scaling.
Origin
List began this approach at baseball-card conventions and expanded it through charities, governments, schools, Uber, and Lyft.
Core principles
- 01Rules of thumb without data are hypotheses, not evidence
- 02The people and circumstances of the real market affect behavior
- 03A useful experiment reveals both what works and why
- 04Replication becomes valuable when an organization has an incentive to avoid waste
How to run it
- 1
Expose the assumption
Write the rule of thumb or causal claim currently guiding action. Ask what data actually supports it.
Pro tip Start with a decision that consumes meaningful money or effort.
Watch out Tradition is not a control condition.
- 2
Enter the real setting
Run the study with the people, incentives, experience, and constraints found in the target market. Preserve normal behavior as far as the design allows.
Pro tip Partner with the organization already operating the process.
Watch out A convenient sample may behave differently from the eventual audience.
- 3
Build the comparison
Assign or identify treatment and comparison groups that differ on the intervention of interest. When random assignment would cause harm, seek a defensible natural comparison.
Pro tip Define the comparison before inspecting outcomes.
Watch out Do not manufacture bad experiences merely to obtain a control group.
- 4
Measure behavior
Choose outcomes tied to the real objective, then record effects, persistence, and unintended consequences. Separate observed behavior from stated preference.
Pro tip Use organizational data to extend the observation window.
Watch out A short-term lift can hide later depreciation or spillovers.
- 5
Replicate and bound
Repeat the test across audiences or places and identify where the effect changes. Scale where the evidence carries and redesign where it does not.
Pro tip Treat heterogeneity as useful information, not an inconvenience.
Watch out One successful market does not establish universal validity.
In the wild
Fundraising specialists told List that a three-to-one match should outperform a one-to-one match but could not provide data. His charitable-giving field experiments found that offering a match matters, while the stated ratio matters less than practitioners assumed.
→ Charities gained evidence for designing appeals instead of relying on inherited rules of thumb.
List did not deliberately give riders bad trips. His team found a rider who experienced a bad trip and a statistical twin traveling over the same route at the same time who received a good trip, then compared their spending over the next 90 days.
→ The analysis showed that bad trips caused substantial lost revenue and created a base for testing remedies.
Common mistakes
Treating anecdotes as evidence
Repeated professional practice can persist without a valid comparison showing that it works.
Generalizing from the wrong setting
A laboratory can change the people, observation, experience, and selection conditions that govern real behavior.
Testing harm on purpose
When a treatment would create a bad customer experience, use naturally occurring data and a credible comparison instead.
Is it for you?
Best for
Organizations able to compare interventions during real customer, worker, donor, or policy activity.
Not ideal for
Situations where assigning the treatment would deliberately expose people to serious harm or where no valid comparison can be formed.
From the transcript
“my research agenda has always been to use the world as my lab”
“if you're ever trying to accomplish something do it scientifically so you can not only figure out what works and why but then you can…”
“in the field you have not only the right people but you also have the right set of circumstances”
From the episode
#566: John List — A Master Economist on Strategic Quitting, How to Practice Theory of Mind, Learnings from Uber, Optimizations to Boost Donations, the Primitives of Decision-Making, and How Field Experiments Reveal Hidden Realities
John List