Reviews are not one dataset. They are a rating, a text, a timestamp, a location, and a reviewer, all describing one thing that already happened. This lesson makes the case that those layers have to be analyzed as a system, and it is equally clear about when the system is the wrong tool.
Lecture video arrives here.
The chapter's own argument, distilled. If you read nothing else on this page, read this.
A survey asks what someone intends to do. A review is written after the experience: the person went, something happened worth reporting, and they composed text about it for strangers. That temporal ordering is where the inferential advantage over stated-preference data comes from, and it is a different advantage from scale. The intention-behavior gap is well documented; even Ajzen's theory of planned behavior, which treats intention as the strongest proximal predictor, concedes that intentions do not fully determine action, and Morwitz and colleagues showed that asking about purchase intent can itself change behavior. None of this makes reviews clean. They carry social desirability from public posting, platform norms that shape what gets written, extremity bias because people with strong opinions post more, and the selection effect of who bothers to review at all.
At the surface is the star rating, a structured judgment you can average, distribute, and turn into a time series. Underneath is the text, which carries the reasoning behind the number: a three-star review saying the food was excellent but the service unbearably slow is not the same three-star as everything was fine but nothing special. Then come the metadata layers that survey work rarely gets. Timestamps let you watch satisfaction evolve. Coordinates let you ask where complaints cluster. Reviewer history, account age, and helpfulness votes expose selection and credibility, and helpfulness voting in particular shows which reviews the community treats as authoritative. The claim of the book is that these are facets of one phenomenon, not five separate datasets.
Luca's work on Yelp associates a one-star increase with 5 to 9 percent higher restaurant revenue, with the size varying by restaurant type and by how uncertain the market is. A 2023 survey put the share of consumers who read reviews for local businesses at 98 percent. The pattern repeats but does not transfer unchanged: in healthcare a rise in a physician's average rating is associated with higher patient volume, and in financial services 78 percent of consumers consult reviews before choosing an institution. The reason to care is not only commercial. When reviews signal quality accurately, markets get more efficient; when they are manipulated, consumers are harmed.
This is the fragmentation problem, and it is the reason the book exists. Econometricians regress ratings on outcomes and ignore the text. Natural language processing researchers find structure in the text and disconnect it from business performance. Spatial analysts surface geographic pattern without temporal context. Behavioral scientists test theory but struggle at scale. Forecasters predict without explaining the driver. Each tradition produces real findings, and each misses what the others see, so the literature accumulates partial answers that do not connect to each other. A practitioner reading across all five cannot assemble them into one account.
Machine learning extracts sentiment, topics, and features from text. Panel regression relates those features to outcomes while holding entity-level differences constant, which is possible because review data naturally forms a panel: the same entities observed over many periods, so a unit can be compared against its own earlier behavior instead of against other units. Time-series analysis exposes the temporal dynamics and any leading indicator. Spatial econometrics detects geographic structure and diagnoses plausible sources of it. Structural equation modeling tests the mechanism underneath. Take a chain of fifty properties: cleanliness complaints concentrate at some properties, are associated with lower revenue, spike weeks before occupancy falls, may or may not genuinely spill across neighbors, and may damage revenue through trust, through ratings, or both. Note the honest boundary: integration of this kind improves explanation, and it is still not a causal claim without the identification strategies of Chapter 7.
Each method has a data requirement. Panel regression and time-series analysis need repeated observations across many periods. Spatial diagnostics need many geographically dispersed units. Structural equation modeling needs a sample large enough to estimate a covariance structure stably. With a handful of entities, a single cross-section, or one point in time, most of the machinery simply does not apply, and the honest move is a careful descriptive analysis paired with one well-chosen method. Forcing the full stack onto thin data produces confident-looking output with nothing underneath it. The same discipline governs how results get read: when clustering appears on a map, similar businesses selecting into the same place is a live rival explanation, and a non-significant residual Moran's I under one specification is consistent with that compositional account without proving it or ruling out every spatial process.
No sample answers here: your output, your interpretation, is the exercise.
Before any method, you need to know what your raw material actually is, and that depends on where it came from. Pick one business you know well that appears on at least two review platforms, for example Google Maps and Yelp, or TripAdvisor and a booking site. Open both in your browser and look at them properly before you ask the AI anything. You are not collecting data yet; you are learning what each platform would and would not give you.
Act as a research methodologist advising me on data sources. I am planning a study using online review data about [type of business, for example independent restaurants / hotels / physicians / banks] and I am comparing two platforms as possible sources: [platform A] and [platform B]. Here is what I can see on each one for the same business. [Platform A]: the rating scale is [describe it], the number of reviews shown is [number], and each review appears to include [list what you can actually see: text, date, reviewer name, reviewer history, photos, helpfulness votes, management response, verification badge, anything else]. [Platform B]: the rating scale is [describe it], the number of reviews shown is [number], and each review appears to include [same list]. Do four things. First, build a table comparing the two platforms on the layers that matter for analysis: the rating, the text, the timestamp, the location, and the reviewer. For each layer, say what is available, what is partially available, and what is absent. Second, name the selection effect specific to each platform. Who ends up writing reviews there, and who does not, given how that platform is used and who goes looking for it. Third, tell me which research questions each platform could support and which it could not, and be concrete about why. Fourth, tell me where you are inferring rather than knowing. You cannot see these pages. Mark clearly which of your claims come from my description, which come from general knowledge of the platform, and which are guesses.
Paste it into your own ChatGPT, Claude, or Gemini window and run it. Replace anything in square brackets with your real numbers first.
This is the lesson's deliverable and the spine of everything that follows. You need one real question, not a topic. A topic is guest satisfaction; a question names an outcome, a unit, and a comparison. Write yours before you open the AI, then hand it over along with an honest account of the data you can actually get, including what you cannot get.
Act as a methods advisor who is willing to tell me my plan does not work. I want to use online review data to answer this question. My question: [write it as one sentence with an outcome, a unit of analysis, and a comparison. For example: does a decline in service sentiment predict a fall in monthly revenue across independent restaurants in one city?] The data I can realistically obtain: [number] entities, observed [across how many time periods, or a single snapshot], with [do you have review text, or only ratings?], located [are they geographically spread, or all in one place?], and with an outcome measure of [what outcome can you actually observe, and where does it come from, or say that you have none]. Work through the five methods in this framework one at a time: machine learning for extracting features from text, panel regression for relating features to outcomes while controlling entity-level differences, time-series analysis for temporal dynamics, spatial econometrics for geographic structure, and structural equation modeling for mechanisms. For each of the five, give me three lines. What it would add to my specific question. Whether my data can actually support it, answered yes or no, with the specific data requirement that decides the answer. If no, what the minimum change to my data collection would be to make it possible. Then do two harder things. First, tell me whether the full framework is warranted here at all, or whether the honest answer is a careful descriptive analysis plus one well-chosen method. Say which single method, if so. Second, tell me the strongest claim my data could support and the weakest claim I would be tempted to overstate it into. Do not be encouraging. If my question cannot be answered with what I have, say that plainly and tell me what question the data could answer instead.
Paste it into your own ChatGPT, Claude, or Gemini window and run it. Replace anything in square brackets with your real numbers first.
The chapter opens with a warning: clustering on a map is easy to see and easy to over-read. When high-rated restaurants sit together it is tempting to conclude that they lift each other, when often similar businesses simply selected into the same place. The habit that protects you is generating the rival explanation yourself, early, while it is still cheap to change the design. Take one pattern you already believe about your own domain.
Act as a hostile but fair reviewer for a methods journal. I am going to give you a claim I believe about review data, and I want you to attack the inference, not the writing. My claim: [state the pattern and the interpretation you attach to it. For example: restaurants whose service sentiment falls in one quarter lose revenue the next quarter, so falling service sentiment causes the revenue loss.] The evidence I would have: [describe what you would actually measure, over what units, over what period, and how the outcome is observed.] Give me the following. First, three rival explanations that would produce exactly the pattern I described without my interpretation being true. Cover at minimum: selection into the sample, composition (the units differ in stable ways I have not modeled), and reverse causality (the outcome moving first, and the sentiment following). Second, for each rival, the specific observable consequence that would distinguish it from my explanation. Something I could actually check, not a suggestion to control for it. Third, the single piece of evidence that would most change your mind, and whether my described data contains it. Fourth, rank my claim on a scale from descriptive, to associational, to causal, and say what would have to be true for it to move up one step. Be specific to my claim. Generic advice about correlation and causation is not useful to me.
Paste it into your own ChatGPT, Claude, or Gemini window and run it. Replace anything in square brackets with your real numbers first.
Every answer comes from this chapter. Pick one; a wrong pick tells you where to look and lets you try again.
01The chapter argues that reviews carry an inferential advantage over survey data. Where does that advantage come from?
02Luca's study of Yelp is cited for which specific finding?
03In the chapter's table of the integrated framework, what is named as panel regression's limitation when it is used without the other methods?
04Maple City, the running dataset for the whole book, is fully synthetic. What reason does the chapter give for that being an advantage?
05In Maple City, a service-improvement program was rolled out to some restaurants and raised ratings by about 0.30 stars. According to the chapter, what makes that effect visible?
06Sentiment in Maple City clusters geographically at a Moran's I near 0.31. After the cuisine, price, and aspect controls are added, residual clustering is small and non-significant. What does the chapter say that result establishes?
07When does the chapter say the integrated framework should not be used?
One question, honestly matched to the methods your data can and cannot support. This is the document the rest of the course revises rather than replaces: Chapter 4 sharpens it into an analysis plan, and the methods chapters fill in the rows you marked as possible.
You now have your research question and method map for your own data. Paste it below and we will read it and write back with what we would tighten first: the question, the identification, or the measurement. No charge and no pitch, and if your design is already sound we will say that instead.
Goes to Jeong-Yeol Park directly. Nothing is published, and your notes stay on this device unless you press send.