E-COMMERCE
Why a Marketplace Star Rating Doesn't Tell the Whole Story
The rating dropped. But why?
Say the weekly report shows that one of your best-selling products has slipped from 4.6 to 4.3 stars on a marketplace over a few weeks. The category team asks the product team, the product team points to shipping, and operations points to the seller. Everyone argues from the handful of reviews they happen to remember. The meeting usually ends with “let's keep an eye on the reviews.”
The short answer: a star average rolls very different causes into a single number. A product defect, a wrong size chart, a late or damaged delivery, an unresponsive seller and an overstated description can all pull the same rating down. Until those causes are separated, the rating does not tell you what needs fixing. To see the cause, you need to read reviews along three axes: which product or variant (SKU), which topic and which period.
In this article we look at what a star average hides, how teams try to fill that gap today, and a practical method for reading reviews more reliably.
One number can send the wrong team into action
The rating belongs to the product page, but its causes are spread across teams. A quality problem belongs to product and sourcing, a wrong size chart or misleading image to the content team, a damaged parcel to logistics, an unanswered customer message to the seller or channel team. Because the average makes no such distinction, it is unclear who owns the problem when the rating falls.
The cost of that ambiguity is misdiagnosis. Debating a product reformulation when the drop comes from delivery, or switching carriers when the issue sits in one production batch, burns time and budget. The real cause stays where it is.
The rating itself still matters. In an analysis spanning more than 40 categories, Northwestern University's Spiegel Research Center found that purchase likelihood peaks when the average rating sits between 4.2 and 4.5, and starts to fall as it approaches 5. So the goal is not to push the stars as high as possible; it is to know what moves them and to route that knowledge to the right team.
Why averages, quick reads and word clouds fall short
Teams usually try to fill the gap in one of three ways: tracking the average, reading the latest reviews by hand, or looking at a sentiment score or word cloud. Each shows something, but none answers the question of why.
An average hides the distribution. Hu, Pavlou and Zhang (2009) showed that online product reviews tend to follow a J-shaped distribution: many 5-star reviews, a noticeable number of 1-star reviews and few in between. A large share of reviewers are either very satisfied or very unhappy. The same 4.2 average can therefore stand for two entirely different situations.
Ratings are aggregated at the product page. Color, size or volume variants often share one page. A defect in one variant gets diluted by good reviews of the others and is hard to spot in the page average. When several sellers sell the same product, the picture gets even muddier.
Seller ratings and product ratings are separate; review text is not. Major marketplaces keep a seller or store score that tracks delivery and operational performance alongside the product rating. Yet customers usually write about a late parcel or a crushed box in the product review and give the product a low score. The product rating is shaped by causes outside the product too.
Reading by hand does not scale and is selective. Nobody can regularly read several hundred reviews. The ones that do get read tend to be the newest and the harshest, and how often a complaint repeats stays invisible.
Sentiment scores and word clouds don't tell you the topic. Frequent mentions of “shipping” can signal praise or complaints. A rising share of negative sentiment doesn't tell you which topic or which variant it is rising in.
The time context disappears. Seasonal sales volume, a new production batch, a packaging change or a new seller on the page can all affect the rating from a specific date. The overall average flattens those breakpoints.
Reading reviews by SKU, topic and period
A more reliable reading looks at reviews not one by one but where three axes intersect. Each axis rules out a different misdiagnosis.
| Axis | The question it asks | The misdiagnosis it rules out |
|---|---|---|
| Product / SKU | Is the drop across the product or in one variant? | Blaming the whole product for one variant's problem |
| Topic | Is the complaint about the product, its presentation, delivery or the seller? | Pinning a delivery- or seller-driven drop on the product |
| Period | When did the topic start, is it ongoing, does it coincide with an event? | Mistaking a short spike for a lasting problem, or the reverse |
The logic of this reading can be summed up in five steps:
- Put the scope in writing. Which products, which channels, which date range? Note the total number of reviews. A finding without a clear scope falls apart at the first question in the meeting.
- Map reviews to products and variants. Group the same product's pages across channels under one product, and keep variant information where it exists.
- Use one shared topic list. Keep product topics (quality, durability, size and fit, use), presentation topics (whether the description and images match reality), delivery topics (delays, damaged or missing items) and seller topics (communication, order handling) apart. One review can belong to more than one topic.
- Break the rating drop down by topic. Among low-rated reviews, which topic's share grew? Always report the percentage alongside the review count.
- Compare against the period and keep the evidence. Set the before and after of a topic next to known events such as a campaign, a new batch or a packaging change. Every finding should sit next to the raw reviews that support it.
A team that reads reviews this way can replace “the rating dropped” with narrower, testable statements. For example: “Most of the drop is concentrated in one variant, in packaging-related reviews, after a specific date.” That sentence already carries who it should go to, what to check and which reviews to read.
Where Prodora fits
On paper, the steps above look simple. The hard part is applying them across thousands of reviews, every week, consistently. Teams rarely run this reading on a regular basis, and the reason isn't a lack of method. It's these practical obstacles:
- Product matching. The same product is listed under different names and different variant structures across channels. Tying each review to the right product and SKU is a job in itself.
- The range of language. Customers describe the same problem in very different ways. Slang, irony, typos and abbreviations leave keyword-based methods short.
- Reviews with several topics. A review like “Great product, but the box arrived crushed and the seller never replied” is positive about the product and negative about delivery and the seller. Each part has to be counted separately.
- Consistency over time. If a complaint tagged “packaging” this month is counted as “shipping” next month, period comparisons lose their meaning.
- The link to evidence. Once a finding is separated from the raw reviews behind it, even the best analysis weakens at the first question in the meeting.
Prodora treats this not as a one-off analysis but as infrastructure that runs continuously. It collects reviews from marketplaces, complaint platforms and social media, maps each review to the relevant product and SKU, and tags it with the same topic taxonomy, so this week's data is genuinely comparable with last month's. Product-driven complaints are kept apart from delivery- and seller-driven ones, and every finding arrives with the raw reviews it rests on.
The result is that teams spend their time interpreting the data rather than preparing it. What reaches the meeting is a finding with a clear scope and evidence, not an impression. The decision still belongs to the team; Prodora keeps the ground under that decision solid.
Limitations: what not to conclude from reviews
Breaking reviews down helps you ask better questions, but it doesn't give a definitive answer to every one. Keep these points in mind before sharing findings.
- Reviewers don't represent all customers. The J-shaped distribution is a reminder that extreme experiences are over-represented and middling ones under-represented. A topic's share of reviews is not its share among all buyers.
- Small numbers mislead. For a variant with few reviews, a handful of reviews can swing the ratio sharply. Set a threshold and always report percentages with the raw count.
- Coinciding is not causing. A rating drop may line up with a campaign or a new batch. That overlap is a hypothesis; confirming it requires internal records.
- A shipping complaint is not always the seller's fault. Delivery problems can come from the carrier, the marketplace's own fulfillment network or packaging standards.
- Automated classification is not error-free. Irony, slang, typos and several topics in one review make classification harder. Before an important decision, read a sample of the relevant reviews.
- Review data is not enough on its own. Findings gain meaning when read alongside the team's own data, such as sales, customer service and quality records.
Watch the rating, look for the cause in the reviews
A star average is a good warning light, but it is not a diagnostic tool. When the rating falls, the question to ask is not “how many points did we lose?” but “in which product, on which topic, since when?” Teams that can answer that with evidence get the problem to the right owner sooner.
We explain how we analyze marketplace reviews by product, topic and period in more detail on our E-Commerce solution page. If you'd like to look at your own products' reviews this way, you can book a call with our team.
Sources
- Hu, N., Pavlou, P. A. and Zhang, J. (2009). “Overcoming the J-shaped distribution of product reviews.” Communications of the ACM, 52(10).
- Medill Spiegel Research Center, Northwestern University. How Do Star Ratings and Review Content Influence Purchase? (analysis based on PowerReviews data).
Let's listen to your brand's consumer voice together
Start with a small pilot and see the value in your own data.
Request a demo