Why do apps get one-star reviews?
Crashes and broken features, ahead of everything else. And iOS users complain about crashes at nearly twice the rate Android users do, which is worth knowing before you decide the app is the cheap half of the project.
We went looking for this while scoping a product where the rating was the whole business case. The useful finding is that most of what sinks a rating is decided before any code is written, and it is decided by people who are not engineers.
What one-star reviewers actually complain about
The most careful source here is a study that hand-tagged 9,902 one and two-star reviews across 19 apps published on both the App Store and Google Play, sampled twice about a year apart. Hand-tagged matters: the categories are read by people rather than inferred by sentiment analysis, which is where most review-mining goes wrong.
On both platforms, the single most common complaint is a functional error, meaning a feature that does not do what it says. App crashing sits immediately behind it. Between them they account for the bulk of what people are angry enough to write about.
The iOS penalty is real and it is large
| iOS | Android | |
|---|---|---|
| Complaints about the app crashing | 23.14% | 13.55% |
For more than 59% of the apps studied, users complained about crashes more often on iOS than on Android. The same app, the same bug, a harsher response on one platform.
There is a sharper version of this. Across the apps whose complaint profiles differed by platform, most did not even share the same leading complaint: Amazon’s reviewers complained most about crashing on iOS and about functional errors on Android. Star ratings alone will not show you this. Two platforms can carry the same score for different reasons, which means a single number cannot tell you what to fix.
The choices that predict a bad rating
The second source rated 1,496 mobile games, drawn from an initial set of 52,111, for manipulative design, producing 85,388 rated instances. It sorted them into 843 the raters called dark and 653 they called healthy.
The gap is not subtle. Counting temporal mechanics, the ones that work by making you show up at a particular moment, daily rewards, appointment play, timers and grinding:
| Rated dark | Rated healthy | |
|---|---|---|
| Games in the set | 843 | 653 |
| Temporal mechanics per game | 22.09 | 4.85 |
| Free to play | 96.8% | 53.0% |
| Includes in-app purchases | 93.6% | 54.0% |
| Carries advertisements | 52.4% | 37.3% |
Roughly 4.5 times the temporal mechanics, and almost every badly-regarded game is free to play. That second number is easy to misread, so it is worth being careful: free to play does not cause a bad rating. Slightly over half the well-regarded games are free too. What the data says is that the mechanics people resent almost all live inside free-to-play, so choosing that model is choosing to work near them.
One more figure from that study is worth sitting with. Only 10.76% of all 1,496 games had no reported manipulative instances at all. Shipping something with none of this in it puts you in a very small group, which is either a warning or an opening depending on what you are building.
Push and pull, which is the distinction that matters
The mechanic is not what decides it. The tuning is. A CHI PLAY study of engagement rewards found players experience them dualistically: the same daily reward reads as motivation to one player and as an obligation or a chore to another, and the clearest harm shows up when someone misses out on something rare or time-limited.
Which gives a usable rule. Reward the outcome of using the thing, finishing something, hitting a goal, rather than rewarding attendance. Make streaks forgiving. Never punish an absence. The difference between a feature people like and the same feature people resent is usually whether it pulls them back or pushes them.
What this changes if you are commissioning an app
Budget for stability as a feature, not as polish. Crashes and broken features are the top two complaint categories, and on iOS the crash penalty is close to double. Regression testing on real devices before each release is not a nice-to-have line item; it is the line item that protects the rating.
Decide the monetization model with the rating in mind, because the mechanics that correlate with resentment cluster in one of them. If the model is free to play, the mechanics to avoid are known in advance and can be written into the brief.
Do not read a single star rating as a diagnosis. The same score on two platforms frequently means two different problems.
Method, and what would change the answer
This is a reading of other people’s measurements, not our own, and it is bounded accordingly. The review study covers 19 cross-platform apps, all free, all popular, in two snapshots. Popular free apps are not all apps, and a small business tool has a different reviewer than a consumer app at scale. The crash percentages quoted are from the first of the two snapshots.
The dark-patterns study covers games only. We think the push-versus-pull finding generalizes to any product with notifications and streaks, but that is our inference and not something the paper tested.
Two figures that appeared in our internal brief are not on this page, because when we went back to the source we could not confirm them. We would rather the page be shorter.
Sources
Khalid et al., Studying the consistency of star ratings and the complaints in 1 and 2-star user reviews (Empirical Software Engineering). Dark patterns in mobile games at scale (MUM 2024). Daily Quests or Daily Pests? The Benefits and Pitfalls of Engagement Rewards in Games (CHI PLAY 2022).
We read this before scoping an app, not after shipping one. If you are weighing something similar, we would rather tell you what it will not do before you pay for it.
Tell us the problemAll research