How this is measured
Every number on this site comes from public app-store reviews. Here is exactly how, including what the method cannot see.
1. Only apps people pay for
Every app in the corpus has a real paid tier. A complaint from someone using a free product proves nothing about willingness to pay, and willingness to pay is the whole question. Apps are seeded by name and resolved to store IDs programmatically, then verified: the resolved title must contain the vendor name, because searching for “Protractor” shop software returns a measuring tool with 10,413 ratings, and searching “Cornerstone” veterinary returns HR software. Twenty-eight such false matches were rejected in one sweep alone.
2. Reviews, across four storefronts
Apple's public review feed caps at roughly 500 reviews per query, so each app is swept across four English storefronts (US, UK, Canada, Australia) and two sort orders, then deduplicated by review ID — which yields around a thousand per app instead of five hundred. Apple returns every star rating, so it is the only source that gives a true low-star share. Google Play is queried per star rating separately, because its paging otherwise buries low-star reviews under the four-and-five-star mass.
3. Substantive reviews only
Percentages are computed over reviews of 60 characters or more. A review reading “Terrible” carries a rating but no diagnosable complaint; counting it would deflate every rate while adding nothing. Roughly 12% of low-star reviews fall below the threshold.
4. A 23-theme complaint taxonomy
Each review is tagged with every theme it matches — compound complaints are the valuable ones. The themes group into three buckets that mean different things commercially:
| Bucket | Themes | What it tells you |
|---|---|---|
| Payer pressure | billing & cancellation, price & value | The reviewer is paying and disputing. The strongest signal there is. |
| Trust breakdown | data loss, sync failure, support, reliability | The incumbent has broken the relationship. |
| Buildable wedge | missing capability, manual entry, complexity, search, reporting, integrations, permissions, platform parity, export | What a small team could actually out-execute. |
The patterns that matter most are the ordinary ones. “Doesn't work” is by far the most common phrasing people use, and an early version of the taxonomy missed it entirely — which cut the match rate from 59% to 46% until it was fixed.
5. Scoring
Niches are ranked on
0.35 · buildable wedge + 0.25 · payer pressure + 0.25 · trust breakdown + 0.15 · low-star
share, with weights renormalised when a component is missing — otherwise a niche whose
collection had not finished scored zero on it and ranked low for a reason that was not real. A
400-review floor applies: below that the ranking is noise, and one niche briefly topped the table
on 304 reviews before evaporating with more data.
6. Then we try to break the finding
A theme found in a narrow sample has to hold its share as the sample widens. If it shrinks, it was an artefact of which apps happened to be sampled. This is not a formality — it killed our best idea, and that story is here.
The four limits on every number here
- About 41% of substantive reviews go untagged. They are idiosyncratic one-off narratives. Every percentage published is therefore a floor, not a true rate. The miss rate is roughly uniform across niches, so comparisons between niches remain valid.
- App-store reviews over-represent the daily user. The buyer complains on G2, Capterra and in churn calls. Those sources block automated collection and are absent here — which matters most for business software, where buyer and user differ.
- No recency weighting. A 2021 complaint counts like last week's. Some cited failures are certainly fixed by now.
- Store coverage is uneven. Some products are iOS-only in practice; some Play packages could not be resolved. Where a niche leans on one store, its numbers inherit that store's bias — Play, for instance, reports reliability complaints several points higher than Apple for the same products.