We Searched 866 App Reviews of an Apple Design Award Winner for “Delight”
Published 3rd August, 2026 by Stuart Hall
We mined 866 reviews of this year’s Apple Design Award winner for “delight.” We found it 5 times.
The feeling was everywhere. The word was almost nowhere. That gap is the whole story.
“Delight” is the most over-worked word in app marketing. It’s in the pitch deck, the onboarding copy, and the investor update. So we ran an experiment on the one app that should be swimming in it.
In this overview you'll learn:
- Nobody Names Delight. They Enact It.
- When “Complexity” Is a Compliment
- Where the Classifier Blinked
- The Tell: Users Start Talking Like grug
- Where the Delight Breaks
- How to Actually Measure a Feeling
- Run This on Your Own Reviews
Want to run this analysis on your app
Try the Appbot MCP, free for 14 days →
grug, made by Netherlands-based studio Ocho, is a daily-wisdom app built around a caveman who dispenses small truths in small words. There are widgets, a drawing feature, and a mascot who talks like he’s chiseling each sentence out of rock. In June 2026 it won an Apple Design Award in the category literally named “Delight and Fun.” Apple put the word on the trophy. If “delight” exists anywhere in the wild, it exists in grug’s reviews.
We pulled every iOS review from April through June 2026: 866 of them, across 72 countries. The app is beloved: a 4.85 average, more than nine in ten reviews five-star, 84.4% positive sentiment. Plenty of delight to go around.

The word “delight” appears in five reviews. 0.6%.
And it gets better. Four of those five aren’t even using the word in the review itself. It’s the one-word title, “Delightful,” slapped on top of a body that never says it again. The bodies talk about quirky design, about smiling, about a hand-drawn battery meter. Even when users reach for the word, it’s a label, not a sentence. The feeling lives somewhere else entirely.
Here’s what 866 people said instead.
A note on the dataset
The award matters. grug collected reviews at a trickle for nine months, then the ADA hit and volume went up roughly 30x in four weeks. Most of this dataset is post-award traffic: people who arrived pre-primed to admire the design, some of whom mention the award by name. That inflates the praise vocabulary somewhat. It doesn’t change the core finding, but honest analysis says it out loud.
Nobody Names Delight. They Enact It.
If you can’t count the word, you count the vocabulary people reach for when they’re feeling the thing. Here’s what grug’s reviewers actually say, counting unique reviews whose title or body contains each term.
These counts come from our analysis snapshot of 851 reviews. A handful of late-arriving reviews have since nudged some of them slightly upward (“love” is now 136, for instance) without changing any share by more than a fraction of a point.
| What users say | Reviews | Share |
|---|---|---|
| love | 129 | 15.2% |
| simple | 106 | 12.5% |
| wisdom / wise | 92 | 10.8% |
| fun | 70 | 8.2% |
| cute | 40 | 4.7% |
| happy | 31 | 3.6% |
| beautiful | 23 | 2.7% |
| funny | 16 | 1.9% |
| smile | 15 | 1.8% |
| warm / cozy / comforting | 12 | 1.4% |
| delight | 5 | 0.6% |
Taken individually, none of these is “delight.” Taken together, 42% of all reviews reach for at least one of them, and that’s measured against all 851 reviews, including the roughly one in five written in a language other than English. Restrict to English reviews and the share climbs past half. (Yes, “love” is the biggest single contributor, and yes, “love” can be generic five-star filler. Drop it entirely and the vocabulary still covers a third of the dataset. The fingerprint survives.)
That’s the point sentiment analysis makes that a keyword search can’t: a feeling isn’t a word, it’s a fingerprint. No single term carries it, so if you’re grepping your reviews for the emotion you’re chasing, you’re measuring your own marketing copy, not your users.
It’s worth separating two lenses that often get treated as one. Sentiment gives you polarity: 84% of grug’s reviews are positive. That number would look much the same for a tax app that never crashes. Emotion tells you which feeling is sitting behind the polarity, where cozy separates from relieved and delighted separates from merely satisfied. Sentiment alone tells you people are happy. The two together tell you what kind of happy, which is the only version a product team can act on.
And the fingerprint has a texture. Look at what’s on the list, and at what isn’t. grug’s delight is cozy, gentle, wholesome, wise. Nobody calls it “powerful,” “premium,” or “game-changing.” The words users chose describe a warm blanket, not a Ferrari. For a product team, that texture is the brief: it tells you which future features will feel on-voice and which will feel like a betrayal.
When “Complexity” Is a Compliment
Here’s where reading topic volume without topic sentiment will steer you straight off a cliff.
The third most-discussed topic in grug’s reviews is Complexity, with 108 mentions, right behind “Satisfied users” (373) and “Design & UX” (178). For almost any app in the store, a Complexity bucket that size is a five-alarm fire: it’s the pile where “confusing,” “bloated,” and “I gave up” go to accumulate.
Not here. Of those 108 Complexity mentions, 101 are positive (94%), with an average rating of 4.98. Users aren’t complaining about complexity; they’re praising its absence. Simplicity is grug’s product, and the topic model is capturing people celebrating it:



A team scanning a bar chart of topic counts would see “Complexity is huge” and start planning a simplification sprint for an app that’s already winning on exactly that axis. The sentiment layer is the difference between reading the topic and reading the room.
Where the Classifier Blinked
One confession, because it proves the larger point. The single best delight testimonial in the entire dataset, a long, considered review praising the hand-drawn battery meter as “a nudge and a wink between me and grug,” got auto-tagged neutral by our sentiment model. Why? Because it’s written in measured, descriptive prose. No exclamation points, no superlatives, just a person carefully explaining why an interface detail made them feel something.

Classifiers read intensity more easily than they read depth. That’s not a reason to skip sentiment analysis. It’s a reason to treat the tags as a map, not the territory, and to actually read the reviews the map points you toward. The tooling’s job is to make 851 reviews navigable, not to spare you from ever opening one.
The Tell: Users Start Talking Like grug
This is the finding no star rating could ever surface.
Roughly 10% of reviewers write their review in grug’s own voice, dropping articles, mangling grammar on purpose, speaking as the caveman:

That reviewer replaced Tolstoy with a cartoon caveman and reported it in the caveman’s grammar. And they’re not alone:



You cannot fake this, and you cannot buy it. When a user adopts your product’s voice unprompted (in a review, the most transactional surface there is), they’ve stopped evaluating the app and started playing inside it. That’s not satisfaction. That’s identification. It’s the highest-grade delight signal in the entire dataset, and it is completely invisible to a rating average, a sentiment score, or a keyword search. It only shows up when something is actually reading the language.
Where the Delight Breaks
An 84% positive dataset makes the other 16% precious, because the friction that survives near-universal love tends to be real. grug’s roughly 70 non-five-star reviews cluster into three clean, fixable groups, and none of them is “the app is bad.”
1. Language Is the Wall
grug’s charm is minimal English, which is a beautiful conceit until you don’t read English. Nineteen percent of reviews arrive in another language, and the unhappiest ones are variations on “Je comprends pas l’anglais,” “没看懂,” “모르겠어요,” plus explicit pleas for a Chinese version. The thing that delights native speakers is the exact thing that locks everyone else out. Given how much of the audience sits in China, India, and Korea, localization isn’t a nice-to-have; it’s the single largest ceiling on the app’s delight.
2. Onboarding Loses a Few at the Door
A small but distinct group never got started: “very difficult to figure out how to work it… I still never got in.” “Don’t know what to do.” Delight requires a first moment, and a handful of users never reach it.
3. Accessibility Is a Real Gap
One three-star review says it all in its title: “Love the design, but not accessible.” A visually impaired user who loves everything grug is doing reports that VoiceOver and Select to Speak don’t work at all. That’s delight blocked by a fixable defect, a would-be evangelist locked out by the very hand-drawn aesthetic everyone else is celebrating. It’s exactly the kind of review that vanishes in a sea of praise unless something is watching the bottom of the distribution for you.
None of these is loud. All of them are actionable. That’s the quiet value of sentiment work on a happy app: it finds the roughly 70 reviews worth acting on inside the 866 that just say “grug good.”
How to Actually Measure a Feeling
grug is an unusually pure test case, a product that sells an emotion and nothing else, so it makes the lessons unusually clean:
- Don’t count the word; map the vocabulary. “Delight” was in 0.6% of reviews (and mostly in the titles at that). The feeling was in 42%. The word your marketing uses is almost never the word your users use.
- Never read topic volume without topic sentiment. grug’s biggest “problem” topic was 94% praise.
- Guard the bottom of the distribution. In an 84%-positive dataset, the unhappy 16% is where the roadmap is: localization, onboarding, and accessibility were all sitting in roughly 70 reviews.
- Read the language, not just the score. The best signal in the whole set has no numeric proxy at all.
Because here’s how this story ends: an app won an Apple Design Award in a category named “Delight and Fun,” 866 people showed up to say how it made them feel, and almost none of them used the word. They didn’t need it. They just quietly started talking like a caveman.
That’s better than “delight.” That’s the review equivalent of a standing ovation, and if your tooling only produces averages, you’ll never hear it.
Run This on Your Own Reviews
One last practical note: nothing in this analysis required a data export, a spreadsheet, or a line of code. It took three questions and five minutes with the Appbot MCP. We ran it through Appbot’s MCP server, a connector that lets an AI assistant like ChatGPT or Claude query your app review data directly in conversation.
Once connected, the assistant can list your tracked apps, pull aggregate app review stats (topics, emotions, sentiment, and words and phrases over any date range), and fetch the reviews themselves, filtered by keyword, star rating, sentiment, language, or version.
If you want to know what your users are enacting rather than what they’re rating, connect the MCP server and ask.
Want to run this analysis on your app
Try the Appbot MCP, free for 14 days →Where to from here?
- Discover effective strategies for app review management to efficiently handle and leverage user feedback.
- Unlock valuable insights into user sentiment with our powerful sentiment analysis tool for informed decision-making.
- Simplify your review tracking process with our efficient review aggregator, providing a centralized view of user feedback.
- Engage with your users effectively by crafting thoughtful responses with our convenient Reply to App Store Reviews feature.
About The Author

Stuart is Co-founder & Co-CEO of Appbot. Stuart has been involved in mobile as a developer, blogger and entrepreneur since the early days of the App Store. He built the 7 Minute Workout app in one night and blogged the story of growing the app to 2.3 million downloads before exiting to a large fitness device company. Previously he was the co-founder of the Discovr series of applications which achieved over 4 million downloads. You can connect with him on LinkedIn.
Enjoying the read? You may also like these
Why Generic AI Isn't Enough for App Ratings and Review Feedback Analysis Learn why generic AI alone isn't enough for customer app feedback analysis and how combining LLMs with structured customer intelligence delivers more reliable product insights.
Learn how Apple’s new Retention Messaging API helps reduce app subscription churn with personalized, in-the-moment messages. Backed by research, this guide shows how to combine App Store feedback, discounts, and smart messaging to boost retention.
A complete guide to App Store product page optimization in 2026 — custom product pages, A/B testing, and metadata that converts more installs.
What the new Apple and Google age-verification requirements mean for your app, and which APIs to use to comply without hurting the user experience.
