Do gamified fitness apps actually work?
Yes, with a real but shrinking effect once the game stops. A 2022 meta-analysis of 16 randomized trials by Mazeas and colleagues, published in the Journal of Medical Internet Research, found gamified programs raised physical activity with a "small to medium" pooled effect (Hedges g = 0.42, 95% CI 0.14-0.69) across 2,407 participants. Two JAMA Internal Medicine trials from the same U.S. research group measured the same pattern directly: daily step counts rose during a 12- to 24-week gamified intervention, then fell after the game ended, in one trial dropping to a difference no longer statistically significant from a plain step tracker.
How four studies measured it
This table is this page's own compilation, built by lining up each source's own reported design, sample, intervention and result; none of the four studies below published a table shaped like this.
| Study | Design & n | Intervention | Outcome | Follow-up |
|---|---|---|---|---|
| Mazeas et al., 2022 (JMIR) | Meta-analysis, 16 RCTs, 2,407 participants | Various gamified physical-activity apps/programs vs. inactive or active (non-gamified) comparators | Pooled Hedges g = 0.42 (95% CI 0.14-0.69); g = 0.58 vs. inactive controls; g = 0.23 vs. active non-gamified controls | Averaging 14 weeks post-intervention: g = 0.15 (95% CI 0.07-0.23), weaker but still measurable |
| Patel et al., 2017, BE FIT (JAMA Intern Med) | RCT, 200 adults / 94 families | 12-week family gamification (points, levels for step-goal days) vs. step tracking with feedback only | Step-goal-achievement days: 0.53 vs. 0.32 (adj. diff 0.27, 95% CI 0.20-0.33, p<.001); daily steps above baseline: +953 vs. control (95% CI 505-1401, p<.001) | 12 weeks post-game: daily-step advantage fell to +494 (95% CI 170-818, p<.01), about half the peak, still significant |
| Patel et al., 2019, STEP UP (JAMA Intern Med) | RCT, 4 arms, 602 adults (151 control, 151 support, 150 collaboration, 150 competition) | 24-week gamification, three social-incentive designs vs. step tracking with feedback only | Daily steps above control: competition +920 (513-1328, p<.001); support +689 (267-977, p<.001); collaboration +637 (258-1017, p=.001) | 12 weeks post-game: only competition stayed significant vs. control (+569, p=.009); support (+428, p=.04) and collaboration (+126, p=.49) were not, per the trial's own conclusion |
| Davis et al., 2021 (J Clin Nurs) | Systematic review, 7 studies, 657 patients at high cardiovascular-disease risk | Gamified apps using 2+ game tactics for secondary prevention (heart disease, hypertension, stroke, type 2 diabetes) | More improvement in physical activity, HbA1c and diabetes self-management empowerment vs. comparators; no added benefit over usual care for blood pressure, BMI, heart-failure self-management or medication adherence | App acceptability "declined with time," though ratings for challenges, leaderboards and badges stayed high; enjoyment highest for "surprise/novelty…" |
What the evidence found
Alexandre Mazeas and colleagues searched five databases for randomized controlled trials of gamified physical-activity interventions published from 2010 to 2020, and pooled 16 trials covering 2,407 participants in a 2022 meta-analysis published in the Journal of Medical Internet Research. The main analysis found what the authors describe as "a small to medium summary effect of gamified interventions on PA behavior" (Hedges g = 0.42, 95% CI 0.14-0.69). The effect was larger against inactive control groups such as waiting lists (g = 0.58, 95% CI 0.08-1.07) than against active control groups built around a non-gamified physical-activity program (g = 0.23, 95% CI 0.05-0.41). The authors found no statistical difference in the pooled effect between adults and adolescents, or between healthy participants and participants with a chronic disease. Measuring follow-up periods that averaged 14 weeks after each intervention ended, the same analysis found a weaker but still measurable effect (g = 0.15, 95% CI 0.07-0.23), which the authors interpret as evidence the initial effect is "not just a novelty effect caused by the playful nature of gamification…" They also call for "future rigorous trials" to confirm the finding.
Mitesh Patel and colleagues ran the Behavioral Economics Framingham Incentive Trial (BE FIT), a randomized clinical trial among 200 adults from 94 families enrolled in the Framingham Heart Study, published in 2017 in JAMA Internal Medicine. Over a 12-week intervention, participants in the gamification arm, who could earn points and progress through levels for hitting step goals, achieved their step goal on a greater proportion of days than controls (0.53 vs. 0.32; adjusted difference 0.27, 95% CI 0.20-0.33, p < .001) and increased mean daily steps over baseline more than controls (adjusted difference +953 steps/day, 95% CI 505-1401, p < .001). Across the 12-week follow-up period after the game ended, the gap narrowed: the daily-step advantage fell to +494 steps/day (95% CI 170-818, p < .01), roughly half its peak, though it remained statistically significant.
A separate trial by Patel and colleagues, the STEP UP randomized clinical trial, tested three social-incentive designs against a control arm among 602 overweight or obese adults (BMI 25 or higher) across 40 U.S. states, published in 2019 in the same journal, JAMA Internal Medicine. Over a 24-week intervention, all three gamified designs increased daily steps significantly more than control: competition by 920 steps/day (95% CI 513-1328, p < .001), support by 689 steps/day (95% CI 267-977, p < .001), and collaboration by 637 steps/day (95% CI 258-1017, p = .001). During the 12-week follow-up, with no gamification running, only the competition arm's advantage stayed statistically significant against control (+569 steps/day, 95% CI 142-996, p = .009); the study's own conclusion states the support arm (+428 steps/day, p = .04) and collaboration arm (+126 steps/day, p = .49) were not.
A 2021 systematic review by A. J. Davis and colleagues in the Journal of Clinical Nursing covered 7 studies and 657 patients at high cardiovascular-disease risk (diagnosed with heart disease, hypertension, stroke or type 2 diabetes), evaluating apps that used two or more game tactics for secondary prevention. It found gamified apps produced more improvement in physical activity, HbA1c and diabetes self-management empowerment than various comparators, and more physical-activity motivation than a neutral-content control app; it also found no added benefit above usual care for blood pressure, body mass index, heart-failure self-management, medication adherence or atrial-fibrillation knowledge. On engagement over time, the review reports that "app acceptability in terms of usage declined with time," even though ratings stayed high for specific game components such as challenges, leaderboards and badges; it also found "enjoyment was highest for elements that featured surprise/novelty…", the same trait that fades once a feature stops being new.
What this evidence doesn't show
None of the four sources above tested a running-specific app, a tiered rank system, or effects past 12 to 14 weeks of follow-up; the longest post-intervention window measured anywhere in this set is Patel et al.'s 12 weeks and Mazeas et al.'s follow-up average of 14 weeks. Davis et al.'s review found no added benefit from gamified apps, above usual care, for blood pressure, body mass index, heart-failure self-management, medication adherence or atrial-fibrillation knowledge, alongside the physical-activity gains it did find. Mazeas et al. found no statistical difference in the pooled effect between adults and adolescents, or between healthy participants and those with chronic disease, but that finding describes only the trials included in their review; it does not extend to every population or every gamified design.
How to apply this
The trials above studied wearable step trackers and points systems; none tested a specific running app, so the honest takeaway is about which design principles tend to hold up. A social-comparison design, the competition arm in Patel et al.'s 2019 trial, held up longer after the game ended than the support or collaboration arms in the same trial. Badges and leaderboard-style elements rated well for enjoyment in Davis et al.'s review, alongside "surprise/novelty…", which the same review flagged as fading with time. None of the four sources measured what happens past 12 to 14 weeks, so a gamified feature is best treated as reliable for getting a routine started, with a separate reason, a race goal, a training block, a habit that has become its own reward, needed to keep it going past that window.
Compare this week against your own last one, not against a study's average participant: in Runked the rank starts everyone at Bronze and only moves on your own runs. See how the running rank works.
Common mistakes
- Treating a "small to medium" pooled effect as proof any single gamified feature works. Mazeas et al.'s own description of their headline number, Hedges g = 0.42, is "small to medium"; it is a summary across 16 differently designed trials, and it does not guarantee any single feature works in any one app.
- Assuming a drop-off after a trial ends means gamification failed. In Patel et al.'s 2017 BE FIT trial, the daily-step advantage over control fell from 953 to 494 steps after the 12-week game ended, but it stayed statistically significant (95% CI, 170-818; p < .01); a smaller effect is not the same as no effect.
- Assuming every gamification design holds up the same way once the game stops. In Patel et al.'s 2019 STEP UP trial, only the competition arm remained significantly ahead of control during the 12-week follow-up; support and collaboration did not, per the study's own conclusion.
- Generalizing one trial's population to everyone. Davis et al.'s review covered patients at high cardiovascular-disease risk; Patel et al.'s STEP UP trial enrolled overweight and obese adults; neither is evidence about a general fitness-app user base.
What to do next
For how points, streaks, leaderboards and rank systems differ across specific running apps, see gamified running apps compared. For a plainer walkthrough of what a rank-based progression system is built to do, see gamified running, explained. For the habit-formation side of keeping a routine going once a game's novelty fades, see how to make running a habit.
Find out where you stand
Runked is free on the App Store. Connect your watch or track 3 runs, and meet your Runk.
Download on App StoreFrequently asked questions
Does gamification actually increase physical activity?
In pooled terms, yes. Mazeas and colleagues' 2022 meta-analysis of 16 randomized trials (2,407 participants), published in the Journal of Medical Internet Research, found a "small to medium" pooled effect of gamified interventions on physical activity (Hedges g = 0.42, 95% CI 0.14-0.69). The effect was larger against inactive comparators such as waiting lists (g = 0.58, 95% CI 0.08-1.07) than against active, non-gamified physical-activity programs (g = 0.23, 95% CI 0.05-0.41), and the authors found no statistical difference in the pooled effect between adults and adolescents, or between healthy participants and those with a chronic disease.
Does the effect disappear once the game ends?
It shrinks, without disappearing entirely, in the trials that measured this directly. Mazeas et al.'s pooled follow-up effect, averaging 14 weeks after an intervention ended, fell from Hedges g = 0.42 to g = 0.15 (95% CI 0.07-0.23). In Patel et al.'s 2017 BE FIT trial, the daily-step advantage over control fell from 953 to 494 steps across a 12-week follow-up (95% CI, 170-818; p < .01). Davis et al.'s 2021 review of gamified apps for cardiovascular-risk patients reported that app acceptability "declined with time," even as ratings for specific features such as challenges and leaderboards stayed high.
Which type of gamification held up best after the game stopped?
Social comparison, in the one trial that compared designs head to head. Patel et al.'s 2019 STEP UP trial tested three social-incentive designs, support, collaboration and competition, against a control group that had only a step tracker, across a 24-week intervention and 12-week follow-up. During the 12-week follow-up, only the competition arm stayed significantly ahead of control (+569 steps/day; 95% CI, 142-996; p = .009); the support arm's advantage (+428 steps/day; p = .04) and the collaboration arm's (+126 steps/day; p = .49) were not, per the trial's own conclusion.
Has Runked's own design been studied in a trial like these?
No. None of the trials or reviews above tested Runked or any specific running app. Runked's free tier includes a Bronze-to-Olympian rank, weekly effort-scored leagues and streaks, the same broad category of points-and-levels mechanic studied above, but that specific design has not itself been the subject of a published trial.