Not All Step Counters Are Equal — Here's Which Wearable Algorithm Actually Works
A rigorous free-living validation study reveals dramatic accuracy gaps among six popular step-counting algorithms, with one wrist algorithm matching gold-standard thigh devices.
Summary
Researchers at NCI and Cal Poly filmed 20 adults during 35 real-world observation sessions and compared six device-based step-counting algorithms against video-coded ground truth. The thigh-worn activPAL and the wrist-worn ActiGraph running the open-source 'stepcount' algorithm were the most accurate, both statistically equivalent to direct observation at a 15% error threshold. Other wrist algorithms ranged from mediocre to disastrously inaccurate — the SDT algorithm showed a mean absolute percent error of 231%, making it nearly useless. Accuracy was highest during straight walking or running, but fell sharply for activities like biking, stroller pushing, and food prep. These findings matter because step counts increasingly appear in longevity research and public health guidelines, and choosing the wrong algorithm could fundamentally distort health associations.
Detailed Summary
Daily step counts have become one of the most widely used proxies for physical activity in epidemiologic research and consumer health tracking, with multiple large cohort studies linking higher step counts to reduced risk of cardiovascular disease, metabolic disorders, cancer, and all-cause mortality. Yet the accuracy of the algorithms that convert raw wrist accelerometer signals into step counts has rarely been rigorously tested against a true gold standard in naturalistic, free-living conditions. This study directly addresses that gap.
The research team recruited 20 adults (mean age 36.1 ± 14.7 years; 50% female) from San Luis Obispo, CA, between July 2019 and March 2020. Each participant wore a thigh-worn activPAL3 and a wrist-worn ActiGraph GT3X+ simultaneously for seven days. A subset completed two separate three-hour sessions in which they were continuously filmed with a GoPro Hero 5 camera. Five trained coders annotated the footage for posture, activity type, intensity, and step counts — with inter-rater reliability exceeding ICC 0.9 in all categories. This yielded 35 matched observation sessions across 20 participants.
Six step-counting methods were evaluated: the activPAL's proprietary VANE/CREA algorithm, plus five methods applied to ActiGraph data — the proprietary ActiLife, and four open-source algorithms (Oak, SDT, Verisense, and stepcount). Among all wrist-based methods, the OxWearables 'stepcount' algorithm — a hybrid machine-learning plus peak-detection model trained on annotated free-living data — performed dramatically better than its competitors. Its mean absolute percent error (MAPE) was just 17.1% compared to direct observation, versus 40.5% for ActiLife, 41.7% for Verisense, 62.4% for Oak, and a staggering 231.5% for SDT. Both activPAL and stepcount were statistically equivalent to direct observation at the 15% equivalence threshold (two one-sided test, p < 0.05), while no other wrist algorithm achieved even the 15% level.
Variance explained (marginal R²) ranged from 0.64 (SDT) to 0.90 (Oak) at the session level, though high R² did not necessarily indicate low bias — a critical methodological distinction. Over the full seven-day wear period, activPAL and stepcount produced highly concordant estimates (R² = 0.87, MAPE = 12.8%), suggesting these two approaches are interchangeable for population-level step count research and could enable pooling of data across studies that used different device placements.
Activity-type stratification revealed a consistent pattern: all algorithms performed best during sustained walking or running (the intended use case), but accuracy degraded substantially — and variably — during biking, modified walking (e.g., carrying loads, ascending stairs), and mixed movements such as pushing a stroller, food preparation, or computer work. The activPAL's VANE algorithm notably misclassified cycling as stepping at the epoch level, though its daily summary software corrected this. SDT's catastrophic overcount appears driven by false positives during non-locomotor upper-limb movements, a known vulnerability of threshold-based wrist algorithms.
For the longevity research community, the implications are significant. Studies that have used SDT, Oak, or ActiLife to quantify daily steps may have substantially mismeasured exposure, potentially attenuating or distorting dose-response relationships between steps and health outcomes. Researchers designing new wearable studies or reanalyzing existing datasets should strongly consider the stepcount algorithm or, where device placement allows, the activPAL as the preferred step-counting method.
Key Findings
- The 'stepcount' wrist algorithm had the lowest error among all wrist methods: MAPE of 17.1% vs direct observation, compared to 40.5% (ActiLife), 41.7% (Verisense), 62.4% (Oak), and 231.5% (SDT)
- Both activPAL (thigh) and stepcount (wrist) were statistically equivalent to direct observation at the 15% equivalence threshold (two one-sided tests, p < 0.05); no other wrist algorithm met even this threshold
- Over 7 days of free-living wear, activPAL and stepcount agreed closely with each other: R² = 0.87, MAPE = 12.8%
- Session-level R² ranged from 0.64 (SDT) to 0.90 (Oak), but high R² did not indicate low bias — Oak had 62.4% MAPE despite the highest R²
- Accuracy was highest during walking/running for all algorithms and lowest during biking, modified walking (stair climbing, load-carrying), and mixed movements like stroller-pushing or food prep
- The activPAL's epoch-level data misclassified cycling pedaling as steps (due to VANE algorithm behavior), though the daily summary CREA algorithm corrected this exclusion
- 20 adults completed 35 valid 3-hour video-recorded sessions; inter-rater reliability for step counting exceeded ICC 0.9 in 20% of dual-coded videos
Methodology
Twenty adults (50% female, mean age 36.1 ± 14.7 years) simultaneously wore a thigh-worn activPAL3 and wrist-worn ActiGraph GT3X+ for seven days; a subset completed two 3-hour GoPro-recorded free-living sessions yielding 35 observation periods. Five trained coders annotated video footage for posture, activity type, and step counts (ICC > 0.9), serving as the gold standard criterion. Six algorithms (activPAL VANE/CREA, ActiLife, Oak, SDT, Verisense, stepcount) were benchmarked using linear mixed-effects models, MAPE, RMSE, and two one-sided equivalence tests at 10% and 15% thresholds, with Bland-Altman plots for visual bias assessment.
Study Limitations
The sample was small (n=20) and relatively young (mean age 36), recruited from a single geographic area, limiting generalizability to older adults or those with gait abnormalities. Direct observation sessions were only three hours each and scheduled partly by participant preference, which may not fully capture the diversity of daily activities. The study did not include consumer wearables (e.g., Apple Watch, Fitbit) or ankle/hip placements, and no conflicts of interest are disclosed by the authors, who are affiliated with NCI and Cal Poly.
Enjoyed this summary?
Get the latest longevity research delivered to your inbox every week.
Enter your email to subscribe:
