How to track running progress without racing
Racing gives a clean number, but a training block without a race on the calendar still needs its own check. Five checks fill that gap: a short maximal time trial, a maximal-effort run on a fixed benchmark route, heart rate at a familiar pace, a rolling training-load trend, and perceived effort at a familiar pace. A four-minute time trial reproduced a runner's pace more consistently than a lab step test or its post-test verification stage, coefficient of variation 1.8% against 4.5% and 9.7%, while VO2max readings were similarly reproducible across all three protocols, in a 2017 reliability study of ten recreational runners by Kerry McGawley of Mid Sweden University. The table below lines up what each check measures, how often to run it, and what its source states.
Five ways to check progress between races
| Method | What it measures | How often | What the source says |
|---|---|---|---|
| Time trial | Current sustainable pace and an estimated VO2max, from a short maximal effort | Repeat the identical protocol; a fixed test matters more than a fixed calendar date | Kerry McGawley, 2017, Frontiers in Physiology: a 4-minute time trial's performance measure had a coefficient of variation of 1.8% (ICC 0.95), versus 4.5% (0.83) for a graded step test and 9.7% (0.59) for its verification stage, in 10 recreational runners tested 5 times each. |
| Benchmark route | Finish time on one course whose distance never changes | Every other week inside a 1- to 2-month plan; roughly monthly inside a longer one | Marnie Kunz, USATF- and RRCA-certified coach, Runstreet: a fixed distance is what makes results comparable run to run; test again 1–2 weeks before race day. |
| HR at a fixed pace | Aerobic decoupling: how much heart rate climbs relative to pace across one steady session | Any steady aerobic run of real length, read as a trend across several weeks | Joe Friel, TrainingPeaks: decoupling under 5% signals strong aerobic endurance at that intensity; 5–10% signals moderate limitation; over 10% signals the effort likely sat above aerobic threshold. |
| Training-load trend | Whether accumulated training is building capacity without excess strain, as a chronic-versus-acute trend | Reviewed weekly at most; the chronic side is built to move slowly | Calvert, Banister, Savage and Bach, 1976, IEEE Transactions on Systems, Man, and Cybernetics: training modeled as daily impulses feeding separate, slower-building fitness and faster-fading fatigue terms. |
| RPE at a fixed pace | Whether the same felt effort now produces a faster pace than it used to | Compare against the same route about four weeks back | Jeff Gaudette, RunnersConnect: a pace drift of 15–30 seconds per mile at the same perceived effort over four weeks signals aerobic development. |
This table is Runked's own, built by lining up the five sources above by method; none of them published a comparison table like this on its own.
What the evidence says
Kerry McGawley, of the Swedish Winter Sports Research Centre at Mid Sweden University, put ten recreational runners (five men, five women, mean age 32) through five repeats each of a graded treadmill step test, a post-test verification stage, and a self-paced four-minute running time trial, publishing the comparison in Frontiers in Physiology in 2017. VO2max readings were similarly reproducible across all three protocols, coefficient of variation 1.8–2.2%, ICC 0.97 for each. Performance reproducibility split further apart: the time trial's coefficient of variation was 1.8% (ICC 0.95), against 4.5% (ICC 0.83) for the step test and 9.7% (ICC 0.59) for the verification stage. The time trial also produced VO2max values that ran lower than the step test's, by about 1.6 mL/kg/min on average with individual variation of plus or minus 3.6, a statistically significant gap (P = 0.008) McGawley attributes to differences between a fixed treadmill gradient and a runner's own pacing. McGawley is explicit that this does not make the time trial a substitute for a lab-grade VO2max reading: she states that if obtaining a true VO2max estimate is the main concern, the time-trial protocol used in her study would not be recommended for that purpose.
Marnie Kunz, a NASM-certified trainer and a USATF- and RRCA-certified running coach who founded Runstreet, describes a benchmark run as a timed run at maximum effort over a distance that stays fixed every time, since holding the distance constant is what makes one result comparable to the next. She recommends running one at the start of a training plan, then repeating it every other week inside a one- to two-month plan or roughly monthly inside a longer one, with a final check one to two weeks before race day depending on race distance.
Joe Friel, author of The Triathlete's Training Bible and a coach writing for TrainingPeaks, describes aerobic decoupling: the rise in heart rate at an unchanged pace or power across a single steady session, driven by heat, dehydration and accumulating fatigue. He puts decoupling under 5% at a given intensity as a sign of strong aerobic endurance at that intensity, 5–10% as a sign of moderate limitation or fatigue, and over 10% as a sign the effort likely sat above aerobic threshold, or that endurance at that pace is not yet there. The article cites research by Wingo and Cureton on cardiovascular drift and by Vautier and colleagues on predicting exhaustion from heart-rate drift, without restating their figures here. Friel frames decoupling as a trend to track across weeks: as aerobic fitness improves, decoupling at the same pace should shrink.
Thomas Calvert, Eric Banister, Margaret Savage and Timothy Bach modeled training and performance as a systems problem in a 1976 paper in IEEE Transactions on Systems, Man, and Cybernetics, treating a runner's daily training as a series of impulses feeding two components: a slower-building fitness term and a faster-building, faster-fading fatigue term. The model's central proposal is to compare the two trends together instead of reading either number by itself; the paper tests the approach against one swimmer's 100-meter performances over a season. The chronic-versus-acute framing still used across endurance coaching traces back to this model.
Jeff Gaudette, who holds a master's degree from Johns Hopkins University and co-founded RunnersConnect, recommends comparing an easy run on the same route at the same perceived effort about four weeks apart: “A drift of 15 to 30 seconds per mile (9 to 18 seconds per kilometer) over four weeks at the same effort means your aerobic system has developed.”
A separate study complicates how much weight a felt sense of progress alone should carry, on a different timescale. Eneko Larumbe-Zabala and colleagues followed sixteen recreational marathoners across five check-ins over sixteen weeks, publishing in Frontiers in Psychology in 2020. They modeled the runners' measured aerobic and anaerobic threshold running speeds as predictors of each runner's own later-rated perceived fitness. Across the full sixteen weeks, both threshold speeds were significantly associated with perceived fitness over time (aerobic threshold: p = 0.004, R2 = 0.14; anaerobic threshold: p = 0.001, R2 = 0.14). When the authors instead modeled the percent change between one check-in and the next, none of the physiological variables, including both threshold speeds, showed a statistically significant predictive relationship with the matching change in perceived fitness; the authors describe those percent-change effect sizes as trivial to small.
How to apply this to your own training
Pick one test format and hold it steady: a four-minute maximal effort on a track or treadmill, a fixed-distance outdoor course, or a familiar 20 to 40 minute route at an easy, conversational effort. Any of these works, as long as the distance, the surface and the effort target stay the same every time. Log the date, the conditions (heat, wind, hills) and the result together; a fast time on a cool, flat day and a slower one into a headwind are not the same data point.
Space repeats to match the method. A maximal time trial or benchmark run fits roughly every two weeks to a month, per Kunz's cadence above. Heart-rate decoupling and perceived effort at a fixed pace read better as a trend across several easy runs a week or two apart than as a single test. A training-load trend needs a run history of several weeks before its slower-moving side means anything, following the chronic-versus-acute logic above.
Compare this month's result to your own number from a month or two back, not to another runner's time on the same course; in Runked the rank starts everyone at Bronze and only moves on your own runs. See how the running rank works.
Common mistakes
- Changing the course, the distance or the effort level between checks. Kunz's benchmark-run method holds the distance fixed specifically so results compare across sessions; McGawley's reliability numbers above describe one fixed protocol, repeated exactly, without swapping in a new course each time.
- Reading a single felt-effort check as the whole picture. Larumbe-Zabala and colleagues' finding above, that check-to-check changes in measured threshold speed did not reliably predict the matching change in perceived fitness, even though the raw threshold speeds and perceived fitness tracked each other significantly across the full sixteen weeks, is a reason to pair a felt sense of progress with at least one measured number.
- Checking a slow-moving trend every day and reacting to the noise. Calvert and Banister's model treats fitness as a trend built over weeks of training; a single day's number, good or bad, is not that trend.
What to do next
Two of these methods have a fuller page of their own. Training load, fitness and fatigue goes deeper on turning a week of runs into a chronic-versus-acute trend line. What is RPE in running? covers the 0-to-10 effort scale behind the felt-effort check above. If a time trial or benchmark route needs a target beyond "faster than last time," age-graded running times compares a result against runners of other ages, and the race time predictor turns training data into a projected finish time.
Find out where you stand
Runked is free on the App Store. Connect your watch or track 3 runs, and meet your Runk.
Download on App StoreFrequently asked questions
How often should progress checks happen if there's no race on the calendar?
It depends on the check. Marnie Kunz, a USATF- and RRCA-certified coach and founder of Runstreet, recommends a benchmark run at the start of a plan, then every other week inside a one- to two-month plan or roughly monthly inside a longer one, with a final check one to two weeks before race day. Joe Friel, writing for TrainingPeaks, treats heart-rate decoupling at a fixed pace as a trend, something to read across several weeks of steady runs, since any single run can be thrown off by heat, wind or a bad night's sleep.
Which single check gives the most reliable number?
Of the methods here, a short maximal time trial has the clearest reliability data behind it. Kerry McGawley, of the Swedish Winter Sports Research Centre at Mid Sweden University, tested ten recreational runners five times each on a graded treadmill step test, its post-test verification stage, and a self-paced four-minute time trial, publishing the comparison in Frontiers in Physiology in 2017. The time trial's performance measure had a coefficient of variation of 1.8% and an ICC of 0.95, tighter than the step test's 4.5% and 0.83 or the verification stage's 9.7% and 0.59. VO2max readings were similarly reproducible across all three protocols, at a coefficient of variation of roughly 2% each.
Does feeling fitter always match the numbers?
It depends on the timescale, per one study. Eneko Larumbe-Zabala and colleagues followed sixteen recreational marathoners across five check-ins over sixteen weeks, publishing in Frontiers in Psychology in 2020. Across the full sixteen weeks, runners' measured aerobic and anaerobic threshold running speeds were significantly associated with their own later-rated perceived fitness (aerobic threshold: p = 0.004; anaerobic threshold: p = 0.001). Check-to-check, though, the percent change in either threshold speed did not significantly predict the matching percent change in perceived fitness; the authors describe those effect sizes as trivial to small. That gap between the two timescales is one reason to pair a felt-effort check with at least one measured number.
Does Runked replace racing as a way to track progress?
Not by itself, and it is not built to. PRO's race-time predictions, a projected finish time recalculated from training data, are the one Runked feature aimed at the same question in race-clock terms: roughly how a runner would finish if a race were on the calendar right now.