Whoop built a company on a single number. Every morning its band hands you a recovery score, green, yellow or red, before your feet touch the floor, and tells you how hard you're allowed to train. Oura put the same promise in a ring and called it readiness; Garmin, Coros and Apple have all since shipped their own. The pitch is compelling: sleep on it, and your body wakes you with a verdict.
This is not the first time we've watched a real piece of physiology get flattened into a number. In our last piece we followed VO2 Max from a 1923 lab bench into the middle of running culture, and we left one thread deliberately hanging: these readiness scores, we said, deserve their own scrutiny.
Look underneath the hood and you find the same measurement underneath: heart rate variability. Cardiologists have used it for decades, as a survival predictor after heart attacks in the 1980s and with formal measurement standards by the 1990s. Only later did it migrate from the clinic to the wrist, repackaged as a same-morning verdict on whether to train. HRV is real, measurable, peer-reviewed physiology. And it has become a useful, yet overloaded and gamified metric in many wearables.
What HRV Actually Measures
Your heart beats almost like a metronome, but not quite. Even at rest, the gap between one beat and the next stretches and shrinks by a few milliseconds, and that jitter is heart rate variability. It isn't measurement noise, it's the fingerprint of your autonomic nervous system, the involuntary layer that runs your heart and your stress response without asking you.
The interval between beats is never identical, even at rest. That beat-to-beat variation, measured in milliseconds, is HRV.
Two forces pull on that system. The parasympathetic branch, carried by the vagus nerve, is "rest and digest": when it dominates, heart rate slows and the variation widens. The sympathetic "fight or flight" branch takes over under stress, illness, poor sleep or fatigue, and the variation narrows. HRV is one of the few non-invasive windows into that system, but a one-sided one. The metric worth knowing, rMSSD (log-transformed to lnRMSSD to tame its skew), is the most reliable day-to-day index, and it is specifically a measure of parasympathetic (vagal) activity. No simple HRV metric cleanly isolates the sympathetic side. The frequency-domain LF/HF ratio was once sold as a "sympathovagal balance," but that interpretation has not held up. So HRV reads your "rest and digest" branch well and your "fight or flight" branch only by inference: real physiology, with a real blind spot.
HRV as Daily Readiness
The Finnish physiologist Antti Kiviniemi and colleagues first tested the idea in 2007: measure HRV each morning, schedule a hard session only when it says the body is ready. Since then a handful of randomized trials have tested the idea, with results that are real but modest.
In 2016, Vesterinen and colleagues took 40 recreational runners through an eight-week block: one group followed a fixed plan, the other had its hard days timed by morning HRV. The HRV group improved its 3,000 m time significantly (+2.1%, versus a non-significant +1.1% for the fixed plan) while running roughly 25% fewer hard sessions (13 versus 18). HRV guidance didn't build a bigger engine, it reached the same or better result on less hard work by putting the quality sessions on the right days. The fixed-plan group's VO2 Max actually rose slightly more, which sounds like a contradiction until you separate the engine from the race: their extra hard sessions bought a little more raw aerobic ceiling, but it didn't translate into a faster 3,000 m. The HRV group raced quicker on less hard work, and the race is the outcome that counts.
Zoom out from any single trial and the picture gets muddier. Three meta-analyses have now pooled the studies, and they disagree, which is itself revealing. Granero-Gallegos and Plews (2020), across six trials and 195 athletes, found HRV-guided training gave a small but real VO2 Max gain (effect size 0.40), statistically larger than a predefined plan but with enormous heterogeneity. Manresa-Rocamora and Flatt (2021), across eight studies, found the opposite: no significant performance advantage. The one thing HRV guidance reliably improved there was the athletes' own vagal HRV, a bit like a diet that improves your cholesterol reading without changing your weight.
So how much evidence is there? Enough to take seriously, not to worship: six to eight small trials, mostly cyclists and mixed endurance athletes rather than runners, high variability, quality the reviewers rate "unclear." The effect is real but small, and what HRV guidance actually does is not mysterious: it nudges you to train hard when you're fresh and ease off when you're not.
That is simply auto-regulation, and it isn't unique to a wristband. A good coach, or an athlete honestly reading their own legs, sleep and motivation, tends to reach much the same decisions without any device. HRV can make that judgment a little more consistent, but it is one input among several, not a new source of fitness.
The Paradox That Breaks the Daily Readiness Score
A readiness score rests on one simple assumption: high HRV means recovered, low HRV means tired. The problem is that in trained endurance athletes, the very people most likely to be wearing the device, that assumption breaks down.
The research has a name for it: parasympathetic saturation, mapped out by Plews, Stanley and Buchheit in a 2013 review. Their finding is awkward for any daily score. In fit athletes, HRV going up and HRV going down have both been linked to trouble, and getting fitter can actually push HRV lower. High-good, low-bad is simply the wrong rule at the top end.
The mechanism is counterintuitive. You'd expect heavy training to tip toward "fight or flight" and pull HRV down, and acutely it can. But under sustained overload the body often overcorrects into a parasympathetic overdrive that pushes resting HRV back up. And well-trained athletes sit near the ceiling of vagal activity, where HRV stops tracking it in a straight line, the "saturation" the researchers named. A high or rising number can mean deep recovery or the early cost of digging too far, and the reading alone can't tell which.
The same change in HRV can point two opposite ways. Direction alone can't tell adaptation from breakdown.
It gets stranger. When Le Meur and colleagues (2013) overloaded 21 trained triathletes into functional overreaching, the fatigued athletes' resting HRV went up, not down, and only came back down when they tapered. Run it the other way and you get the mirror image: when Nuuttila and colleagues (2024) pushed 24 recreational runners through a hard overload block, their overnight HRV also rose, but this time it reflected positive adaptation and their 3,000 m times improved. Same signal, same direction, opposite meanings. A 2016 meta-analysis by Bellenger tied the bow: post-exercise HRV rises in both good adaptation and bad overreaching, while resting HRV is often barely moved by overreaching at all.
For the number on your wrist, that's damning. No single HRV value, and often not even its direction, cleanly separates "adapting well" from "breaking down." A daily readiness percentage off one morning's reading asserts a certainty the physiology does not support.
One Reading Is Mostly Noise
Even setting the paradox aside, a single morning HRV reading is noisy. When Bellenger's group tracked elite athletes for 16 weeks (2022), the week-to-week wobble in their HRV averaged about 5%, independent of how they'd trained. That's the noise floor: any single day can land 5% above or below where your body really is for reasons unrelated to fitness or fatigue.
The scattered dots are raw daily readings, bouncing about 5% for no meaningful reason; one low morning is just a rough night. The blue line is the seven-day rolling average against your own normal-range band, and the signal worth acting on is that average drifting out of the band, not any single dot.
This is why every HRV researcher says the same thing: don't read the day, read the trend. Le Meur's overreaching signal was invisible in any single measurement and showed up only in the weekly average. The rule that falls out is a seven-day rolling average against your own baseline, not yesterday's number against the day before. Plews and Buchheit have even proposed watching the variability of the variability as an early warning of overreaching, though that still rests on a single two-athlete case study and should be treated as promising, not proven.
Then there are the confounders, which are brutal. Morning HRV is pushed around by poor sleep, alcohol, caffeine, a coming illness, menstrual-cycle phase, heat, altitude, even your breathing rate and body position. Many move the number as much as training does. A single red morning is as likely to mean "two glasses of wine and bad sleep" as "overtrained." That isn't the number failing: as Marco Altini, who has validated these measurements for over a decade, notes, a drop after alcohol or lost sleep is HRV doing its job, detecting real stress. What it can't do is name the source.
So How Should You Actually Use It
None of this makes HRV useless. It makes it a trend instrument, not a daily oracle. Used that way, a simple protocol:
- Measure consistently or not at all. Same time, position, tool. Morning and overnight both work with no real accuracy edge, as long as the condition stays consistent. What matters is the tool: a chest strap, a validated fingertip app, or a ring capturing the full night are all reliable; a sporadic daytime wrist reading is the weak link.
- Ignore the single day. Read three layers. The daily value, a seven-day baseline, and a normal-range band from a month or two of your own data (the framework Altini's tools popularized). A low morning inside the band is noise; when the baseline or daily values fall below it, that's the threshold worth acting on, especially if your legs agree.
- Time hard sessions, don't police rest. The best-supported use is Vesterinen's: load when the trend is stable or rising, hold the quality work when it's suppressed. A nudge, not a command.
- Never let it overrule your body. If it says "recovered" but you feel wrecked, or "rest" but you feel strong and your training's been sound, trust the athlete. HRV is one input, and how the run feels is still the one sports science leans on hardest.
The Verdict: A Good Signal, a Bad Oracle
Is HRV a good metric? Yes and no, and the distinction is the point. As a smoothed, individualized trend, it's a legitimate window into how your nervous system is coping with the sum of your training and your life, and timing hard days by it has real if modest support.
As the thing your watch sells you, a precise daily score that greenlights or benches your workout off one morning's reading, it asks a signal that's useful as a trend to be a same-day verdict, then folds in sleep and activity data that only muddy it. The paradox and the noise are real, but the sharper problem is the packaging: the same move as VO2 Max, one watch generation later, a real measurement gamified into false certainty. Read the trend, stay honest about how the run feels, keep the athlete, not the algorithm, in charge. The number is an instrument, not the verdict.
FAQ
Should I skip a workout because my HRV is low this morning?
Not on one reading. A single low morning is within the normal 5% day-to-day noise and is often explained by sleep, alcohol, or a late meal rather than fatigue. If your seven-day rolling average has been trending down for a week and your legs agree, that's a reason to ease off. One red morning on its own isn't.
Is Whoop, Oura, or Garmin HRV accurate?
For the raw measurement, mostly yes, especially the overnight readings those devices favor, with rings and chest straps ahead of wrist optical. The accuracy question was largely solved. The harder question is what the readiness score built on top of it claims to know, which is more than the physiology supports.
References
- Kiviniemi, A. M., Hautala, A. J., Kinnunen, H., & Tulppo, M. P. (2007). Endurance training guided individually by daily heart rate variability measurements. European Journal of Applied Physiology, 101(6), 743–751. doi.org/10.1007/s00421-007-0552-2 — the original HRV-guided training study.
- Vesterinen, V., Nummela, A., et al. (2016). Individual endurance training prescription with heart rate variability. Medicine & Science in Sports & Exercise, 48(7), 1347–1354. pubmed.ncbi.nlm.nih.gov/26909534 — 40 recreational runners; the HRV-guided group improved 3,000 m time while doing roughly 25% fewer hard sessions.
- Granero-Gallegos, A., Plews, D., et al. (2020). Effectiveness of training prescription guided by heart rate variability versus predefined training: a systematic review and meta-analysis. pmc.ncbi.nlm.nih.gov/articles/PMC7663087 — 6 RCTs, 195 athletes; HRV-guided VO2 Max gain small but significantly greater than predefined training (ES 0.40), very high heterogeneity (I²≈94%).
- Manresa-Rocamora, A., Flatt, A. A., et al. (2021). Heart rate variability-guided training for enhancing cardiac-vagal modulation, aerobic fitness, and endurance performance: a systematic review with meta-analysis. pmc.ncbi.nlm.nih.gov/articles/PMC8507742 — 8 studies; no significant between-group performance advantage; the only robust win was vagal HRV itself (SMD 0.50).
- Plews, D. J., Laursen, P. B., Stanley, J., Kilding, A. E., & Buchheit, M. (2013). Training adaptation and heart rate variability in elite endurance athletes: opening the door to effective monitoring. Sports Medicine, 43(9), 773–781. pubmed.ncbi.nlm.nih.gov/23852425 — the foundational review of parasympathetic saturation and why single indices mislead.
- Le Meur, Y., et al. (2013). Evidence of parasympathetic hyperactivity in functionally overreached athletes. Medicine & Science in Sports & Exercise, 45(11), 2061–2071. pubmed.ncbi.nlm.nih.gov/24136138 — overreached triathletes' resting HRV rose, not fell, and normalized on taper.
- Nuuttila, O.-P., et al. (2024). Nocturnal heart rate variability and training responses in recreational runners. pmc.ncbi.nlm.nih.gov/articles/PMC11541970 — overnight HRV rose during a hard overload block, reflecting positive adaptation.
- Bellenger, C. R., et al. (2016). Monitoring athletic training status through autonomic heart rate regulation: a systematic review and meta-analysis. Sports Medicine, 46(10), 1461–1486. doi.org/10.1007/s40279-016-0484-2 — post-exercise HRV rises in both adaptation and overreaching; resting HRV is largely unaffected by overreaching.
- Plews, D. J., Laursen, P. B., Kilding, A. E., & Buchheit, M. (2012). Heart rate variability in elite triathletes, is variation in variability the key to effective training? A case comparison. European Journal of Applied Physiology, 112(11), 3729–3741. pubmed.ncbi.nlm.nih.gov/22367011 — origin of the "variation in variability" idea; a two-athlete case study.
- Bellenger, C. R., et al. (2022). Evaluating the typical day-to-day variability of WHOOP-derived heart rate variability in Olympic water polo athletes. Sensors, 22(18), 6723. pmc.ncbi.nlm.nih.gov/articles/PMC9505647 — quantifies the roughly 5% weekly coefficient of variation (5.4%) that is HRV's noise floor.
- Plews, D. J., Scott, B., Altini, M., Wood, M., Kilding, A. E., & Laursen, P. B. (2017). Comparison of heart rate variability recording with smartphone photoplethysmography, Polar H7 chest strap, and electrocardiography. International Journal of Sports Physiology and Performance, 12(10), 1324–1328. pubmed.ncbi.nlm.nih.gov/28290720 — smartphone PPG and chest strap both agree near-perfectly with ECG for rMSSD at rest.
- Altini, M. On heart rate variability and readiness. Medium. medium.com/@altini_marco/on-heart-rate-variability-hrv-and-readiness-394a499ed05b — argues a readiness score's weakness is bundling HRV with behavioral data, which double-counts stress and dilutes the signal; raw HRV against your own baseline is the trustworthy part, and a drop after alcohol or poor sleep is the signal working, not failing.
- Altini, M. The ultimate guide to heart rate variability (HRV), part 2. Medium. medium.com/@altini_marco/the-ultimate-guide-to-heart-rate-variability-hrv-part-2-323a38213fbc — the three-layer interpretation used in practice: daily value, seven-day baseline, and a normal-range band from one to two months of your own data; act when the baseline or daily values fall below the band.
- Altini, M. Morning or night for heart rate variability (HRV) measurement? marcoaltini.com. marcoaltini.com/blog/morning-or-night-for-heart-rate-variability-hrv-measurement — morning and overnight measurement carry no clear accuracy advantage over each other; consistency of condition and full-night capture matter more than the timing.
- Kleiger, R. E., Miller, J. P., Bigger, J. T., & Moss, A. J. (1987). Decreased heart rate variability and its association with increased mortality after acute myocardial infarction. The American Journal of Cardiology, 59(4), 256–262. pubmed.ncbi.nlm.nih.gov/3812275 — the classic finding that reduced HRV predicts survival after a heart attack; HRV's clinical roots long predate wearables.
- Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology (1996). Heart rate variability: standards of measurement, physiological interpretation and clinical use. Circulation, 93(5), 1043–1065. pubmed.ncbi.nlm.nih.gov/8598068 — the formal measurement standards that codified HRV for clinical use.
- Billman, G. E. (2013). The LF/HF ratio does not accurately measure cardiac sympatho-vagal balance. Frontiers in Physiology, 4, 26. pubmed.ncbi.nlm.nih.gov/23431279 — why the LF/HF ratio can't be read as sympathovagal balance; rMSSD remains a valid vagal index.