← Measurement Science
Program templates and the program builder · In every tier

Periodization

Planned training phases are near-universal in serious lifting. The evidence that they beat equally hard unstructured training is much thinner than you’ve been told.

TL;DR

Periodization — organising training into planned phases — is in almost every serious strength program, and it isn’t going anywhere. But the claim that a periodized program beats an equally hard, equally high-volume unperiodized one is far weaker than the industry implies: small for maximal strength, absent for muscle growth, and never once tested against the comparator that would actually settle it. The tighter the study design, the smaller the effect gets. Between the models themselves, the pooled evidence leans slightly toward undulating over linear for maximal strength — two meta-analyses point that way, with confidence intervals that nearly touch zero — and shows nothing whatsoever for muscle growth. Progressive overload and volume are doing most of the work; tapering for a competition is real; everything past that is structure, preference, and whether you actually run the thing.

See it in the app

In every tier
  1. Open the Programs tab

  2. Tap Browse Templates to run an established program as written — or Build Your Own to design one from scratch

  3. Whatever you pick lands in My Programs and drives your workouts from there

  4. Nothing in the app ranks one program model above another — because, per this page, the evidence isn’t nearly strong enough to

Even this is in Bronze — the whole tracker is $1/mo, and nothing here is locked behind a higher tier.

Why it matters

Every program you could follow makes the same implicit promise: that the arrangement of the work is worth something on top of the work itself. So here is the question, stated the way a trial would have to state it — if two lifters do the same number of hard sets at the same intensities, and one arranges them into planned phases while the other doesn’t, does the planner end up stronger? That is the whole argument, and most program marketing walks straight past it.

The answer changes what you should spend your attention on. If structure is worth a lot, picking the right model matters and getting it wrong costs you. If structure is worth a little, then volume, progression, and adherence are the game, and the model is a logistics decision — which is a much less stressful way to train. Powerhaus ships program templates and a builder, so we owe you a straight answer about what choosing one actually buys.

How Powerhaus handles it

There is no periodization score in Powerhaus, and there isn’t going to be one. We don’t rank the templates by model, we don’t tell you daily-undulating will build strength faster than linear, and we don’t compute a number that grades how well-periodized your training is. The pooled evidence does lean slightly toward undulating for maximal strength — we say so below rather than pretending otherwise — but the lean is small, it’s strength-only, its confidence intervals nearly touch zero, and the best-controlled single trial goes the other way. That is nowhere near enough to steer your programming, and grading your training against it would be inventing precision we don’t have.

What we ship instead: the Programs tab carries templates you can run as written, plus a builder for designing your own. Model choice is presented as a fit decision — how your week is shaped, how much exercise variety you want, whether you’re training toward a date — not as an efficacy ranking.

The one periodization-adjacent number we do compute is deliberately retrospective, and it says so: Deload Detection flags weeks where your volume fell more than 30% below your rolling average. It recognises structure you already ran; it doesn’t prescribe structure you should run. Meanwhile Best Volume and the Muscle Heat Map track the variables the evidence genuinely supports — how much hard work you did, and where it went.

The science

The rigor ladder

The same claim — periodized beats non-periodized — measured under progressively tighter study designs. The pattern is not any single number. It is that as methodological control goes up, the effect goes down. Read the small print under each rung: they are not all measuring quite the same thing, and we have flagged where they differ rather than smoothing it over.

  1. Rhea & Alderman 2004k = 11 · no volume equating · every control arm a constant program · not all trials randomised · pools strength AND power outcomes together
  2. Williams et al. 2017 — as reportedk = 18, 81 effects, n = 612 · equating not required: 45 of 81 effects (55.6%) volume-equated, 36 (44.4%) not · 1RM only
  3. Williams et al. 2017 — outliers removedsame trials, minus the 14 effects (from 5 studies) falling outside the funnel plot’s 95% CI — mostly single-set controls against multi-set periodized arms
  4. Moesgaard et al. 2022 — strengthk = 35 RCTs, n ≈ 1,187 · sets × reps matched between arms as an inclusion requirement · 1RM only
  5. Moesgaard et al. 2022 — hypertrophysame volume-equated trials, muscle growth instead of strength

Bars are drawn proportional to the reported effect size. Two things this chart is not: the top and bottom rungs are not strictly like-for-like — Rhea & Alderman’s 0.84 pools strength *and* power outcomes, while Moesgaard’s 0.31 is strength only, and no strength-only figure from Rhea is retrievable. And the 0.23 rung is a sensitivity analysis with funnel-plot outliers removed, not a formal publication-bias correction — no trim-and-fill was run, and Williams also reports a fail-safe N of about 1,038, meaning that many null results would be needed to erase its effect. The direction of the ladder is real. Its precision is not.

The best-controlled answer we have. Moesgaard et al. (2022), in *Sports Medicine*, pooled 35 randomised trials (n ≈ 1,187, 6–36 weeks) and — critically — required that sets and reps be matched between the periodized and non-periodized arms as a condition of inclusion. For maximal strength it found a small but statistically real advantage: ES = 0.31 (95% CI 0.04–0.57, p = 0.02). For muscle growth it found nothing: ES = 0.13 (95% CI −0.10 to 0.36, p = 0.27). Grgic et al. (2017) arrives in the same place from a different direction — linear versus daily-undulating, volume equated, hypertrophy — at d = −0.02 (95% CI −0.25 to 0.21, p = 0.848).

Now look at what happens as the designs tighten. The number that made periodization famous is Rhea & Alderman’s (2004) ES = 0.84 — reported with a standard deviation of ±1.41, wider than the effect itself. It is also the most compromised: all eleven of its included studies compared a periodized program against a constant, unvarying one, and a later systematic review found that not all of them were even randomised. One more thing about that number that almost never travels with it — **0.84 pools strength *and* power outcomes together, so it isn’t a like-for-like comparison with the strength-only figures below, and no strength-only figure from that paper is publicly retrievable. [Williams et al. (2017)](https://link.springer.com/article/10.1007/s40279-017-0734-y) shows what inflates these numbers, and it is worth being precise about, because this is the figure most often mangled. Its reported effect was d = 0.43 (0.27–0.58). Its funnel plot and Egger’s test flagged asymmetry at p = 0.006, and a sensitivity analysis that stripped out the 14 effects (from five studies) falling outside the funnel plot’s 95% CI dropped the number to d = 0.23 (0.13–0.33) — roughly half. That is an outlier-removal analysis, not a formal bias correction: no trim-and-fill was run, and Williams also reports a fail-safe N of about 1,038**, meaning roughly a thousand unpublished null results would be needed to erase the effect entirely. Both facts belong in the same sentence. On volume: Williams didn’t *require* equating — 45 of its 81 effects (55.6%) were volume-equated and 36 (44.4%) weren’t — and the removed outliers sit squarely in that unequated 44.4%, where single-set non-periodized arms were compared against multi-set periodized ones. That portion measures doing more work, not organising work differently.

Undulating versus linear: a real but small lean, for strength only. Take the individual trials first. Five 12-week head-to-head comparisons, 20 to 42 subjects each, split three ways. Rhea (2002) favours daily-undulating with clear significance (bench +28.8% vs +14.4%; leg press +55.8% vs +25.7%). Simão (2012) favours nonlinear, significant on two lifts. Prestes (2009) and Miranda (2011) lean undulating but with no significant between-group difference. And Apel, Lacey & Kell (2011) — the one trial that equated volume and intensity — significantly favours linear. That is a set of trials noisy enough to produce whichever conclusion you go looking for.

Pooled, though, the needle does move — and we’re not going to hide it because it complicates the story. Moesgaard et al. (2022)’s comparison of linear against undulating across its volume-equated trials found an overall effect on 1RM favouring undulating: ES = 0.31 (95% CI 0.02–0.61, p = 0.04). Williams’ meta-regression independently found a periodization-model term pointing the same way (b = 0.51, p = 0.001). Two separate pooled analyses, same direction. Inside Moesgaard the effect is carried almost entirely by trained lifters (ES = 0.61, 95% CI 0.00–1.22, p = 0.05) with essentially nothing in untrained ones (ES = 0.06, 95% CI −0.20–0.31, p = 0.67). So the honest sentence is: for maximal strength, the evidence weakly favours undulating over linear. Weakly is doing real work there — both confidence intervals have a lower bound at or barely off zero, Harries et al. (2015) (k = 17, n = 510, volume equated) found nothing at all, and the single most tightly controlled trial in the set went the *other* way. It is a lean, not a verdict, and nowhere near strong enough for an app to tell you which model to run.

For muscle growth there is no lean at all, and that null is the clean one. Moesgaard’s linear-vs-undulating hypertrophy comparison lands at ES = 0.05 (95% CI −0.20–0.29, p = 0.72) — sitting alongside Grgic’s d = −0.02 above. Klemp et al. (2016) found even the internal wave shape of undulating programming — high-rep waves against low-rep waves — made no difference to either strength or size once volume was matched. If you train for size, the model genuinely is not the variable.

The sharpest point, and it is rarely said out loud: the experiment has never been run. Afonso et al. (2019), reviewing the meta-analyses and the 21 studies underneath them, found that every single “non-periodized” control arm was a constant program — the same sets, the same reps, the same relative load, week after week. But periodization’s actual claim was never “planned training beats doing the identical workout for twelve weeks.” Its claim is that planned variation beats unplanned variation. No located trial has ever used a varied-but-unplanned comparator. Afonso also found that 8 of those 21 studies were comparing something else entirely — total volume, or coaching supervision — rather than periodization models, and warns that pooling weak primary trials doesn’t produce a good answer; it produces a confident-looking one.

There is a second explanation for the strength effect that nobody has ruled out. Periodized programs typically end their cycle at low reps and high intensity — precisely the training that best rehearses a heavy single — immediately before the study measures a 1RM. Nunes and colleagues (2018), in a published *Sports Medicine* comment on the Williams meta-analysis, raise this as a specificity confound: part of periodization’s apparent strength edge may be the outcome test being flattered by the last few weeks of training rather than any difference in adaptation. They also propose the fix — give the non-periodized arm a heavy, low-rep loading zone too, and see whether the gap survives (the argument is relayed in Evans’ 2019 mini-review). Notice how cleanly it fits the pattern above — the effect shows up for strength, where a 1RM test can be rehearsed, and vanishes for hypertrophy, where it can’t. That isn’t proof. It is the most parsimonious alternative explanation on the table, and no trial we located has controlled for it. It is also badly under-reported in popular coverage, which is why it gets a paragraph here instead of a footnote.

The theoretical fight is live, and we’re not going to pretend somebody won it. Kiely (2018) argues that periodization’s founding scientific platform — Selye’s General Adaptation Syndrome — has been superseded by allostasis theory, and asks whether the model can be justified once its foundation has disintegrated. Steele, Fisher, Loenneke & Buckner push further, arguing periodization doesn’t meet the standard of a scientific theory at all because it lacks a consensus definition yielding testable predictions, and that plain progressive overload plus adequate recovery explain the results more economically — though that paper is a SportRxiv preprint and has not been peer-reviewed. The defence is substantive too: Stone, DeWeese and colleagues (2018) argue the critics read Selye’s work too narrowly, and the NSCA’s framing raises something the lab literature genuinely doesn’t address — that periodization’s applied claim is about producing a peak inside a narrow competition window through managed fatigue, while the trials measure average adaptation over twelve weeks. Those are different outcomes. What we could not find, in any source, is a pro-periodization theorist answering the varied-comparator or timing-prediction critiques on their own terms. The two camps are talking past each other, and saying so is more useful than manufacturing a resolution.

What does hold up. Two things, and they are not small. First, progressive overload and adequate volume are the load-bearing variables — the apparent periodization advantage concentrates precisely in the designs that failed to equate them, which is what you would expect if the work, not the scheduling, is doing the job. Second, tapering for competition is the one sub-question in genuinely better shape. Travis et al. (2020) reviewed peaking in powerlifting and weightlifting and found a 7-day step taper (volume cut 31.6–67%) producing +6.4%, or +8.1 kg, on the back squat (n = 15), and a 14-day exponential taper at 50% volume reduction producing +5.3% to +9.5% improvements (+8.2 to +14.8 kg, n = 9–12). Bazyler et al. (2021), a randomised trial in 16 competitive powerlifters, found step and exponential tapers essentially equivalent (back squat g = 0.54 both; Wilks g = 0.55 both). Travis’ practical read: cut volume 30–70% while holding intensity at or above 85% of 1RM across one to two weeks, then take 2–7 days off — but not much beyond two weeks, where measurable loss starts. Note the irony: the fatigue-management mechanism underneath periodization has better evidence than periodization’s organisational claims do.

So what should a lifter actually do? Pick a structure you will genuinely run, and then run it. The model matters less than the execution, and the execution is mostly volume, progression, and showing up — a program that fits your week and holds your interest for six months beats a theoretically superior one you abandon in three, and on this evidence “theoretically superior” is doing a great deal of unearned work in that sentence. If you’re training toward a meet, plan the taper; that part is real and worth getting right. Otherwise, use a structure because it makes progression legible and keeps you consistent, which is a genuinely good reason. It just isn’t the reason you were sold one.

Sources

  1. Moesgaard et al. (2022), Sports Medicine — periodization in volume-equated programs (the anchor meta-analysis)
  2. Williams et al. (2017), Sports Medicine — periodized vs non-periodized 1RM; removing funnel-plot outliers halved the effect
  3. Nunes et al. (2018), Sports Medicine — comment on Williams; the specificity / 1RM-testing-artifact confound
  4. Grgic et al. (2017), PeerJ — linear vs daily-undulating, hypertrophy, volume equated
  5. Harries et al. (2015), JSCR — linear vs undulating meta-analysis (k=17, n=510)
  6. Rhea & Alderman (2004) — the original ES=0.84 periodized vs non-periodized meta-analysis
  7. Afonso et al. (2019), Frontiers in Physiology — systematic review of meta-analyses; the missing varied comparator
  8. Kiely (2018), Sports Medicine — Periodization Theory: Confronting an Inconvenient Truth
  9. Steele, Fisher, Loenneke & Buckner — The Myth of Periodisation (SportRxiv PREPRINT, not peer-reviewed)
  10. Stone, DeWeese et al. (2018), Sports Medicine — the GAS defence of periodization
  11. NSCA — central concepts related to periodization (fitness-fatigue and the competition peak)
  12. Evans (2019), Frontiers — balanced mini-review; where the 1RM-test specificity confound is discussed
  13. Travis et al. (2020), Sports — tapering and peaking for powerlifting performance
  14. Bazyler et al. (2021), Frontiers — step vs exponential taper RCT in competitive powerlifters
  15. Apel, Lacey & Kell (2011), JSCR — volume- AND intensity-equated; favoured linear
  16. Rhea et al. (2002), JSCR — DUP vs linear, 12 weeks
  17. Klemp et al. (2016) — volume-equated high- vs low-rep daily undulating programming

The whole tracker. One dollar a month.

iOS beta launching shortly on TestFlight. Be first in line.

Join the iOS beta →