A recent review of methods used in clinical trials supported by NIH identified 12 common errors of alignment between 1) the methods proposed for assignment of participants to study arms and for delivery of intervention to participants and 2) the methods used for sample size estimation and proposed for data analysis (Moyer et al., 2026; Murray et al., 2026). Those errors are presented here, summarized by study design.
Randomized Control Trials (RCTs)
Among RCTs, the primary error was not accounting for multiplicity. Multiplicity occurs when there are multiple hypothesis tests, multiple primary outcomes, or multiple comparisons each tested at the same nominal type 1 error rate; failure to address multiplicity can increase the type 1 error rate (Food and Drug Administration, 2022). However, instances where multiplicity may not be a concern include:
- If the study requires that all primary outcomes be significant at the nominal alpha level in order to declare that the intervention was effective, there is no multiplicity issue as there is only one path that leads to a successful outcome for the study.
- Similarly, if several outcomes are combined into a single composite, there is no multiplicity issue.
- Finally, if one primary outcome is a safety measure and the other primary outcome is an efficacy or effectiveness outcome, there is no multiplicity issue.
Individually Randomized Group Treatment (IRGT) Trials
Among IRGT trials, the primary errors were not accounting for the shared intervention agents or group-formatted intervention components, not accounting for multiple membership, and not accounting for multiplicity; each of those errors can inflate the type 1 error rate.
- The classic error in an IRGT trial is not accounting for the small groups or shared interventions used to deliver at least part of the intervention (Pals et al., 2008). This error ignores the intraclass correlation and is expected to develop for participants in the same small group or who interact with the same shared agent. The solution is to include the small groups or agents as sources of random variation, e.g., as random effects in a mixed model.
- A more subtle form of this error is not accounting for multiple membership. Multiple membership exists when participants receive some of their intervention from more than one shared intervention agent. For example, a weight loss trial might have both a dietician and an exercise physiologist work with each participant, focusing on diet and exercise respectively. Failure to account for multiple membership can increase the type 1 error rate even if the analysis accounts for the primary shared intervention agent (Moyer et al., 2024).
- The multiplicity issues are the same as for RCTs.
- A few studies used a single-agent design; that design is not recommended because it completely confounds the agent with study arm making it impossible to evaluate the effect of the intervention (Moyer et al., 2024; Varnell et al., 2001).
Group- or Cluster-Randomized Trials (GRTs)
Among GRTs, the primary errors were not accounting for multiplicity, not accounting for groups or clusters, not accounting for cross-classification, not accounting for a time x cluster component of variance in a repeated measures analysis, and including only a single group or cluster in each study arm.
- The multiplicity issues are the same as for RCTs.
- The classic error in a GRT is not accounting for groups or clusters; doing so will increase the type 1 error rate whenever the correlation among observations taken on members of the same group or cluster is positive, as it almost always is (Cornfield, 1978; Donner & Klar, 2000; Murray, 1998).
- A variation on that error is not accounting for agents working across groups or clusters. This creates cross-classification – participants are nested within more than one non-hierarchical group or cluster (Cafri et al., 2015; Goldstein, 1994). Consider a school-based trial where the intervention is delivered primarily by teachers but in part by a navigator. The teachers will be unique to each school but the navigators work with several schools in the same arm. Contact with the navigator threatens the independence of the schools and so must be accounted for in the sample size and in the data analysis. Failure to do so can result in an inflated type 1 error rate.
- In a parallel GRT with two or more time periods in the analysis, it will be important to allow not only for variation due to groups or clusters but also variation over time due to groups or clusters. This requires a component of variance for the time x group or time x cluster interaction. The main effect for group or cluster allows for correlation among participants in the same group or cluster – the familiar intraclass correlation (ICC). The interaction with time allows that correlation to vary at each time. Failure to account for the time x group or time x cluster component of variance can lead to an inflated type 1 error rate (Moyer et al., 2022).
- A few GRTs proposed a one-group-per-arm design (Moyer et al., 2026; Murray et al., 2026). That design is not recommended because it completely confounds the group or cluster with study arm making it impossible to evaluate the effect of the intervention (Varnell et al., 2001).
Stepped Wedge Group- or Cluster-Randomized Trials (SWGRTs)
Among SWGRTs, the primary errors were not allowing for a time-varying intervention effect when there were two or more follow-up periods, not accounting for the groups or clusters randomized to sequences, multiplicity, and cross-classification. Each of these errors can result in an inflated type 1 error rate.
- Time-varying intervention effects can render the results from a standard SWGRT analysis highly inaccurate (e.g., Hughes et al., 2024; Kenny et al., 2022; Lee et al., 2025; Maleyeff et al., 2023). The prudent course is to allow for such effects in the analytic plan if there are at least two follow-up periods.
- Failure to account for groups or clusters randomized to sequences is a variation on the classic error in a GRT. Doing so will increase the type 1 error rate whenever the correlation among observations taken on members of the same group or cluster is positive, as it almost always is.
- The multiplicity issues are the same as for RCTs.
- The cross-classification issues are the same as for GRTs.
