A contract research site running its first peptide-obesity study enrolls 22 subjects, splits them 11 and 11, runs the protocol for 16 weeks, and gets a p-value of 0.11 on the primary endpoint. The compound may work. The trial simply never had a mathematical chance of proving it. This is the single most common failure mode in small-scale GLP-1 and next-generation multi-agonist peptide research: a study built around a hypothesis, a budget, and an enrollment deadline, but never around a power calculation. The drug doesn't fail. The design does.
This matters more in 2026 than it did five years ago, because the pipeline behind semaglutide and tirzepatide has expanded into triple-receptor agonists, amylin co-agonists, and oral small-molecule GLP-1 candidates, most of which are first being tested in exactly the kind of small, resource-constrained cohorts where power problems are most common. Getting the statistics right at the design stage is not a formality — it is the difference between a study that can be published and cited, and one that produces a number nobody can interpret.
The Small-Trial Statistical Power Problem in GLP-3-Class Research
"GLP-3" is not a receptor classification recognized by the FDA, EMA, or the peer-reviewed pharmacology literature. It is used informally, including on this site, as shorthand for the emerging class of dual- and triple-receptor agonists — compounds acting on GLP-1, GIP, and glucagon receptors, such as retatrutide — that sit conceptually one step beyond single-target GLP-1 receptor agonists like semaglutide. Investigators writing protocols should name the actual receptor targets in the study documents; regulatory reviewers will not accept an undefined mechanism class in an IND submission.
The statistical power problem is the same regardless of receptor target. Power is the probability that a trial detects a true effect if one exists, and it is a function of four inputs locked together: sample size, effect size, variance, and the alpha threshold. Move any one without recalculating the others, and the study is either overpowered (wasting subjects and budget) or underpowered (unable to detect a real signal). Retatrutide's phase 2 data (PMID 37366315) showed mean weight loss up to 24.2% at the highest dose over 48 weeks in n=48-49 per arm — a large effect size that made a comparatively small per-arm size adequate. A trial testing a more incremental effect, such as a 3-4% weight-loss difference between two dosing regimens of the same compound, needs a substantially larger cohort to detect it at the same confidence level.
What Small Peptide Research Actually Means for Effect Size Assumptions
Effect size is the number every underpowered trial gets wrong first. Investigators frequently borrow an effect size from a pivotal phase 3 trial with a healthier, larger, more homogeneous population and apply it to a phase 1/2 population with more comorbidity, more variance, and less rigorous adherence monitoring. The result is a power calculation that looks adequate on paper and is not adequate in the actual patient population being enrolled.
A more defensible approach pulls the variance estimate from a trial population as close as possible to the one being studied. The SELECT trial (NCT03574597), which enrolled over 17,600 subjects with cardiovascular disease and overweight or obesity without diabetes, is a useful variance reference for cardiometabolic secondary endpoints precisely because its population resembles many real-world small-trial cohorts more closely than a lean, diabetes-free phase 1 population does. Details on how SELECT's four-year follow-up data breaks down by subgroup are covered in a separate review of the SELECT trial's cardiovascular outcome data.
Three questions should be answered in the protocol before a single subject is screened:
- What is the smallest effect size that would be clinically meaningful, not just statistically detectable?
- What variance estimate comes from a population comparable to the one being enrolled?
- What alpha and power thresholds are being used, and are they consistent with what a peer-reviewed journal will expect at submission?
Calculating Sample Size Before Enrollment Opens, Not After
The standard two-sample power formula for a continuous endpoint (n = 2(Z_alpha/2 + Z_beta)^2 * sigma^2 / delta^2) is not complicated, but it is routinely skipped or run backward — plugging in the sample size the budget allows and solving for the effect size the study can detect, then quietly hoping the true effect is at least that large. Julious' commonly cited rule of thumb (PMID 15898135) sets n=12 per arm as a floor for a pilot study intended to estimate variance and feasibility, not to confirm efficacy. Whitehead et al. (PMID 26908536) extended this with formal sample-size formulas for pilot trials feeding into a larger confirmatory study, which is the more defensible framework for early peptide research that will eventually scale to a phase 2b or 3 program.
For a percent-body-weight-change endpoint with a standard deviation of roughly 6-7 percentage points (consistent with variance reported across the STEP and SURMOUNT program publications), detecting a 5-percentage-point between-group difference at 80% power and alpha=0.05 requires approximately n=55-65 per arm before any attrition adjustment. A trial run at n=20 per arm under those same assumptions has power closer to 25-30% — it is more likely to miss a true effect than to find one.
The Real Cost of an Underpowered Trial
An underpowered trial does not fail quietly. It produces a number, a confidence interval, and a p-value, and that result gets cited. A non-significant finding from an underpowered n=20 study gets treated in subsequent literature reviews and meta-analyses as evidence of "no effect," when the correct interpretation is "no conclusion possible." That distinction rarely survives being cited three papers downstream.
The financial math compounds the problem. A typical small peptide trial running 16 weeks with monthly labs, imaging, and a coordinator on-site costs somewhere between $8,000 and $15,000 per enrolled subject once screening failures, coordinator time, and site overhead are included. A study run at n=20 per arm that produces an uninterpretable result has spent $320,000-$600,000 to answer nothing. Adding the additional 20-30 subjects per arm needed to reach adequate power, at the same per-subject cost, adds $160,000-$450,000 — a real number, but one that is usually smaller than the sunk cost of the underpowered study plus the cost of running a second one to correct it.
Reputational cost is harder to price but not smaller. A site or sponsor that publishes two or three underpowered null results develops a track record that makes the next funding conversation — with an IRB, a CRO partner, or an investor — measurably harder.
Adaptive and Bayesian Designs for Small Cohorts
FDA's 2019 guidance on adaptive designs for drugs and biologics formalized what many small-trial statisticians had already been doing informally: building pre-specified rules into the protocol that allow sample size, randomization ratio, or stopping decisions to change based on accumulating data, without inflating the trial's Type I error rate. This is particularly relevant in peptide research where early dose-ranging work often runs on constrained budgets.
Group-sequential designs allow a trial to stop early for overwhelming efficacy or futility at pre-planned interim analyses, which can reduce total exposure and cost by 20-30% in trials where the true effect turns out to be larger or smaller than assumed. Bayesian response-adaptive randomization shifts allocation toward the better-performing arm as data accumulate, which is attractive in multi-arm dose-finding studies — the kind increasingly used for triple-receptor agonist candidates — because it exposes fewer subjects to doses that are clearly underperforming.
The tradeoff is operational complexity. Adaptive designs require an independent data monitoring committee, a locked statistical analysis plan before unblinding, and a biostatistician available for each interim look, which raises the up-front cost of protocol development. For a single-site trial under n=100 total, that added cost needs to be weighed against the smaller total enrollment adaptive designs typically require. The Retatrutide phase 2 program's dose-ranging structure, discussed in more detail in a review of its 48-week weight and glycemic endpoint data, illustrates how a multi-arm design can extract dose-response information from a moderate total sample.
Choosing Endpoints That Maximize Power Without Inflating Type I Error
Endpoint choice has as much influence on required sample size as effect size does. Continuous endpoints — percent body weight change, HbA1c reduction, liver fat fraction on MRI-PDFF — carry more statistical information per subject than binary endpoints like "achieved 5% weight loss, yes/no," and generally require 30-40% fewer subjects to reach the same power. A trial with a binary responder endpoint that could be run as a continuous change-from-baseline analysis is leaving power on the table for no analytical benefit.
Composite endpoints introduce a different risk. Combining weight loss, glycemic control, and a liver or renal marker into a single composite can appear to improve power by capturing more events, but it complicates interpretation of which component drove the result, and reviewers increasingly ask for pre-specified hierarchical testing to control the overall Type I error rate when multiple endpoints are tested. The SYNERGY-NASH trial's approach to liver-specific histologic and imaging endpoints in tirzepatide research, covered in a 52-week results analysis, is a useful reference for how a single well-chosen primary endpoint with clearly ranked secondary endpoints avoids the multiplicity problem.
Surrogate endpoints deserve particular scrutiny in small trials. A biomarker change that has not been validated against a hard clinical outcome in the specific population being studied can produce a statistically significant result that does not translate to the outcome a clinician or patient actually cares about. Renal function markers analyzed in the FLOW trial's kidney endpoint work provide one example of how surrogate and hard outcome data can diverge, detailed in a kidney endpoint analysis of the FLOW trial.
Budgeting a Small Trial for Adequate Power, Not Just Adequate Funding
Running a research site is a business, and the budget conversation for a trial has to happen before the protocol is finalized, not after. A common mistake is setting the enrollment target to match the available budget, then reverse-engineering a power justification that fits — a sequence that regulatory statisticians and peer reviewers can usually identify on sight because the assumed effect size ends up implausibly large relative to prior literature.
A more defensible sequence starts with the power calculation, then prices the trial against that number. If the power calculation returns n=70 per arm and the budget only supports n=40 per arm, the honest options are: narrow the primary endpoint to one with a larger expected effect size, seek co-funding or a CRO partnership to cover the gap, or explicitly re-scope the study as a pilot with a variance-estimation goal rather than a confirmatory one. All three are legitimate. Quietly running the underpowered version and reporting the p-value as if it answered the confirmatory question is not.
Line-item budgeting for a 16-week, two-arm peptide trial at n=60 per arm (n=120 total) typically breaks down as follows: subject stipends and travel reimbursement ($400-$800 per subject), lab panels at 4-6 timepoints ($150-$300 per panel), coordinator and PI time (roughly 35-40% of total budget), IRB and regulatory fees ($15,000-$30,000 flat), and a 10-15% contingency for protocol amendments. Total cost in this range commonly lands between $900,000 and $1.6 million — a number that should be on the table before enrollment targets are set, not discovered after the first monthly burn report.
Managing Attrition and Missing Data Without Losing Power
Attrition is the second-most-common power killer after a bad effect-size assumption, and it is the cheapest to fix. Trials of GLP-1-class and multi-agonist peptides running 12 weeks or longer commonly see 15-25% dropout, driven predominantly by gastrointestinal tolerability issues — nausea, vomiting, and constipation reported in the STEP and SURMOUNT program safety data. A power calculation that assumes zero dropout and enrolls exactly the calculated n will finish the trial underpowered by definition.
The fix is arithmetic: if the power calculation says n=50 per arm and 20% attrition is expected, enroll n=63 per arm (50 / 0.8), not n=50. This single adjustment, applied at the design stage, is inexpensive relative to discovering mid-trial that the completer analysis is underpowered and there is no budget left to add subjects.
Missing data handling also needs to be pre-specified rather than decided after unblinding. Last-observation-carried-forward analyses are increasingly viewed skeptically by reviewers because they can bias results toward the null or away from it depending on the dropout pattern; mixed-model repeated-measures (MMRM) analysis, which uses all available data points rather than a single carried-forward value, has become the more accepted standard in recent GLP-1 program publications and should be named explicitly in the statistical analysis plan before enrollment begins.
Reporting Standards That Determine Whether Small-Trial Data Gets Taken Seriously
A well-powered small trial can still fail to be useful if it is reported poorly. CONSORT reporting guidelines, which most peer-reviewed journals now require for randomized trial submissions, ask for the exact power calculation inputs — assumed effect size, variance, alpha, and power — to be stated in the methods section, along with the actual enrollment, attrition, and completion numbers in a flow diagram. Omitting the power calculation entirely, still common in smaller peptide research publications, is one of the fastest ways to get a manuscript sent back for revision or rejected outright.
Pre-registration on ClinicalTrials.gov before enrollment begins, with the primary and secondary endpoints locked, has also become close to a de facto requirement for publication in indexed journals and is required outright for any study intending to support a future IND submission. A protocol that is registered after data collection has already started, or that changes its primary endpoint after unblinding, invites the kind of scrutiny that can sink an otherwise sound dataset. The oral orforglipron program's ACHIEVE trial reporting, discussed in a review of its phase 3 weight-loss and glycemic endpoints, and the CagriSema REDEFINE-1 program's phase 3 data reporting structure are both useful models for how endpoint hierarchy and pre-specified analysis plans are laid out in current large-scale GLP-1 program publications, even though most small research trials operate at a fraction of that scale.
Building the Protocol: A Practical Sequence
The sequence that keeps a small peptide trial out of the underpowered category is straightforward, even if it is frequently skipped under time pressure. First, define the smallest clinically meaningful effect size with input from a clinician who will actually use the result, not just a statistician working from a spreadsheet. Second, pull variance estimates from a population as close as possible to the one being enrolled, favoring RCT or meta-analysis data over observational or animal-model figures. Third, run the power calculation with a realistic alpha (0.05 unless multiplicity requires adjustment) and target at least 80% power, then inflate the resulting n for expected attrition using site-specific or published dropout rates for the compound class.
Fourth, decide before enrollment whether an adaptive or fixed design fits the budget and infrastructure, and lock the statistical analysis plan, including the missing-data method, before any interim data is reviewed. Fifth, register the trial on ClinicalTrials.gov with the primary endpoint specified, and budget for the full calculated n rather than the number that happens to fit this quarter's operating budget. A site that runs this sequence consistently will produce fewer trials, but every one of them will produce a number that means something — which is the entire point of running the study.
This article summarizes research and does not constitute medical advice. Consult a licensed clinician for diagnosis, treatment, or any decisions about medications or supplements.