Designing Effective Patient Inclusion Criteria for GLP-1/GLP-3 Clinical Trials: A Framework for Diversity and Validity

A trial coordinator running enrollment for a mid-size Phase 3 obesity study once described the problem bluntly: of the first 400 screening calls, fewer than 60 produced eligible participants who were also male, over age 65, or Black — despite those groups making up a large share of the target treatment population once the drug reached market. The inclusion criteria on paper looked standard: BMI 30-45 kg/m2, age 18-75, no history of pancreatitis. In practice, the recruitment funnel, the site locations, and the eligibility thresholds combined to produce a study population that looked nothing like the patients who would eventually receive a prescription. This is not a hypothetical failure mode. It is the documented pattern across the GLP-1 and GLP-3 receptor agonist trial literature, and it is the central design problem this article addresses.

Why Inclusion Criteria Are a Validity Lever, Not Just a Screening Checklist

Inclusion and exclusion criteria are usually treated as a compliance artifact — a list to satisfy an institutional review board. In reality, every threshold chosen (BMI cutoff, eGFR minimum, HbA1c range, psychiatric comorbidity exclusion) is a lever that trades internal validity against external validity. Narrow criteria reduce the number of confounding variables and produce a cleaner effect-size estimate. Broad criteria produce a messier signal but one that generalizes to the population clinicians will actually treat.

The pivotal semaglutide trial STEP 1 (NCT03548935, n=1,961) illustrates the tradeoff. The trial required BMI ≥30 kg/m2 (or ≥27 with a weight-related comorbidity), age 18 and older, and no diabetes diagnosis. Those criteria produced a clean 14.9% mean placebo-subtracted weight reduction at 68 weeks (NEJM, Wilding et al., 2021) — a strong internal validity result. But the enrolled population was approximately 74% female and about 75% White, non-Hispanic, per the published baseline characteristics table.

That demographic skew is not necessarily a flaw in the eligibility criteria themselves — BMI and comorbidity thresholds were not written to exclude men or non-White participants. It is a downstream effect of where and how sites recruited against those criteria. Designing inclusion criteria effectively means anticipating that downstream effect at the protocol-writing stage, not correcting for it after unblinding.

Baseline Criteria in Pivotal GLP-1 Obesity Trials

Reviewing the eligibility frameworks used in STEP 1, SURMOUNT-1 (NCT04184622, tirzepatide, n=2,539), and the cardiovascular outcomes trial SELECT (NCT03574597, semaglutide, n=17,604) gives a useful baseline for what "standard" inclusion criteria look like before diversity-oriented adjustments are applied.

  • BMI threshold: Typically ≥30 kg/m2 alone, or ≥27 kg/m2 with at least one weight-related comorbidity (hypertension, dyslipidemia, obstructive sleep apnea).
  • Age range: Most obesity trials cap at 75, though SELECT extended to age 75 with no upper limit waiver process for older adults with stable cardiovascular disease.
  • Renal function: eGFR cutoffs commonly exclude below 15-30 mL/min/1.73m2, which disproportionately removes older adults and populations with higher rates of undiagnosed chronic kidney disease.
  • Psychiatric history: Active suicidal ideation within a defined lookback window and history of severe depression are common exclusions, which can disproportionately affect enrollment from lower-income populations with higher background rates of untreated depression.
  • Personal/family history of MEN2 or medullary thyroid carcinoma: A hard exclusion across essentially all GLP-1 receptor agonist trials, driven by the rodent C-cell tumor signal described in the FDA label and nonclinical pharmacology data.

These criteria are clinically defensible on safety grounds. The diversity problem emerges less from the criteria themselves and more from how recruitment funnels interact with them across different populations.

Widening Eligibility Bands Without Eroding Safety Signal Detection

Sponsors evaluating whether to widen an eligibility band should treat each threshold change as a separate risk-benefit calculation, not a single blanket policy. Lowering the eGFR floor from 30 to 15 mL/min/1.73m2, for example, expands the eligible population meaningfully — chronic kidney disease prevalence rises sharply with age and is disproportionately represented in Black and Hispanic populations in U.S. epidemiological data — but it also introduces a population at higher baseline risk for volume depletion and acute kidney injury from GI-related fluid loss.

The methodologically sound response is not to avoid the broader eGFR band but to pair it with a compensating monitoring protocol: more frequent renal function labs (e.g., at weeks 4, 8, 16 rather than only at baseline and study end), a lower threshold for temporary dose interruption, and explicit criteria for independent Data Safety Monitoring Board review of renal adverse events in that subgroup.

The same logic applies to psychiatric history exclusions. Rather than a binary exclusion for "history of depression," a tiered approach — structured screening instrument (e.g., PHQ-9) at baseline, exclusion only above a defined severity threshold, and scheduled reassessment during the trial — retains a meaningful safety guardrail while not systematically excluding populations with higher background rates of treated, stable depression. This is a methodological recommendation, not an RCT-validated protocol; sponsors adopting it should document it prospectively in the statistical analysis plan.

The Documented Racial, Ethnic, and Sex Representation Gap

The representation gap in GLP-1 trials is well documented in the published baseline tables, not merely alleged. STEP 1's population was roughly 75% White; SURMOUNT-1's population was approximately 70.6% White and 66.4% female (NEJM, Jastreboff et al., 2022). SELECT, with its much larger n=17,604 and cardiovascular-outcomes design, achieved somewhat broader geographic spread across more than 40 countries but still reported a White-majority population and a male-skewed sample of roughly 72% male — the inverse skew from the obesity trials, reflecting cardiovascular trial recruitment patterns rather than obesity trial patterns.

The clinical significance of this gap is that subgroup analyses by race or sex in these trials are underpowered by design. A trial built to detect a population-level effect size with 90% power at alpha=0.05 is rarely built to detect the same effect size within a subgroup that represents 10-15% of total enrollment. Published subgroup forest plots from these trials should be read as hypothesis-generating, not confirmatory, for smaller demographic strata — a distinction the FDA guidance on diversity action plans explicitly addresses.

Socioeconomic and Geographic Barriers to Enrollment

Eligibility criteria on paper interact with practical barriers that are rarely written into the protocol but shape who actually completes screening. Trial visit schedules requiring in-person visits every 2-4 weeks for 68 weeks impose a substantial burden on participants without paid time off, reliable transportation, or childcare — burdens that are not evenly distributed across income levels.

Site selection compounds this. Academic medical centers concentrated in a handful of metropolitan areas produce convenience samples skewed toward populations within a reasonable driving radius of a tertiary care center, which correlates with higher household income and insurance coverage. A trial with 25 sites concentrated in five metropolitan areas will structurally underrepresent rural populations regardless of how the BMI or age criteria are written.

Practical mitigations documented in the trial methodology literature include: reimbursing transportation and lodging costs directly rather than through delayed reimbursement, offering visit windows outside standard business hours, and pre-specifying enrollment targets by site type (rural vs. urban, academic vs. community) rather than leaving site mix to opportunistic contracting. None of these are criteria changes in the strict sense, but they determine whether criteria that look inclusive on paper translate into an inclusive enrolled sample.

Decentralized and Hybrid Trial Elements That Improve Diversity

Decentralized clinical trial (DCT) elements — remote consent, home health visits for vital signs and injections, telehealth follow-up visits, and local lab draw partnerships — have been piloted across multiple therapeutic areas as a structural response to the geographic access problem. For GLP-1/GLP-3 trials specifically, home-administered injectable dosing with remote monitoring reduces the in-person visit burden that disproportionately excludes working participants and those without transportation.

A hybrid design that keeps in-person visits for safety-critical assessments (baseline labs, ECG where required, initial injection training) while shifting routine follow-up to telehealth and local phlebotomy partners can meaningfully widen the geographic catchment area without relaxing any clinical eligibility criteria. This is a trial infrastructure decision, not an inclusion criteria decision, but it directly affects which eligible patients actually enroll and complete the study.

Evidence tier for DCT effectiveness specifically in obesity pharmacotherapy trials is still limited — most published data comes from oncology and cardiovascular DCT pilots — so sponsors should treat this as an operationally reasonable but not yet RCT-confirmed strategy for this drug class specifically.

Statistical Design: Stratified Randomization and Subgroup Power

Once a sponsor decides diversity of enrollment is a prespecified objective rather than an incidental outcome, the statistical analysis plan needs to reflect that decision before enrollment opens, not after. Stratified randomization by race/ethnicity, sex, baseline BMI class, and renal function ensures that treatment and placebo arms remain balanced within each subgroup, which prevents a scenario where, for example, 80% of Black participants land in the placebo arm by chance.

Powering a trial to detect a clinically meaningful effect size within the smallest prespecified subgroup — rather than only at the whole-population level — typically requires increasing total enrollment by roughly 8-15%, depending on the anticipated subgroup allocation ratio and expected effect size. For a trial targeting an overall n=2,000, this might mean enrolling closer to n=2,200-2,300, with explicit enrollment caps preventing any single demographic stratum from filling disproportionately early and closing out enrollment for underrepresented strata.

This approach has direct precedent in oncology trial design, where FDA has increasingly requested subgroup-powered analyses for approvals affecting genetically or demographically distinct populations. Applying the same statistical discipline to GLP-1/GLP-3 obesity and metabolic trials is a design choice sponsors can make now, ahead of any regulatory mandate to do so.

Regulatory Context: FDA Diversity Action Plans Under FDORA

The Food and Drug Omnibus Reform Act (FDORA), signed into law in December 2022, requires sponsors of certain Phase 3 (and some Phase 2) clinical trials to submit a diversity action plan to FDA describing enrollment goals by race, ethnicity, and sex, along with the rationale for those goals and the specific strategies for reaching them. FDA finalized guidance implementing this requirement in 2024, describing expectations for how sponsors should identify the disease prevalence in affected populations and set enrollment targets accordingly.

This regulatory shift changes the calculus for sponsors designing GLP-1/GLP-3 trials going forward. A diversity action plan is not simply a narrative statement of intent — it requires sponsors to connect specific inclusion/exclusion criteria and site selection decisions to a quantified enrollment target, and to explain deviations if targets are not met by the time of database lock. Sponsors that build this into protocol design from the outset, rather than retrofitting a narrative after enrollment closes, are better positioned for a smoother FDA review process, though the guidance does not specify enrollment shortfalls as an automatic basis for rejection.

A Practical Framework for Writing Inclusion Criteria

Bringing the preceding sections together, a workable framework for protocol teams drafting GLP-1/GLP-3 inclusion criteria includes five concrete steps. First, map each proposed exclusion criterion against its epidemiological prevalence in the target population, using published NHANES or equivalent registry data, to identify which criteria will disproportionately exclude specific demographic groups. Second, for any criterion with a disproportionate exclusion effect (renal function, psychiatric comorbidity, BMI floor), design a compensating monitoring protocol rather than defaulting to the narrowest threshold.

Third, prespecify enrollment targets by race, ethnicity, sex, and site type in the statistical analysis plan before first patient enrolled, and recalculate total sample size to preserve power within the smallest stratum. Fourth, evaluate site mix and DCT elements specifically for their effect on geographic and socioeconomic access, treating site selection as an extension of the inclusion criteria rather than a separate operational decision. Fifth, document all of this in a diversity action plan aligned with FDA's 2024 guidance, submitted with enough lead time to incorporate agency feedback before enrollment opens.

None of these steps requires abandoning rigorous safety-driven exclusion criteria. The evidence from STEP, SURMOUNT, and SELECT shows that a trial can produce a strong, internally valid effect-size estimate and still fall short on external validity — and that gap is addressable at the design stage, not just at the interpretation stage. For readers exploring related trial design questions, additional context on how these compounds are compared across pivotal trials, how their receptor pharmacology is characterized, and how safety monitoring protocols are structured in practice can inform how inclusion criteria decisions interact with the broader evidence base for this drug class.

This article summarizes research and does not constitute medical advice. Consult a licensed clinician for diagnosis, treatment, or any decisions about medications or supplements.

Frequently asked questions

What are inclusion and exclusion criteria in a GLP-1 clinical trial?

Inclusion criteria define the population eligible for enrollment (e.g., BMI ≥30 kg/m2, age 18-75, HbA1c within a defined range); exclusion criteria remove participants for safety or confounding reasons (e.g., personal or family history of medullary thyroid carcinoma, eGFR <15 mL/min/1.73m2, active pancreatitis). Together they determine both internal validity and generalizability of the trial results.

Why do GLP-1 trials underrepresent certain racial and ethnic groups?

Historical GLP-1 pivotal trials recruited primarily through academic medical centers in high-income countries with limited language and transportation support, which favors White, insured, urban participants. STEP 1 (NCT03548935) reported a study population that was approximately 75% White, a pattern documented across multiple semaglutide and tirzepatide trials.

How does the FDA diversity action plan affect trial design?

Under the Food and Drug Omnibus Reform Act (FDORA) of 2022, sponsors of certain Phase 3 trials must submit a diversity action plan describing enrollment goals by race, ethnicity, and sex, along with the rationale and strategies to meet them. FDA finalized guidance on this requirement in 2024, and non-compliance can affect application review timelines.

Does broadening inclusion criteria compromise trial safety?

Not inherently. Broadening criteria such as eGFR thresholds or psychiatric comorbidity limits requires added monitoring — more frequent labs, structured psychiatric screening, or independent safety monitoring board review — rather than removing safeguards. Evidence tier for these adaptations is largely observational and methodological, not RCT-confirmed for every criterion change.

What sample size adjustment is needed for stratified enrollment?

Stratifying randomization by demographic or clinical subgroup generally requires increasing total enrollment by roughly 8-15% to preserve statistical power (typically 80-90%) within the smallest prespecified subgroup, depending on the anticipated effect size and subgroup allocation ratio.

Related references on this site

guide

Peptides 101: A Clinician's Reference

Reference guide on this site.

View →
guide

Handling, Reconstitution, and Storage of Research Peptides

Reference guide on this site.

View →
guide

How Weight-Loss Peptides Compare

Reference guide on this site.

View →