A site coordinator running a phase 2 metabolic trial pulls the enrollment tracker on a Friday afternoon: 260 candidates pre-screened over eight weeks, 91 consented to formal screening, and only 54 randomized. The math does not close the gap to the contracted target of 80 by month three. This is not a recruitment problem in the marketing sense — it is an inclusion criteria problem, and it is one of the most common reasons metabolic and obesity trials fall behind their enrollment timeline. Getting inclusion and exclusion criteria right before the first site activates is cheaper than fixing them with a protocol amendment six months in.
What "GLP-3" Actually Refers to in Current Research
Before criteria can be written, the mechanism has to be named correctly. GLP-3 is not a receptor recognized in current endocrine pharmacology alongside GLP-1 and GIP receptors. The term circulates informally, usually as shorthand for triple-hormone-receptor agonists that engage GIP, GLP-1, and glucagon receptors concurrently — the class that includes retatrutide, currently in phase 3 development. Protocol documents, informed consent forms, and IRB submissions should state the actual pharmacological targets rather than relying on the shorthand, since regulatory reviewers and ethics committees will expect precision on mechanism.
This distinction matters operationally. A trial studying a genuine triple agonist has a different adverse-event profile — more pronounced gastrointestinal effects and a glucagon-receptor-driven component affecting hepatic glucose output — than a single or dual agonist. Inclusion criteria written around outdated or imprecise mechanism descriptions tend to under-specify exclusions that turn out to matter, such as pre-existing hepatic impairment relevant to glucagon receptor activity.
Data from the retatrutide phase 2 program (NCT04881760) and the corresponding NEJM publication (Jastreboff et al., PMID 37366315) is the closest available reference point for what a well-specified triple-agonist inclusion framework looks like in practice, and it is worth reviewing before drafting a new protocol.
Core Demographic and Metabolic Inclusion Bands
Age and BMI bands are the first filter and the one most candidates fail on. SURMOUNT-1 (NCT04184622) enrolled adults with BMI ≥30, or ≥27 with at least one weight-related comorbidity, ages 18 and older without type 2 diabetes. That band was chosen deliberately: wide enough to reflect the population the drug would eventually be prescribed to, narrow enough to keep baseline variance in body weight from diluting the effect-size estimate on the primary endpoint.
A common design mistake is setting the BMI floor too low relative to the trial's statistical power calculation. If the study is powered to detect a percent-body-weight-change difference of roughly 10 percentage points between arms, admitting candidates at BMI 25–27 without a qualifying comorbidity adds heterogeneity that the sample size was never built to absorb. The fix is not more subjects — it is a tighter band matched to the power calculation from the start.
Age ceilings deserve the same scrutiny. Many phase 2/3 obesity and incretin trials cap enrollment around age 75 not for arbitrary reasons but because renal clearance, polypharmacy exposure, and cardiovascular comorbidity burden all rise sharply past that point, each of which independently confounds the safety analysis.
Comorbidity Criteria That Drive Screen-Fail Rates
Screen failure is where the enrollment funnel actually breaks, and it is rarely the BMI or age line that causes it. In practice, three comorbidity-related exclusions account for a disproportionate share of screen fails in incretin and multi-agonist trials:
- Personal or family history of medullary thyroid carcinoma or MEN2 syndrome — a standard exclusion tied to rodent C-cell tumor findings observed with GLP-1-class compounds.
- History of acute pancreatitis within the prior 6–12 months, given the class-wide signal for pancreatic enzyme elevation.
- Prior bariatric surgery within a defined lookback window (commonly 12–24 months), which confounds weight-trajectory endpoints independent of drug effect.
Across published incretin and obesity trials, aggregate screen-fail rates have ranged from roughly 20% up to 45% of candidates who reach formal screening, depending on how tightly the comorbidity exclusions are drawn and how motivated the referral population is. A site running a diabetes-endocrinology referral base will see a different fail pattern than one recruiting from a general primary care panel — the former screens out more on prior bariatric history, the latter on undiagnosed thyroid nodules found incidentally during baseline ultrasound.
Sponsors that underestimate this rate when setting per-site enrollment targets create a recurring problem: sites hit their contracted screening number but fall well short of randomization, and the study timeline absorbs the difference.
Baseline Laboratory and Biomarker Requirements
Every lab value on a baseline panel should map to a specific safety or eligibility rationale, not sit on the requisition form as boilerplate. For GLP-1 and multi-agonist trials, the standard panel typically includes fasting plasma glucose or HbA1c (to confirm the glycemic inclusion band or exclude undiagnosed diabetes in obesity-only cohorts), a comprehensive metabolic panel for renal and hepatic function, a fasting lipid panel, and serum calcitonin alongside a structured thyroid history questionnaire.
Renal function thresholds matter more than they are often given credit for in protocol design meetings. Reduced eGFR changes drug clearance and elevates hypoglycemia risk when a GLP-1 or triple-agonist compound is combined with background diabetes medication, so most protocols set an eGFR floor around 30–45 mL/min/1.73m² depending on the specific compound's renal elimination pathway.
Fasting C-peptide is a less universal but increasingly common addition, used to screen out participants with substantial beta-cell failure who would not respond to an incretin-based mechanism the same way — including them without this filter adds noise to a glycemic secondary endpoint even in an obesity-primary trial.
Pancreatic enzyme levels (amylase, lipase) at baseline give the study a reference point for adjudicating any on-study pancreatitis signal, which is essential given how closely regulators scrutinize this adverse event class for the entire GLP-1 and multi-agonist family.
Writing Criteria That Serve the Primary Endpoint
Inclusion criteria are not a generic eligibility checklist — they are an extension of the statistical analysis plan. If the primary endpoint is percent change in body weight at week 48, the inclusion band on baseline weight and BMI needs to produce a population where that endpoint is measurable with the assumed variance. If the primary endpoint is HbA1c reduction in a mixed obesity/type 2 diabetes population, the glycemic inclusion band has to be tight enough that baseline HbA1c does not itself become a major source of between-subject variance.
A frequent design error is treating secondary endpoints as an afterthought when setting criteria, then discovering during analysis that the enrolled population was never homogeneous enough to detect a meaningful signal on a secondary measure like waist circumference or lipid panel change. If a secondary endpoint matters enough to report, it needs its own eligibility consideration at the design stage, not a post-hoc subgroup analysis built on a population that was never selected with it in mind.
Run-in periods are the other lever available here. A 2–4 week placebo or diet-stabilization run-in before randomization filters out candidates who cannot maintain protocol adherence, which improves the signal-to-noise ratio on the primary endpoint without changing the nominal inclusion criteria at all. This is a design choice worth defending to a sponsor pushing for a faster enrollment timeline, because a bad run-in shortcut shows up later as excess dropout in the modified intention-to-treat population.
Enrollment Funnel Math and Site Operations
Enrollment planning has to work backward from the randomization target through the expected screen-fail rate, not forward from an optimistic assumption about how many people will walk in the door. Take a trial targeting n=200 randomized, with a 35% screen-fail rate observed in comparable prior incretin studies and a roughly 90% conversion rate from initial informed consent to completed formal screening. The pre-screening pool required works out to approximately 340 candidates identified and consented into the screening process — before accounting for further attrition during any run-in or washout period, which commonly removes another 5–10%.
Site-level throughput has to be sized against that number, not against the randomization target alone. A site that can realistically pre-screen 15 candidates per month against those assumptions will take roughly 23 months to deliver its share of a 340-candidate pool working alone — the reason multi-site and often multi-country designs are standard for anything beyond a small phase 1b cohort.
Coordinators who track screen-fail reasons by category, not just as an aggregate pass/fail number, generate the data a sponsor needs to decide whether a criterion is too restrictive relative to the population actually available. If 60% of screen fails at a given site are hitting the same exclusion — say, prior bariatric surgery — that is actionable information for a protocol amendment discussion, not just an operational footnote.
Protocol Amendments and Adaptive Criteria
Sponsors facing slow enrollment sometimes propose loosening inclusion criteria mid-trial — widening the BMI band, extending the bariatric surgery lookback window, or relaxing the HbA1c ceiling. Any such change requires a formal amendment through the IRB or central ethics committee, and it is not a purely administrative step. The statistical analysis plan was built around the originally defined population, and a material shift in eligibility partway through enrollment can create a cohort that is meaningfully different in the first half of enrollment versus the second.
The more defensible version of this adjustment happens at the design stage through adaptive or Bayesian elements built into the protocol from the outset, where pre-specified interim analyses allow eligibility refinement without introducing an unplanned discontinuity in the enrolled population. Retrofitting adaptivity onto a fixed protocol after enrollment has stalled is a much weaker position to argue to a regulator.
Documentation discipline matters here more than almost anywhere else in trial conduct. Every amendment needs a clear rationale tied to observed screening data, not simply "enrollment is behind schedule," because reviewers and downstream publication reviewers will ask why the population changed and whether it affects the comparability of the pooled dataset.
Regulatory and IRB Considerations for Vulnerable Populations
Obesity and metabolic trial populations frequently intersect with groups requiring additional protections — pregnant or lactating individuals are near-universally excluded given the absence of reproductive toxicology data for most compounds in this class, and reproductive-age participants are typically required to use a specified form of contraception and undergo periodic pregnancy testing throughout the study.
Adolescent and pediatric extensions, where they exist, require separate protocol sections with distinct dosing, growth-monitoring, and consent/assent procedures, and IRBs generally hold these to a higher scrutiny standard given the FDA's specific guidance on pediatric obesity drug development.
Participants with psychiatric comorbidities, particularly a history of suicidal ideation or behavior, warrant a structured screening instrument (commonly the Columbia-Suicide Severity Rating Scale) rather than a simple yes/no history question, given past regulatory attention to mood-related signals across weight-management drug classes historically. Building this into baseline and ongoing assessment schedules from the start avoids a costly amendment later if a safety signal emerges during the trial.
Monitoring, Retention, and Data Quality After Enrollment
Inclusion criteria do not stop mattering once randomization closes — they set the baseline against which every subsequent protocol deviation and dropout gets interpreted. A participant who develops an exclusionary condition (new pancreatitis diagnosis, new pregnancy, eGFR decline below the study floor) after randomization triggers a discontinuation and adjudication process, and the rate at which this happens is itself a data point on whether the exclusion criteria were calibrated correctly at entry.
Retention tracking should be segmented by the same demographic and comorbidity variables used for inclusion, because differential dropout by baseline characteristic — for example, higher discontinuation among participants at the upper end of the BMI band due to more pronounced gastrointestinal tolerability issues — can bias the modified intention-to-treat analysis even when the overall retention rate looks acceptable in aggregate.
A site running this kind of trial well treats the screening log as a living dataset, not a compliance record filed away after activation. Reviewing screen-fail categories monthly, flagging any criterion responsible for more than a quarter of fails, and raising that pattern to the sponsor's medical monitor before it becomes a six-month enrollment shortfall is the operational habit that separates sites that hit their randomization target from sites that need a no-cost extension.
This article summarizes research and does not constitute medical advice. Consult a licensed clinician for diagnosis, treatment, or any decisions about medications or supplements.