GLP-1 vs. GLP-3 Mimetic Peptides for Weight Loss: A Systematic Comparison of Clinical Trial Outcomes

A clinic operator running a medical weight-management program gets the same question every week from patients who have read about newer peptides online: "Should I wait for the GLP-3 drug instead of starting semaglutide now?" The honest answer requires pulling the actual trial data, because the marketing framing of "GLP-3" as a straightforward upgrade over GLP-1 glosses over real differences in trial design, dosing timelines, and tolerability data that matter for anyone making a near-term treatment decision. This GLP-1 vs GLP-3 mimetic peptides comparison works through what the published clinical trial data actually shows, where the numbers are directly comparable, and where they are not.

Defining the Two Drug Classes

GLP-1 receptor agonists — semaglutide (Ozempic/Wegovy), liraglutide (Saxenda) — act on a single receptor target, the glucagon-like peptide-1 receptor, slowing gastric emptying and increasing satiety signaling. Semaglutide's terminal half-life is approximately 165-184 hours, supporting once-weekly dosing.

What the field informally calls "GLP-3" mimetics are not single-receptor drugs at all. Retatrutide, the most advanced compound in this category, is a triple agonist engaging GIP, GLP-1, and glucagon receptors, each with EC50 values in the low nanomolar range (Jastreboff et al., NEJM 2023, PMID 37366315). Tirzepatide (Mounjaro/Zepbound) sits between the two categories as a dual GIP/GLP-1 agonist, and its trial data is useful as a middle reference point in this comparison.

The added glucagon receptor activity in retatrutide is the mechanistic reason its trial data shows larger effect sizes — glucagon receptor agonism increases hepatic energy expenditure through a pathway GLP-1 monotherapy does not touch. That mechanistic difference is real; what it means for an individual patient's outcome is a separate question this comparison addresses through trial-level data rather than mechanism alone.

Trial Design: STEP 1 vs. Retatrutide Phase 2

STEP 1 (NCT03548935) randomized 1,961 participants 2:1 to semaglutide 2.4 mg weekly or placebo, run over 68 weeks with a co-primary endpoint of percent weight change and achievement of ≥5% weight loss (Wilding et al., NEJM 2021, PMID 33567185). It is one of the largest and longest-running trials in this comparison set.

Retatrutide's Phase 2 trial (NCT04867785) randomized 338 participants across five dose arms (1 mg, 4 mg, 8 mg, 12 mg) plus placebo, with data reported at 24 and 48 weeks (Jastreboff et al., NEJM 2023, PMID 37366315). The sample size is roughly one-sixth of STEP 1's, and the trial duration is 20 weeks shorter, both facts that limit direct comparability of headline percentages.

SURMOUNT-1 for tirzepatide (NCT04184622) randomized 2,539 participants across three dose arms over 72 weeks (Jastreboff et al., NEJM 2022, PMID 35658024), the largest trial in this set and the longest follow-up window. Anyone comparing these three trials needs to hold trial size and duration constant in their reasoning, not just the topline weight-loss figure, because a 48-week Phase 2 signal and a 72-week Phase 3 confirmed result do not carry equivalent evidentiary weight.

Primary Endpoint Results: Percent Weight Loss Head-to-Head

The topline numbers, read side by side: semaglutide 2.4 mg produced a mean weight loss of approximately -14.9% at 68 weeks versus -2.4% for placebo in STEP 1. Tirzepatide's highest dose (15 mg) produced approximately -20.9% at 72 weeks in SURMOUNT-1. Retatrutide's 12 mg dose produced approximately -24.2% at 48 weeks in its Phase 2 trial.

On raw percentages, the multi-receptor compounds outperform GLP-1 monotherapy by a wide margin — roughly 6 to 9 percentage points between tirzepatide and semaglutide, and another 3 to 4 percentage points between retatrutide and tirzepatide. That gradient tracks with the number of receptors engaged, which is consistent with the mechanistic rationale.

The caveat that matters clinically: retatrutide reached its -24.2% figure at 48 weeks, twenty weeks earlier than STEP 1's 68-week endpoint. Some of that apparent advantage may narrow, hold, or widen with longer follow-up, and Phase 3 data will be the test of whether the Phase 2 magnitude replicates in a larger, longer trial. Treating a 48-week Phase 2 number as directly equivalent to a 68-week Phase 3 number overstates the certainty of the comparison.

Secondary Endpoints and Metabolic Markers

Weight change alone does not capture the full clinical picture, and all three trials reported secondary metabolic endpoints. STEP 1 reported significant reductions in waist circumference, systolic blood pressure, and fasting glucose alongside the primary weight endpoint, with a favorable shift in patient-reported physical functioning scores.

Retatrutide's Phase 2 trial reported improvements in triglycerides (up to approximately -25% at the highest dose) and blood pressure reductions in the same range as tirzepatide's SURMOUNT-1 data. Among participants with prediabetes at baseline, a meaningfully higher proportion reverted to normoglycemia on higher-dose retatrutide arms compared to lower-dose arms — a sub-population signal worth tracking into Phase 3, though the current sample size (n=338 across five arms) makes that subgroup finding preliminary.

None of the three trials used an identical secondary-endpoint battery, which is itself a limitation for anyone trying to build a clean comparison table. A clinician evaluating these compounds for a specific patient profile — for example, someone with prediabetes and hypertension rather than obesity alone — needs to check which secondary endpoints each trial actually measured rather than assuming parity across the three datasets.

Attrition and Tolerability Data

Attrition is where cross-trial comparison gets more complicated, because dropout during dose titration is where much of the difference in tolerability shows up. STEP 1 reported approximately 17% discontinuation of study drug due to adverse events in the semaglutide arm over 68 weeks.

SURMOUNT-1 reported adverse-event-driven discontinuation in the range of 6-7% for lower tirzepatide doses, rising toward the mid-teens at the highest dose. Retatrutide's Phase 2 data showed a wider spread, with discontinuation due to adverse events ranging from roughly 13% at lower doses to approximately 27% at the highest dose tested — the widest range of the three trials.

This pattern matters for anyone weighing effect size against real-world adherence: a drug with a larger mean effect size but a higher discontinuation rate produces a smaller population-level benefit than the topline percentage suggests, because participants who drop out are typically excluded from or handled differently within the completer analysis versus the intention-to-treat analysis. Reading which analysis a headline number comes from — intention-to-treat or completers-only — changes the real-world interpretation substantially.

Dose-Response and Titration Schedule Differences

Titration pace differs meaningfully across the three compounds and directly affects the tolerability numbers above. Semaglutide's STEP 1 protocol used a 16-20 week titration to reach the 2.4 mg maintenance dose, a relatively gradual schedule that likely contributed to its comparatively lower discontinuation rate.

Retatrutide's Phase 2 protocol titrated over a similar 16-20 week window depending on the assigned maximum dose, using four-week dose-increase intervals. Despite a titration pace roughly comparable to STEP 1's, retatrutide's higher-dose arms still showed higher discontinuation, suggesting the added glucagon receptor engagement itself carries a tolerability cost independent of titration speed.

A practical takeaway from this pattern: the dose-response curve for GI adverse events tracks receptor engagement breadth more closely than it tracks titration pace alone. This has implications for how future GLP-3-class trials should structure their titration protocols — slower titration helps, but it does not fully offset the added GI burden of engaging a third receptor pathway.

Safety Profile Comparison Beyond GI Effects

Gastrointestinal adverse events dominate both drug classes' safety profiles, but the comparison should also cover less common signals. Heart rate increases of 2-4 beats per minute above placebo have been reported with both GLP-1 monotherapy and multi-receptor agonists, a class effect worth monitoring in patients with existing arrhythmia history.

Gallbladder-related events (cholelithiasis, cholecystitis) occurred at higher rates in the active-treatment arms versus placebo across all three trials, consistent with rapid weight loss itself as a risk factor independent of the specific receptor mechanism. Pancreatitis was reported as a rare event across all three trials, occurring in a small number of participants without a statistically clear difference from placebo in any single trial, though the absolute numbers are small enough that meta-analytic data across multiple trials will be more informative than any single trial's rate.

None of the three compounds discussed here has long-term (multi-year) cardiovascular outcomes data published yet for the obesity indication specifically, which limits what can be said about durable safety beyond the trial window. This is an evidence gap that applies to GLP-1 and GLP-3-class compounds alike, not a distinguishing factor between them.

What the Effect-Size Difference Means for Clinical Decision-Making

An 8-to-10-percentage-point difference in mean weight loss is clinically meaningful for patients targeting metabolic thresholds tied to specific outcomes — reduction in obstructive sleep apnea severity, remission of prediabetes, or meeting BMI criteria for a planned surgical procedure. For those patients, the larger effect size associated with multi-receptor mimetics is a relevant clinical variable.

For a patient whose goal is a moderate, sustainable weight reduction with the lowest tolerability burden, the calculus looks different: semaglutide's lower discontinuation rate and its established, FDA-approved long-term safety and prescribing data represent real, non-mechanistic advantages that a larger Phase 2 effect size does not offset.

The investigational status of retatrutide is itself a decision-relevant fact, not a footnote. It is not FDA-approved and remains available only through clinical trial enrollment or, in some cases, compounding pathways with more limited manufacturing oversight than an approved drug. A treatment decision that weighs Phase 2 effect-size data against an FDA-approved drug's Phase 3 and post-market safety record is comparing two different evidentiary categories, and that distinction should be explicit in any clinical discussion with a patient.

Common Misconceptions When Comparing These Trials

The most common error is treating cross-trial percentage comparisons as equivalent to a head-to-head randomized result. No trial has randomized participants directly between a GLP-1 monotherapy and a triple-receptor mimetic within the same protocol; every comparison in this article, and in the broader literature, is indirect. That does not make the comparison worthless, but it does mean the effect-size gap should be treated as a hypothesis-generating signal rather than a confirmed head-to-head result until a direct trial is run.

A second common error is quoting a completer-analysis percentage as if it were the intention-to-treat result, which inflates the apparent effect size by excluding the participants most likely to have had a poor response or poor tolerability. Checking which analysis a cited number comes from takes thirty seconds and materially changes the interpretation.

A third error is assuming trial duration doesn't matter when comparing percentages — a 48-week result and a 72-week result are not interchangeable, and weight loss trajectories for these compounds are not necessarily linear across time. The practical next step for anyone evaluating these compounds, clinician or researcher, is to pull the primary publication for each trial being compared, confirm the analysis population (ITT versus completers) and the follow-up duration, and only then compare the percentage figures side by side.

This article summarizes research and does not constitute medical advice. Consult a licensed clinician for diagnosis, treatment, or any decisions about medications or supplements.

Frequently asked questions

Is a GLP-3 mimetic more effective than a GLP-1 drug for weight loss?

Cross-trial data suggest larger mean weight loss with triple-receptor mimetics like retatrutide (approximately -24% at 48 weeks) than GLP-1 monotherapy like semaglutide (approximately -14.9% at 68 weeks). No head-to-head RCT exists, so this comparison is indirect and trial durations differ, limiting a direct effectiveness claim.

What is the difference between GLP-1 and GLP-3 mimetic peptides?

GLP-1 mimetics like semaglutide act on a single receptor (GLP-1R). The peptides informally called 'GLP-3' mimetics, such as retatrutide, engage GIP, GLP-1, and glucagon receptors simultaneously, which is the mechanistic basis for their larger observed effect size in trial data.

Are GLP-3 mimetic peptides FDA-approved?

As of current published data, triple-receptor agonists such as retatrutide remain investigational and have not received FDA approval for any indication. They are studied under active clinical trial protocols, unlike semaglutide and tirzepatide, which hold FDA approval for chronic weight management.

Why can't GLP-1 and GLP-3 trial results be compared directly?

The trials used different study populations, trial durations (48 versus 68 weeks), endpoint timing, and dosing schedules. Cross-trial comparison is an established but imperfect method; only a randomized head-to-head trial with a shared control arm removes these confounders.

What are the most common side effects when comparing these peptide classes?

Both classes show gastrointestinal adverse events as the dominant tolerability issue — nausea, vomiting, diarrhea, and constipation — with rates generally rising alongside dose and the number of receptors engaged. Discontinuation due to adverse events runs from roughly 15% to 27% across published trials.

Related references on this site

guide

GLP-3 / Retatrutide: The Triple Agonist Explained

Reference guide on this site.

View →
guide

Peptides 101: A Clinician's Reference

Reference guide on this site.

View →
guide

Handling, Reconstitution, and Storage of Research Peptides

Reference guide on this site.

View →