- Same-day UK dispatch on orders before 3pm
- Certificate status shown on every product
- Shop all research peptides
- Learn: guides and research tools
Incretin & Amylin Analogues
What Is Semaglutide? Structure and Chemistry
For research use only. Not for human or veterinary use. Sold strictly for in-vitro laboratory research; not for diagnosis or treatment.
British Peptide LabsPublished Updated
Key facts
- CAS number
- 910463-68-2
- Molecular formula
- C187H291N45O59
- Average molecular weight
- 4113.58 g/mol
- Monoisotopic mass
- 4111.12 Da
- Residues
- 31, numbered 7–37 by convention
- Sequence
- H-His-Aib-EGTFTSDVSSYLEGQAAKEFIAWLVRGRG-OH
- Substitutions
- Aib8 and Arg34
- Modification
- Lys26 side chain acylated with a C18 fatty diacid through γGlu and two OEG spacers
- C-terminus
- Free carboxylic acid (Gly37)
- PubChem CID
- 56843331
Semaglutide is a synthetic lipidated peptide of 31 amino acid residues, with the molecular formula C187H291N45O59 and an average molecular weight of 4113.58 g/mol. Its chain differs from its parent sequence at two positions: α-aminoisobutyric acid (Aib) at position 8 and arginine at position 34. The lysine at position 26 carries an 18-carbon fatty diacid attached through a γ-glutamic acid and two short ethylene-glycol spacers. At the molecular level, semaglutide is a GLP-1 receptor agonist peptide.
The 31-residue sequence
Semaglutide is a linear chain with a free N-terminal amine on histidine and a free C-terminal carboxylic acid on glycine. The C-terminus is not amidated. The table sets the sequence out in three-letter code, with both the chain position and the conventional position number that structural papers use.
| Chain positions | Conventional numbering | Residues |
|---|---|---|
| 1–10 | 7–16 | His Aib Glu Gly Thr Phe Thr Ser Asp Val |
| 11–20 | 17–26 | Ser Ser Tyr Leu Glu Gly Gln Ala Ala Lys |
| 21–31 | 27–37 | Glu Phe Ile Ala Trp Leu Val Arg Gly Arg Gly |
In one-letter code the chain reads H-His-Aib-EGTFTSDVSSYLEGQAAKEFIAWLVRGRG-OH. His and Aib are written in three-letter form, and H- and -OH mark the free amine and the free acid. They are terminal groups, not residues, so the string holds 31 residues: His, Aib and 29 one-letter codes. Writing His out keeps it from being read as the H- of the free amine. The only lysine is at chain position 20, conventional position 26, and it carries the side chain. Peptide sequence notation explains how one-letter and three-letter codes and terminal groups such as H- and -OH are written.
Why the positions run from 7 to 37
Semaglutide's backbone is an analogue of the native peptide ligand of the GLP-1 receptor, a 31-residue sequence. Sequence databases and structural papers number that sequence from 7 to 37, because it corresponds to residues 7–37 of a longer 37-residue form of the same peptide. Semaglutide inherits the convention, which is why its defining positions are written Aib8, Lys26 and Arg34 rather than 2, 20 and 28.
| Conventional position | Chain position | Parent residue | In semaglutide |
|---|---|---|---|
| 7 | 1 | His | His, free N-terminal amine |
| 8 | 2 | Ala | Aib |
| 26 | 20 | Lys | Lys, side chain acylated |
| 34 | 28 | Lys | Arg |
| 37 | 31 | Gly | Gly, free C-terminal acid |
Every other position matches the parent sequence. That makes semaglutide a close structural analogue: two substitutions and one side-chain modification across 31 residues.
Two substitutions: Aib8 and Arg34
Aib8. α-Aminoisobutyric acid, also called 2-methylalanine, is a non-coded amino acid: alanine with a second methyl group on the α-carbon. The two identical methyl groups make it achiral and leave it without an α-hydrogen. The gem-dimethyl α-carbon restricts the backbone torsion angles the residue can adopt, and Aib is a well-known helix-favouring residue in peptide chemistry. At position 8 it replaces alanine, so the substitution adds exactly one methyl group, next to the N-terminal histidine.
Arg34. At position 34 the parent lysine is replaced by arginine. Both side chains are basic and positively charged at neutral pH, but arginine's guanidinium group remains protonated across a wider pH range than lysine's primary amine. The substitution has a structural consequence: Lys26 becomes the only lysine in the chain. Before the side chain is attached, the peptide therefore has just two primary amines, the N-terminal α-amine of His7 and the ε-amine of Lys26, and only the Lys26 amine carries the side chain in semaglutide.
The lipidated side chain on lysine 26
The ε-amino group of Lys26 is acylated by a linker of three units, and the linker ends in an 18-carbon fatty diacid. Every junction is an amide bond.
| Unit, from the lysine outward | Chemical identity | Structural role |
|---|---|---|
| First spacer | 8-amino-3,6-dioxaoctanoic acid (OEG, also written AEEA) | Short, flexible ethylene-glycol chain |
| Second spacer | 8-amino-3,6-dioxaoctanoic acid | Identical spacer that extends the linker |
| Linker acid | L-γ-glutamic acid (γGlu), bonded through its side-chain carboxyl | Leaves a free α-carboxylic acid on the linker |
| Lipid | Octadecanedioic acid (C18 fatty diacid), amidated at one end | Sixteen methylene groups ending in a free carboxylic acid |
PubChem's systematic synonym spells the same thing out residue by residue. It describes Lys26 as N6-[N-(17-carboxy-1-oxoheptadecyl)-L-γ-glutamyl-2-[2-(2-aminoethoxy)ethoxy]acetyl-2-[2-(2-aminoethoxy)ethoxy]acetyl]-L-lysyl. "17-Carboxy-1-oxoheptadecyl" is the C18 diacid attached through one of its carboxyl groups. Each "2-[2-(2-aminoethoxy)ethoxy]acetyl" is one OEG spacer.
With the side chain in place, semaglutide is an amphiphilic molecule: a peptide chain carrying a long hydrocarbon tail that ends in a carboxylic acid.
Formula, mass and expected ions
The formula C187H291N45O59 describes the complete neutral molecule, including the linker and diacid. Building it up from the 31 residues, the two spacers, the γGlu unit and the diacid reproduces the formula PubChem lists. PubChem rounds the average molecular weight to 4114 g/mol and gives the monoisotopic mass as 4111.1154 Da.
The calculated average mass shifts slightly with the atomic-weight table used: 4113.58 g/mol with the older standard atomic weights (C 12.0107, H 1.00794, N 14.0067, O 15.9994) and 4113.64 g/mol with the current IUPAC abridged values (C 12.011, H 1.008, N 14.007, O 15.999). The difference is under 0.01% and reflects the reference table, not the molecule.
In electrospray ionisation mass spectrometry, a peptide of just over 4.1 kDa appears as a series of multiply protonated ions. Calculated from the average mass, the main charge states fall at:
| Ion | Calculated m/z |
|---|---|
| [M+3H]³⁺ | 1372.20 |
| [M+4H]⁴⁺ | 1029.40 |
| [M+5H]⁵⁺ | 823.72 |
| [M+6H]⁶⁺ | 686.60 |
Deconvoluting that series gives a single neutral mass to compare with the expected value. Mass spectrometry for peptide identity explains how charge states and isotope envelopes are read.
Receptor binding at the molecular level
Semaglutide is a GLP-1 receptor agonist peptide. The GLP-1 receptor is a class B1 G protein-coupled receptor (GPCR). Receptors in this family have a large N-terminal extracellular domain (ECD) in addition to the seven-helix transmembrane domain (TMD), and they bind their peptide ligands through both.
Two structures show semaglutide in that binding site:
- PDB 4ZGM (2015, X-ray, 1.8 Å). The semaglutide peptide backbone, without its side chain, is bound to the isolated extracellular domain. Of its 31 residues, 28 are modelled, and they form a continuous α-helix from residue 13 to residue 33.
- PDB 7KI0 (2021, cryo-EM, 2.5 Å). Full semaglutide is bound to the complete receptor in its active state, coupled to the heterotrimeric Gs protein. Zhang et al. report peptide interactions similar to those of the receptor's native ligand, with different motions within the receptor and the bound peptide.
The lipidated side chain is only partly visible. In the deposited 7KI0 model, the two OEG spacers are built onto the Lys26 side-chain nitrogen, while the γGlu unit and the diacid are not modelled.
How semaglutide is characterised analytically
Identity and purity of a lipidated peptide are normally established with two complementary methods.
- Reversed-phase HPLC separates semaglutide from related substances and reports purity as the main peak's share of total peak area, with UV detection at 214–220 nm. The C18 diacid gives the molecule strong retention on reversed-phase columns compared with a non-lipidated chain of similar length.
- Mass spectrometry confirms identity by matching the deconvoluted mass to the expected mass above.
Related substances from solid-phase synthesis include deletion sequences, truncated chains and incompletely deprotected species. For a lipidated peptide they also include chains with an incomplete linker or no side chain. A certificate of analysis may also report properties of the solid, such as water content and counter-ion content. Because the solid contains both, a vial's net peptide content is lower than its gross powder mass.
A certificate of analysis for semaglutide is published in the COA Library. Certificate status is shown on every product page, and a batch certificate is available on request. How to read a peptide certificate of analysis walks through each field.
Form and storage of the sealed vial
Semaglutide is supplied as a white to off-white lyophilised powder in a sealed vial, in 15 mg and 20 mg presentations. Identifiers, the purity specification and the certificate status are listed on the semaglutide product page.
Store the sealed vial at 2–8 °C for short-term storage, or at −20 °C and below for long term. Protect it from light and avoid repeated freeze-thaw cycles. Handle the material as a laboratory chemical, following your institution's chemical-safety procedures. Storing lyophilised peptides explains why a freeze-dried solid is kept cold, dark and dry.
In our catalogue, semaglutide sits in the Incretin & Amylin Analogues research area, which groups lipidated peptide analogues by molecular class. Terms such as lyophilisation, deconvolution and monoisotopic mass are defined in the glossary.
Frequently asked questions
Semaglutide is a 31-residue synthetic lipidated peptide with the formula C187H291N45O59 and an average molecular weight of 4113.58 g/mol. It carries α-aminoisobutyric acid at position 8 and arginine at position 34, and an 18-carbon fatty diacid is attached to the side chain of lysine 26.
Semaglutide is a linear 31-residue chain, H-His-Aib-EGTFTSDVSSYLEGQAAKEFIAWLVRGRG-OH, numbered 7 to 37 by convention. The ε-amine of Lys26 is acylated by two 8-amino-3,6-dioxaoctanoic acid spacers, a γ-glutamic acid unit and octadecanedioic acid, in that order from the lysine outward. The H- and -OH are the free N-terminal amine and the free C-terminal carboxylic acid, not residues.
Semaglutide is a GLP-1 receptor agonist peptide. The receptor is a class B1 G protein-coupled receptor with a large extracellular domain and a seven-helix transmembrane domain. An X-ray structure from 2015 shows the helical peptide backbone bound to the extracellular domain, and a cryo-EM structure from 2021 shows full semaglutide bound to the complete receptor.
Purity is measured by reversed-phase HPLC with UV detection and reported as the main peak's share of total peak area. Identity is confirmed by mass spectrometry against the expected mass of 4113.58 g/mol (average) or 4111.12 Da (monoisotopic). A certificate of analysis for semaglutide is published in the COA Library.
Store the sealed vial at 2–8 °C for short-term storage, or at −20 °C and below for long term. Protect it from light and avoid repeated freeze-thaw cycles.
References
- PubChem: Semaglutide, CID 56843331 (formula, computed masses, systematic sequence name) (pubchem.ncbi.nlm.nih.gov)
- Zhang X. et al. (2021), Cell Reports 36, 109374 (cryo-EM structure of semaglutide bound to its receptor–Gs complex) (doi.org)
- RCSB Protein Data Bank: entry 7KI0, cryo-EM structure with bound semaglutide (rcsb.org)