- Same-day UK dispatch on orders before 3pm
- Certificate status shown on every product
- Shop all research peptides
- Learn: guides and research tools
Fundamentals
Peptide Sequence Notation: Reading an Amino Acid Sequence
For research use only. Not for human or veterinary use. Sold strictly for in-vitro laboratory research; not for diagnosis or treatment.
British Peptide LabsPublished Updated
Key facts
- Direction
- N-terminus (left) to C-terminus (right)
- Hyphen
- Stands for the peptide bond between two residues
- Default configuration
- L, unless a D- prefix is written
- Standard
- IUPAC-IUB JCBN Recommendations 1983
An amino acid sequence lists a peptide's residues in order from the N-terminus, the end with the free amino group, written on the left, to the C-terminus, the end with the free carboxyl group, written on the right. Each residue is shown by a three-letter symbol such as Gly or a one-letter code such as G, hyphens between three-letter symbols stand for the peptide bonds, and prefixes and suffixes such as Ac-, -NH2 and D- record the modifications that the letters alone cannot show.
The rules below follow the IUPAC-IUB recommendations on amino acid and peptide symbolism, illustrated with sequences from the catalogue. For the underlying chemistry, start with what peptides are.
Direction: N-terminus to C-terminus
A sequence is always read from the N-terminal residue to the C-terminal residue, and residues are numbered from 1 at the N-terminal end. BPC-157 can be written three equivalent ways:
Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-ValH-Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val-OHGEPPPGKPADDAGLV
Glycine is residue 1, at the N-terminus, and valine is residue 15, at the C-terminus. Reversing the order describes a different compound, just as glycylalanine (Gly-Ala) and alanylglycine (Ala-Gly) are different dipeptides.
In the three-letter system each hyphen has a precise meaning. A hyphen to the right of a symbol removes the OH of that residue's carboxyl group, and a hyphen to the left removes one H from its amino group, so the hyphen itself represents the peptide bond. That is why Gly-Glu is a complete dipeptide, while -Gly-Glu- is a fragment inside a longer chain. Writing H- and -OH at the ends states explicitly that both termini are unmodified.
Three-letter and one-letter codes
The twenty common amino acids each have a three-letter symbol and a one-letter code:
| Amino acid | Three-letter | One-letter |
|---|---|---|
| Alanine | Ala | A |
| Arginine | Arg | R |
| Asparagine | Asn | N |
| Aspartic acid | Asp | D |
| Cysteine | Cys | C |
| Glutamic acid | Glu | E |
| Glutamine | Gln | Q |
| Glycine | Gly | G |
| Histidine | His | H |
| Isoleucine | Ile | I |
| Leucine | Leu | L |
| Lysine | Lys | K |
| Methionine | Met | M |
| Phenylalanine | Phe | F |
| Proline | Pro | P |
| Serine | Ser | S |
| Threonine | Thr | T |
| Tryptophan | Trp | W |
| Tyrosine | Tyr | Y |
| Valine | Val | V |
Six one-letter codes are simply initials (C, H, I, M, S, V). Where several amino acids share an initial, the letter went to the simplest and most frequent one (A, G, L, P, T), and the rest were assigned by association: F for phenylalanine, R for arginine, W for the bulky double ring of tryptophan. Three more codes cover uncertainty: B (Asx, aspartic acid or asparagine), Z (Glx, glutamic acid or glutamine) and X (Xaa, unknown or other); U is assigned to selenocysteine (Sec).
Converting between systems is mechanical for unmodified chains. Selank, Thr-Lys-Pro-Arg-Pro-Gly-Pro, becomes TKPRPGP; MOTS-c, Met-Arg-Trp-Gln-Glu-Met-Gly-Tyr-Ile-Phe-Tyr-Pro-Arg-Lys-Leu-Arg, becomes MRWQEMGYIFYPRKLR.
IUPAC recommends the three-letter system for ordinary text and keeps the one-letter code for long sequences, tables and alignments, because the one-letter code is harder to read for anyone not fluent in it. The three-letter form also has room for prefixes and suffixes that record modifications.
Terminal modifications: Ac- and -NH2
Groups attached to either end are written as substituents on the terminal symbol:
| Notation | Meaning | Change in formula | Average mass change |
|---|---|---|---|
H- … -OH | Free amine and free acid (explicit) | None | None |
Ac- | N-terminal acetyl group | + C2H2O | + 42.04 g/mol |
-NH2 | C-terminal carboxamide | + NH, − O | − 0.98 g/mol |
Catalogue examples show both. SS-31 is D-Arg-Dmt-Lys-Phe-NH2: its C-terminal phenylalanine ends in an amide, not a free acid, and its formula, C32H49N9O5, only adds up with that amide included. TB-500 carries an N-terminal acetyl group on its first serine. Other acyl groups follow the same pattern; tesamorelin, for instance, is described in the catalogue as bearing an N-terminal trans-3-hexenoyl group.
The position of the substituent matters. Ala-NH2 means alaninamide, a C-terminal amide, whereas Ala(NH2) would mean an extra amino group on the side chain. The IUPAC text warns about exactly this confusion.
Stereochemistry and non-standard residues
Every symbol denotes the L configuration unless a D- prefix is written in front of it. Ipamorelin, Aib-His-D-2-Nal-D-Phe-Lys-NH2, contains two D-residues and an achiral one, and shows several less common symbols:
| Symbol | Residue | Where it appears |
|---|---|---|
| Aib | α-Aminoisobutyric acid (2-methylalanine), achiral | Ipamorelin, position 1 |
| 2-Nal | 3-(2-Naphthyl)alanine | Ipamorelin, position 3, as D-2-Nal |
| Dmt | 2,6-Dimethyltyrosine | SS-31, position 2 |
| Nle | Norleucine, a straight-chain isomer of leucine; IUPAC now prefers the symbol Ahx, but Nle remains in wide use | MT2, position 1 |
Symbols like these are not part of the standard set, and IUPAC asks for them to be defined in every document that uses them. Record formats do not always agree: some databases spell Dmt out as a substituted tyrosine, Tyr(2,6-diMe), and PubChem writes ipamorelin as H-Aib-His-D-2Nal-D-Phe-Lys-NH2.
The one-letter code struggles here. PubChem's one-letter form of ipamorelin is XHXFK, which loses the Aib, the naphthylalanine, both D configurations and the C-terminal amide. For modified peptides, the three-letter form is the one to trust.
Cyclisation, lactam bridges and disulfides
Rings are the hardest feature to write in one line, and there are two families.
A homodetic cyclic peptide closes its ring through backbone peptide bonds only. It is written with the prefix cyclo and the sequence in brackets, with hyphens at both ends inside the bracket to show that the last residue bonds back to the first.
A heterodetic cyclic peptide closes its ring through at least one other bond, such as an isopeptide, disulfide or ester bond, usually between side chains. IUPAC draws these on two lines, with a line joining the linked residues; in running text, catalogues and databases use one-line conventions instead. Two common types:
- Lactam bridge. An amide formed between an acidic side chain (Asp or Glu) and a basic one (Lys or Orn). MT2 is an example: the catalogue writes it
Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2, while PubChem writes the same moleculeAc-Nle-Asp(1)-His-D-Phe-Arg-Trp-Lys(1)-NH2, where the matching (1) labels mark the two residues joined through their side chains. - Disulfide bridge. An S–S bond between two cysteine side chains, written with matching labels on the two Cys residues or stated separately, for example as a Cys2–Cys7 disulfide. Forming it removes two hydrogen atoms, so the mass falls by 2.02 g/mol compared with the free thiols.
Side-chain substituents in general go in parentheses straight after the residue symbol, which is why a lipid chain on a lysine is written in the form Lys(substituent). Metal complexes are written after the chain, as in the catalogue's Gly-His-Lys · Cu(II) for GHK-Cu.
Worked example: from sequence to formula
Because each residue has a fixed formula, a correctly written sequence determines the molecular formula. MT2 shows each modification changing it in a predictable way:
| Step | Structure | Formula | Average mass (g/mol) |
|---|---|---|---|
| Linear chain | H-Nle-Asp-His-D-Phe-Arg-Trp-Lys-OH | C48H68N14O10 | 1001.16 |
| Add N-acetyl | Ac-…-OH | C50H70N14O11 | 1043.20 |
| Add C-terminal amide | Ac-…-NH2 | C50H71N15O10 | 1042.21 |
| Close the Asp–Lys lactam | Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2 | C50H69N15O9 | 1024.20 |
The final line matches the catalogue formula, C50H69N15O9, and PubChem's monoisotopic mass of 1023.54 Da. The average mass is the figure used for weighing and molar calculations, such as those in the molarity calculator; the monoisotopic mass is the one compared with a high-resolution mass spectrum, as explained in mass spectrometry and peptide identity.
Databases also store sequences in machine-readable formats such as HELM, which lists every monomer, including modified ones, and records ring-forming bonds as separate connections. For reading a certificate of analysis, the three-letter form with explicit termini remains the clearest; see how to read a peptide certificate of analysis. Individual terms are defined in the glossary.
Frequently asked questions
From the N-terminus to the C-terminus. The residue with the free amino group is written on the left and the residue with the free carboxyl group on the right, in both the three-letter and the one-letter systems, and residues are numbered from 1 at the N-terminal end.
It shows a C-terminal amide: the carboxyl group of the last residue has been converted to a carboxamide, CONH2. The formula gains NH and loses O compared with the free acid, so the average molecular mass is about 0.98 g/mol lower. SS-31, written D-Arg-Dmt-Lys-Phe-NH2, is an example.
It marks a residue with the D configuration at its α-carbon. Amino acid symbols denote the L configuration unless a D prefix is written, so D-Phe is D-phenylalanine and Phe on its own is L-phenylalanine.
They name the same residues. Three-letter symbols such as Gly and Trp are easier to read and can carry prefixes and suffixes for modifications, so they are preferred in text. One-letter codes such as G and W are compact and suit long sequences and alignments, but they have no standard way to show most modifications.
A ring made only of backbone peptide bonds is written, following IUPAC, with the prefix cyclo and the sequence in brackets. For a ring closed through side chains, such as a lactam or a disulfide bridge, IUPAC uses a two-line drawing with a line joining the linked residues. In one line, catalogues and databases either bracket the ring residues after cyclo or give the two linked residues the same number in brackets, for example Asp(1) and Lys(1).
References
- IUPAC-IUB JCBN. Nomenclature and Symbolism for Amino Acids and Peptides (Recommendations 1983), 3AA-14 to 3AA-16: three-letter symbols, configuration and the meaning of the hyphen (iupac.qmul.ac.uk)
- IUPAC-IUB JCBN. Nomenclature and Symbolism for Amino Acids and Peptides (Recommendations 1983), 3AA-17: substituted amino acids (iupac.qmul.ac.uk)
- IUPAC-IUB JCBN. Nomenclature and Symbolism for Amino Acids and Peptides (Recommendations 1983), 3AA-18 and 3AA-19: substituents and peptide symbolism (iupac.qmul.ac.uk)
- IUPAC-IUB JCBN. Nomenclature and Symbolism for Amino Acids and Peptides (Recommendations 1983), 3AA-20 and 3AA-21: the one-letter system (iupac.qmul.ac.uk)
- Zhang T, Li H, Xi H, Stanton RV, Rotstein SH. HELM: a hierarchical notation language for complex biomolecule structure representation. J. Chem. Inf. Model. 2012, 52, 2796–2806 (doi.org)
- PubChem CID 9941957: BPC-157 sequence notations and formula (pubchem.ncbi.nlm.nih.gov)
- PubChem CID 92432: MT2 sequence notation, formula and monoisotopic mass (pubchem.ncbi.nlm.nih.gov)
- PubChem CID 9831659: Ipamorelin sequence notation and formula (pubchem.ncbi.nlm.nih.gov)