Skip to content

Amino Acid Code Converter

Convert between one-letter and three-letter amino acid codes. Includes a complete reference table with chemical properties for all 20 standard amino acids.


Converter

Separate multiple codes with spaces or newlines. Mixed case is accepted.

Conversion Result


What This Calculator Does

Paste amino acid codes in either notation and convert between the one-letter and three-letter forms used across peptide work. A three-letter list such as Ala Gly Lys collapses to the one-letter string AGK; a one-letter string such as AGK expands to the hyphen-separated three-letter form Ala-Gly-Lys. The converter covers the 20 standard proteinogenic amino acids and maps exact code matches only — there is no guessing and no chemical interpretation.

Both directions run entirely in the browser; nothing is uploaded, stored, or logged. The sections below set out the matching rules, walk through a conversion, and note where the simple lookup model stops being enough.

How the Conversion Works

The converter is a lookup table, not an algorithm: each residue symbol has exactly one counterpart in the other notation, and the change happens one residue at a time.

Three-letter → one-letter. The input is split on spaces, tabs, newlines, and commas. Each token is lowercased and matched against the standard three-letter codes (Ala, Arg, Asn … Val). Matched tokens are replaced by their one-letter symbols, which are concatenated in order with no separators:

\[ \text{Ala Gly Lys} \;\longrightarrow\; \text{A} + \text{G} + \text{K} = \text{AGK} \]

One-letter → three-letter. The input is scanned for the 20 standard one-letter symbols; any other character — spaces, digits, punctuation, non-standard letters — is ignored. Each symbol is replaced by its three-letter code, and consecutive residues are joined with hyphens:

\[ \text{A} + \text{G} + \text{K} \;\longrightarrow\; \text{Ala-Gly-Lys} \]

Matching is case-insensitive in both directions, so ALA, ala, and Ala all resolve to the same residue. Tokens that do not match a standard three-letter code are returned unchanged (uppercased) rather than dropped — helpful for spotting a typo, but it means invalid input produces invalid-looking output rather than an error.

Worked Example

Convert the tripeptide Gly-His-Lys (GHK, the copper-binding peptide) between notations.

Step 1 — Start from the three-letter form: Gly His Lys.

Step 2 — Look up each token:

Token Residue One-letter
Gly Glycine G
His Histidine H
Lys Lysine K

Step 3 — Concatenate in sequence order: G + H + K = GHK.

Step 4 — Convert back. The reverse conversion expands each one-letter symbol to its three-letter code and joins with hyphens: Gly-His-Lys.

Result: the same tripeptide in both notations — GHK and Gly-His-Lys. The round trip is lossless: a notation change, not a chemical one.


Amino Acid Reference Table

One-Letter Three-Letter Name Side Chain Class MW (Da) pKa (R group) Hydropathy
A Ala Alanine Nonpolar, Aliphatic 71.08 +1.800
C Cys Cysteine Polar, Sulfur 103.14 8.37 +2.500
D Asp Aspartic Acid Acidic 115.09 3.90 -3.500
E Glu Glutamic Acid Acidic 129.12 4.07 -3.500
F Phe Phenylalanine Aromatic 147.18 +2.800
G Gly Glycine Nonpolar, Aliphatic 57.05 -0.400
H His Histidine Basic, Aromatic 137.14 6.00 -3.200
I Ile Isoleucine Nonpolar, Aliphatic 113.16 +4.500
K Lys Lysine Basic 128.17 10.50 -3.900
L Leu Leucine Nonpolar, Aliphatic 113.16 +3.800
M Met Methionine Nonpolar, Sulfur 131.19 +1.900
N Asn Asparagine Polar, Amide 114.10 -3.500
P Pro Proline Cyclic (Imino) 97.12 -1.600
Q Gln Glutamine Polar, Amide 128.13 -3.500
R Arg Arginine Basic 156.19 12.48 -4.500
S Ser Serine Polar, Hydroxyl 87.08 -0.800
T Thr Threonine Polar, Hydroxyl 101.10 -0.700
V Val Valine Nonpolar, Aliphatic 99.13 +4.200
W Trp Tryptophan Aromatic 186.21 -0.900
Y Tyr Tyrosine Aromatic 163.18 10.10 -1.300

Amino Acid Classification Guide

Category Amino Acids Properties
Nonpolar Aliphatic Gly, Ala, Val, Leu, Ile, Pro Hydrophobic, buried in protein cores
Aromatic Phe, Tyr, Trp Absorb UV at 280nm (important for spectroscopy)
Polar Uncharged Ser, Thr, Cys, Asn, Gln Hydrophilic, surface-exposed
Acidic (Negative) Asp, Glu Negatively charged at pH 7; pKa ~4
Basic (Positive) Lys, Arg, His Positively charged at pH 7 (except His, pKa 6)
Special Cys (disulfide), Pro (rigid), Gly (flexible) Unique structural roles

Amino Acid Properties at a Glance

  • Hydrophobicity: Most hydrophobic = Ile, Val, Leu; Most hydrophilic = Arg, Asp, Lys
  • Molecular weight: Smallest = Gly (57 Da); Largest = Trp (186 Da)
  • Charge at pH 7: Positive = Arg, Lys; Negative = Asp, Glu; Neutral = all others

Common Usage Patterns

  • In peptide sequences: single-letter codes are standard for sequences over 5 AA
  • In publications: three-letter codes for individual residues in text
  • In databases: single-letter codes for sequence storage
  • In patents: both formats used interchangeably

Frequently Asked Questions

**Why are there both one-letter and three-letter codes?**

The three-letter system was developed first and is more descriptive, making it useful in written text. The one-letter system was introduced later to support computer sequence analysis and compact representation of long sequences. Both standards are maintained by IUPAC.

**How do I remember the one-letter codes?**

A few mnemonic aids: A = Alanine (first letter), C = Cysteine (first letter), G = Glycine (first letter), H = Histidine (first letter), P = Proline (first letter), S = Serine (first letter), V = Valine (first letter). For others: F = Phenylalanine (ph → F), R = aRginine (letter R in name), Y = tYrosine, K = lysine (sounds like K), D = asparDic acid, E = glutamEic acid, M = M (contains M), T = T (contains T), W = tW (tryptophan, double ring), I = I (contains I), L = L (contains L), N = N (contains N), Q = Q (sounds like "cue" → glutamine).

**What about non-standard amino acids?**

Non-standard amino acids (e.g., selenocysteine Sec/U, pyrrolysine Pyl/O, hydroxyproline Hyp) do not have universally accepted one-letter codes. In sequences, they are often represented by special letters (U for selenocysteine, O for pyrrolysine) or written out in full. The converter above handles only the 20 standard amino acids.

**Why are I and L different if they have the same mass?**

Isoleucine (I) and Leucine (L) are structural isomers — they share the same molecular formula (C₆H₁₃NO₂) and therefore the same molecular weight (113.16 Da). However, they differ in the arrangement of their side chains: Leu has an unbranched isobutyl side chain, while Ile has a branched sec-butyl side chain. This structural difference gives them distinct biochemical properties and roles in proteins.

**What does the U stand for in some sequences?**

The letter U stands for selenocysteine (Sec), the 21st proteinogenic amino acid. It is a cysteine analog where sulfur is replaced by selenium. Selenocysteine is encoded by a special UGA codon (normally a stop codon) in the presence of a selenocysteine insertion sequence (SECIS) element. It is found in several selenoproteins important for antioxidant defense.


Assumptions and Rounding

  • Standard residue set only. Only the 20 proteinogenic amino acids are mapped; selenocysteine (U), pyrrolysine (O), hydroxyproline, D-amino acids, and other non-standard residues have no entry — see Limitations.
  • Exact matches only. Three-letter matching requires the complete code, case-insensitively. Abbreviations, truncated spellings, and full residue names are not recognized.
  • Separator handling. Three-letter input may be separated by spaces, tabs, newlines, or commas; one-letter input may be written continuously or spaced, since non-standard characters are stripped before mapping.
  • Output formatting. One-letter output is concatenated without separators; three-letter output is hyphen-joined to match conventional sequence notation.
  • No rounding. The conversion is a one-to-one symbol substitution — nothing is calculated, averaged, or rounded, and output order always follows input order.

Input Definitions

Input What it means Units Allowed values
Conversion Direction Selects which of the two conversion paths runs "Three-Letter → One-Letter"; "One-Letter → Three-Letter"
Amino Acid Codes Residue codes to convert; mixed case accepted The 20 standard codes in either notation — e.g. Ala Gly Lys, Ala,Gly,Lys, or AGK

The result panel updates when the Convert button is pressed, and the Clear button empties the input and hides the panel. No inputs are stored between visits.

Output Interpretation

There is one output — a converted string — so interpretation is mostly about confirming you received the format you expected:

Output How to read it
Conversion Result The input rewritten in the other notation. One-letter results are continuous (AGK); three-letter results are hyphen-joined (Ala-Gly-Lys). Token order is preserved.
Text that looks unchanged Unmatched tokens pass through uppercased — if something did not convert, cross-check the spelling against the reference table above.
"No valid amino acid codes found" No standard residue symbols were detected in the input; check for stray characters.

Limitations

  • Standard 20 only. Non-standard residues and modified forms — D-amino acids, selenocysteine (U), pyrrolysine (O), hydroxyproline, norleucine — pass through untranslated.
  • No validation. The converter does not check that a sequence exists, is chemically plausible, or matches any reference molecule. It is a notation translator, not a sequence validator.
  • Text-box input only. There is no file upload, batch conversion, or export.
  • No chemical output. Nothing beyond the notation is produced — no mass, charge, or property values. Those live in the site's Molecular Weight Calculator and Peptide Properties Calculator.
  • Verify before use. For publication, database submission, or order forms, confirm converted sequences against your source records; a typo in the input is reproduced faithfully in the output.

The Author's Take

Position — in my view, the two notation systems are best treated as different jobs: one-letter codes for the machine, three-letter codes for the reader, and labs that standardize that split make fewer sequence mix-ups.

Reasoning. One-letter strings are compact and easy to diff, which makes them ideal for databases, alignment tools, and order systems. Three-letter codes are self-checking to the human eye — a reader cannot misread "Glu" for "Gln" the way two similar single characters can blur, and copy errors are easier to catch in review. In my experience the errors that reach the bench are hand-copying errors, not lookup errors. Store in one-letter, review in three-letter, and let a checked table do the conversion rather than memory.

Disclosure. This is the author's working opinion from laboratory practice, not a verified fact; follow your institution's documentation conventions for sequences.

The notation is only the entry point — these references pick up where the conversion ends: