BioDeviceHub

IUPAC Ambiguity Code Tool

Decode IUPAC nucleotide ambiguity codes (R, Y, S, W, K, M, B, D, H, V, N), count the ambiguous positions and the distinct sequences they stand for, and list them.

Formula

Nsequences=∏ibi,bi=bases allowed at position iN_{\text{sequences}} = \prod_i b_i,\qquad b_i = \text{bases allowed at position } i
A,C,G,T=1;  R,Y,S,W,K,M=2;  B,D,H,V=3;  N=4\mathrm{A,C,G,T}=1;\ \ \mathrm{R,Y,S,W,K,M}=2;\ \ \mathrm{B,D,H,V}=3;\ \ \mathrm{N}=4
R, YR,\ Y
purine (A or G), pyrimidine (C or T)
S, WS,\ W
strong (C or G), weak (A or T)
K, MK,\ M
keto (G or T), amino (A or C)
B, D, H, VB,\ D,\ H,\ V
not A, not C, not G, not T
NN
any base

How it works

Sequencing software, consensus sequences and degenerate primers use one letter to stand for several possible bases. The standard codes were proposed by the IUPAC and IUBMB nomenclature committees and published by Cornish-Bowden in 1985.

The number of distinct sequences is the product of the number of possibilities at each position. It grows fast: ten N positions stand for over a million sequences. For a degenerate primer this matters because the concentration of any single sequence in the pool falls in proportion.

Worked example

The sequence ATGRYNCGTWSA.

  1. R = 2 bases, Y = 2, N = 4, W = 2, S = 2; A, T, G, C, G, T, A = 1 each.
  2. 2 × 2 × 4 × 2 × 2 = 64.

Five ambiguous positions of 12 and 64 distinct sequences, the first 32 of which are listed.

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Only the standard IUPAC single-letter nucleotide codes are read. U is read as T.
  • Each ambiguous position is independent of the others.

Common mistakes

  • Using N in a primer's 3′ end, where a mismatch most damages extension.
  • Forgetting that the total primer concentration is shared among all sequences in a degenerate pool.
  • Reading N in a sequencing result as a base that is absent rather than one that could not be called.