BioDeviceHub

FASTA Sequence Analyzer

Paste one sequence or a multi-FASTA file and get the length, GC content, base counts, ambiguous-base count and CpG observed/expected ratio of every record.

Formula

GC %=G+CA+C+G+T×100\mathrm{GC}\,\% = \dfrac{G + C}{A + C + G + T}\times 100
CpGo/e=nCpG×LnC×nG\mathrm{CpG}_{o/e} = \dfrac{n_{\mathrm{CpG}}\times L}{n_{C}\times n_{G}}
nCpGn_{\mathrm{CpG}}
number of adjacent C followed by G on the strand entered
LL
number of characters in the cleaned sequence

How it works

The text is split into records at each line starting with >, and digits, spaces and gap characters are removed. Each record is then counted base by base. Percentages are calculated over the unambiguous bases only, and any other character (N, R, Y and so on) is counted separately so it does not distort them.

The CpG observed/expected ratio compares how often CG occurs with how often it would occur by chance given the C and G content. It is depleted in most vertebrate DNA, so a high ratio over a stretch of a few hundred bases is one of the features used to identify CpG islands. For a single sequence a dinucleotide table is also shown.

Worked example

Two short records: the first is an illustrative 39-base coding sequence, the second a 36-base sequence with one N.

  1. Record 1: 9 A, 8 C, 14 G, 8 T gives GC = (14 + 8) / 39 = 56.4%.
  2. Record 2: 8 A, 8 C, 10 G, 9 T and one N gives GC = 18 / 35 = 51.4%.

Two sequences, 75 bases in total, with the N in record 2 reported under Other.

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Sequences are nucleic acids. A protein FASTA file will be counted letter by letter and is not meaningful here.
  • CpG o/e is reported for the strand as entered and is most informative for sequences of several hundred bases or more.

Common mistakes

  • Pasting a protein FASTA, which reads as mostly ambiguous characters.
  • Comparing the GC of sequences of very different length without looking at the lengths. Short sequences vary more.
  • Reading CpG o/e on a very short sequence, where one CG more or less changes it a lot.