FASTA Sequence Analyzer
Paste one sequence or a multi-FASTA file and get the length, GC content, base counts, ambiguous-base count and CpG observed/expected ratio of every record.
Formula
- number of adjacent C followed by G on the strand entered
- number of characters in the cleaned sequence
How it works
The text is split into records at each line starting with >, and digits, spaces and gap characters are removed. Each record is then counted base by base. Percentages are calculated over the unambiguous bases only, and any other character (N, R, Y and so on) is counted separately so it does not distort them.
The CpG observed/expected ratio compares how often CG occurs with how often it would occur by chance given the C and G content. It is depleted in most vertebrate DNA, so a high ratio over a stretch of a few hundred bases is one of the features used to identify CpG islands. For a single sequence a dinucleotide table is also shown.
Worked example
Two short records: the first is an illustrative 39-base coding sequence, the second a 36-base sequence with one N.
- Record 1: 9 A, 8 C, 14 G, 8 T gives GC = (14 + 8) / 39 = 56.4%.
- Record 2: 8 A, 8 C, 10 G, 9 T and one N gives GC = 18 / 35 = 51.4%.
Two sequences, 75 bases in total, with the N in record 2 reported under Other.
These are the values the calculator opens with, so you can check its output against this example.
Assumptions
- Sequences are nucleic acids. A protein FASTA file will be counted letter by letter and is not meaningful here.
- CpG o/e is reported for the strand as entered and is most informative for sequences of several hundred bases or more.
Common mistakes
- Pasting a protein FASTA, which reads as mostly ambiguous characters.
- Comparing the GC of sequences of very different length without looking at the lengths. Short sequences vary more.
- Reading CpG o/e on a very short sequence, where one CG more or less changes it a lot.