Sequence Comparison and Alignment Tool
Compare two short sequences by global alignment (Needleman-Wunsch) with your own match, mismatch and gap scores, or position by position, and read the percent identity.
Formula
- scores for each kind of column in the alignment; the algorithm finds the alignment with the highest total
How it works
Two related sequences are rarely the same length, so comparing them base for base fails as soon as there is an insertion or deletion. The Needleman-Wunsch algorithm finds the alignment, allowing gaps, with the highest total score. It is an exact dynamic-programming method over the whole of both sequences (a global alignment).
The scoring is yours to set. With a gap penalty that is harsh relative to a mismatch, the alignment avoids gaps; with a mild one, it inserts them freely to line up more bases. Position-by-position comparison needs sequences of equal length and simply counts the differences.
Worked example
Two 35-base sequences differing at one base and with a one-base insertion and a one-base deletion between them, scored +1, −1 and −2.
- The optimal global alignment is 36 columns long.
- 33 columns match, 1 mismatches and 2 contain a gap.
- Score = 33 × 1 + 1 × (−1) + 2 × (−2) = 28.
33 of 36 columns are identical: 91.7% identity over the alignment length.
These are the values the calculator opens with, so you can check its output against this example.
Assumptions
- A linear gap penalty: a gap of length n costs n times the penalty. Affine penalties, which treat opening and extending differently, are not used.
- Sequences up to 1,500 characters. The work grows with the product of the two lengths.
- Percent identity here is matches over the full alignment length, gaps included. Other programs divide by the shorter sequence or by aligned columns, and the results differ.
Common mistakes
- Comparing a forward sequence with a reverse-strand one. Reverse complement one of them first.
- Calling two sequences homologous because they are 70% identical over a short stretch; significance depends on length and composition.
- Comparing percent identity from different programs as though they were the same quantity.