BioDeviceHub

GC Content Profile (Sliding Window)

Plot the GC content along a DNA sequence in a sliding window of a size and step you choose, and find the most GC-rich and GC-poor stretches.

Formula

GCwindow %=G+CA+C+G+T×100(counts in the window)\mathrm{GC}_{\text{window}}\,\% = \dfrac{G + C}{A + C + G + T}\times 100\quad(\text{counts in the window})
windows start at 1, 1+s, 1+2s, … (s=step)\text{windows start at } 1,\ 1 + s,\ 1 + 2s,\ \ldots\ (s = \text{step})
window\text{window}
number of consecutive bases counted at each point
step\text{step}
how far the window moves between points

How it works

GC content is rarely uniform along a sequence. Coding regions, promoters, CpG islands, replication origins and horizontally acquired genes can differ from their surroundings. Sliding a window along the sequence and plotting its GC content shows those differences.

A smaller window shows finer detail but is noisier, because each point rests on fewer bases; a larger window is smoother but blurs short features. The dashed line marks the average over the whole sequence.

Worked example

A 108-base illustrative sequence alternating AT-rich and GC-rich stretches, window 20, step 5.

  1. The first window covers bases 1 to 20, the second 6 to 25, and so on.
  2. There are (108 − 20) / 5 + 1 = 18 windows.

Overall GC of 50.9%, with the most GC-rich window at 75.0% (bases 36 to 55) and the poorest at 25.0% (bases 51 to 70).

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Windows contain only unambiguous bases. A window with no A, C, G or T is skipped.
  • The position plotted is the centre of the window.
  • At most 2,000 windows are drawn.

Common mistakes

  • Interpreting a peak from a very small window as a real feature. Check it with a larger window.
  • Comparing profiles made with different window sizes.
  • Treating a dip in GC as evidence of a gene boundary. It is a hint to investigate, not a conclusion.