BioDeviceHub

Correlation Calculator (Pearson and Spearman)

Calculate Pearson's r or Spearman's rank correlation ρ for paired data, with a confidence interval by the Fisher transform, a test of zero correlation and a scatter plot.

Formula

r=∑(x−xˉ)(y−yˉ)∑(x−xˉ)2 ∑(y−yˉ)2r = \dfrac{\sum (x - \bar{x})(y - \bar{y})}{\sqrt{\sum (x - \bar{x})^2\ \sum (y - \bar{y})^2}}
t=rn−21−r2t = \dfrac{r\sqrt{n - 2}}{\sqrt{1 - r^2}}
interval:tanh⁡ ⁣(atanh⁡r±zn−3)\text{interval:}\quad \tanh\!\left(\operatorname{atanh} r \pm \dfrac{z}{\sqrt{n - 3}}\right)
r, ρr,\ \rho
the coefficient, from −1 (perfect inverse) to +1 (perfect direct)
zz
1.96 for a 95% interval

How it works

Pearson's r measures how closely two variables follow a straight line. Spearman's ρ is the same calculation on the ranks, so it detects any monotonic relationship, straight or curved, and is much less affected by outliers. The interval is computed on the Fisher z scale, where r is approximately normal, then transformed back.

A coefficient near zero does not mean the variables are unrelated, since a curved relationship can give r = 0, and a large one does not mean one causes the other. Plot the data first; Anscombe's famous examples show very different patterns sharing the same r.

Worked example

The same eight points as the regression example.

  1. r = 0.999, r² = 0.998.
  2. Fisher z interval: 0.994 to 1.000; t = 54.1 with 6 degrees of freedom.

A near-perfect linear correlation, r = 0.999 (95% CI 0.994 to 1.000).

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Paired, independent observations.
  • For Pearson, a roughly linear relationship and no extreme outliers, with bivariate normality for the interval.
  • For Spearman, a monotonic relationship; its p-value and interval are approximate for small n.

Common mistakes

  • Inferring causation from correlation.
  • Calculating r on a range restricted by selection, which lowers it.
  • Correlating two quantities that share a component, or pooling distinct groups so that the group difference creates the correlation.