BioDeviceHub

Linear Regression Calculator

Fit a least-squares line to paired data, with slope and intercept and their confidence intervals, R², a slope test, the residuals, a plot and predictions with intervals.

Formula

m=∑(x−xˉ)(y−yˉ)∑(x−xˉ)2;b=yˉ−mxˉm = \dfrac{\sum (x - \bar{x})(y - \bar{y})}{\sum (x - \bar{x})^2};\quad b = \bar{y} - m\bar{x}
R2=1−SSresSStotR^2 = 1 - \dfrac{SS_{\text{res}}}{SS_{\text{tot}}}
s=SSresn−2;SE(m)=s∑(x−xˉ)2s = \sqrt{\dfrac{SS_{\text{res}}}{n - 2}};\quad \mathrm{SE}(m) = \dfrac{s}{\sqrt{\sum (x - \bar{x})^2}}
ss
residual standard deviation, with n − 2 degrees of freedom
95% CI95\%\ \mathrm{CI}
estimate ± t(0.975, n − 2) × standard error

How it works

Ordinary least squares finds the line that minimises the sum of squared vertical distances from the points. The slope says how much y changes per unit of x, and the intercept the value of y at x = 0. The standard errors, from the scatter about the line, give confidence intervals for both, and a t-test of whether the slope is zero.

R² is the fraction of the variation in y accounted for by the line. It does not show that a line is the right model: inspect the residuals, which should scatter evenly around zero with no pattern. The page also predicts y at a chosen x, with a narrower interval for the mean response and a wider one for a single new observation.

Worked example

Eight points: x from 1 to 8, y from 2.1 to 16.5.

  1. Fit: y = 2.031 x − 0.0643; the slope's 95% CI is 1.939 to 2.123.
  2. R² = 0.998; the residual SD is 0.243 with 6 degrees of freedom; t = 54.1 for the slope.
  3. At x = 4.5 the fit gives 9.075; the 95% interval for the mean is 8.865 to 9.285 and for a new value 8.444 to 9.706.

y = 2.031 x − 0.064, R² = 0.998, with the prediction interval wider than the interval for the mean.

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • A linear relationship, with independent errors that are roughly normal and of constant size at every x.
  • x is known much more precisely than y is measured.
  • Predictions are made within the range of the data.

Common mistakes

  • Treating a high R² as proof that the relationship is linear.
  • Extrapolating beyond the range of x, where the line may not hold.
  • Concluding causation from an association.