BioDeviceHub

Descriptive Statistics Calculator

Calculate the mean, median, mode, variance, standard deviation, standard error, coefficient of variation, quartiles, skewness and kurtosis of a list of numbers.

Formula

xˉ=∑xn\bar{x} = \dfrac{\sum x}{n}
s2=∑(x−xˉ)2n−1;s=s2s^2 = \dfrac{\sum (x - \bar{x})^2}{n - 1};\quad s = \sqrt{s^2}
SE=sn;CV=sxˉ×100\mathrm{SE} = \dfrac{s}{\sqrt{n}};\quad \mathrm{CV} = \dfrac{s}{\bar{x}}\times 100
n − 1n\ -\ 1
degrees of freedom of a sample variance; n is used for a whole population
quartiles\text{quartiles}
25th and 75th percentiles, by linear interpolation between order statistics

How it works

Descriptive statistics summarise a set of numbers. The mean and median describe where it is centred, and the standard deviation and the interquartile range how spread out it is. The standard error describes the uncertainty of the mean rather than the spread of the data, and is smaller by the square root of n. The coefficient of variation expresses the spread relative to the mean.

When the data are a sample from a larger population, as they almost always are, the variance divides by n − 1, which corrects the tendency of the sample to underestimate the spread. The population formulas, dividing by n, apply only when every member has been measured. Skewness and excess kurtosis show whether the data are lopsided or heavy-tailed, which matters for deciding whether the mean and SD are good summaries.

Worked example

Ten measurements: 4.2, 5.1, 3.8, 4.9, 5.6, 4.4, 4.9, 6.1, 4.7 and 5.0.

  1. Sum = 48.7, so mean = 4.87; the median is 4.9 and the mode is 4.9.
  2. Variance = 0.4401 (n − 1 = 9), SD = 0.6634, SE = 0.6634 / √10 = 0.2098.
  3. Quartiles 4.475 and 5.075, so the interquartile range is 0.6; CV = 13.6%.

Mean 4.87, SD 0.663, median 4.9, with slight positive skew (0.30).

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Numeric, independent observations of one quantity.
  • For the sample formulas, that the values are a sample from a larger population.
  • Quartiles use the same definition as Excel's QUARTILE.INC and the default in R (type 7).

Common mistakes

  • Reporting the standard error as if it described the spread of the data.
  • Using the population formula on a sample, which understates the spread.
  • Summarising a skewed or multimodal set by the mean and SD alone.