BioDeviceHub

t-Test Calculator (Welch, Student, One-Sample)

Compare two groups by Welch's or Student's t-test, or one group with a reference, with the t statistic, degrees of freedom, p-value, confidence interval and effect size.

Formula

Welch:t=xˉ1−xˉ2s12n1+s22n2\text{Welch:}\quad t = \dfrac{\bar{x}_1 - \bar{x}_2}{\sqrt{\dfrac{s_1^2}{n_1} + \dfrac{s_2^2}{n_2}}}
Student:t=xˉ1−xˉ2sp1n1+1n2\text{Student:}\quad t = \dfrac{\bar{x}_1 - \bar{x}_2}{s_p\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}}
one-sample:t=xˉ−μ0s/n\text{one-sample:}\quad t = \dfrac{\bar{x} - \mu_0}{s/\sqrt{n}}
sps_p
pooled standard deviation of the two groups
ν\nu
Welch-Satterthwaite for Welch's test; n₁ + n₂ − 2 for Student's; n − 1 for one sample

How it works

The t-test compares a difference in means with the variation expected from chance. The statistic is the difference divided by its standard error, and the p-value is the probability, if there were truly no difference and the test's assumptions held, of a statistic at least that large. Student's version assumes equal variances in the two groups. Welch's version does not, at little cost when they are equal, which is why it is recommended as the default.

The p-value alone is a poor summary. The page reports the method, the sample sizes, the degrees of freedom, the test statistic, the p-value, a confidence interval for the difference and an effect size (Hedges' g), and never labels a result 'significant'. The interval shows which differences are compatible with the data, and the effect size how large the difference is relative to the variability.

Worked example

Two groups: seven values with mean 5.371 and eight with mean 6.613, by Welch's test.

  1. Difference = 5.371 − 6.613 = −1.241; its standard error is 0.260.
  2. t = −4.772 with 12.96 degrees of freedom; p = 0.00037.
  3. 95% CI for the difference: −1.803 to −0.679; Hedges' g = −2.31.

Group 1 is lower by 1.24 (95% CI 0.68 to 1.80), a large standardised difference; p = 0.00037 under the test's assumptions.

These are the values the calculator opens with, so you can check its output against this example.

Assumptions

  • Independent observations, and independent groups.
  • Approximately normal values in each group, or large enough samples; a formal normality test has little power in small samples, so look at the data.
  • For Student's test, equal variances.

Common mistakes

  • Using an unpaired test on paired data, or the reverse.
  • Treating p < 0.05 as proof of an effect and p > 0.05 as proof of none.
  • Running many t-tests between several groups instead of an ANOVA with a suitable follow-up.