T-Test Explained: When to Use It (With a Simple Example)
The t-test is the workhorse of medical research: whenever you want to compare two means on a numeric measurement — blood pressure, test scores, recovery days, hemoglobin levels — the t-test is almost always the right first answer. This article explains the idea in plain words, shows the three types, walks through a worked example with real numbers, and tells you when to pick its close cousins instead.
The idea in one paragraph
You have two averages and a question. The averages will always differ a little, even by pure luck. The t-test asks: is the difference bigger than what random luck would produce, given how spread-out the data is? A big mean difference with tight, consistent data = convincing. A small difference with wildly scattered data = probably noise. The test converts this into a t-statistic and a p-value. A p-value below 0.05 means the difference is "statistically significant" — unlikely to be luck alone.
The three flavors of t-test
Picking the wrong flavor is the commonest student mistake. There are three:
| Type | Use it when… | Example |
|---|---|---|
| Independent (two-sample) | Comparing two separate groups of people | Mean hemoglobin: males vs females |
| Paired | The same people measured twice (before/after) | Weight before vs after a diet plan |
| One-sample | Comparing one group's mean to a fixed number | Is mean birth weight different from 2500 g? |
Quick check: same people twice → paired. Two different sets of people → independent. Comparing against a known standard or cutoff → one-sample.
A note on Welch's t-test
Most textbooks teach the "Student" t-test, which assumes both groups have equal spread (variance). Real data rarely cooperates. Welch's t-test does not need that assumption — it adjusts for unequal spreads and unequal sample sizes. It is the safer default. When the groups are similar in size and spread, both versions agree; when they differ, Welch's answer is the more trustworthy one.
Worked example: paired t-test, step by step
Six patients follow a one-month diet plan. Their weights (kg) before and after:
| Patient | Before | After | Difference |
|---|---|---|---|
| 1 | 72 | 70 | −2 |
| 2 | 65 | 64 | −1 |
| 3 | 80 | 78 | −2 |
| 4 | 62 | 62 | 0 |
| 5 | 75 | 73 | −2 |
| 6 | 78 | 76 | −2 |
Step 1 — work with the differences. The paired t-test analyzes the differences, not the raw weights: −2, −1, −2, 0, −2, −2. Their mean is −1.5 kg.
Step 2 — the test result. The t statistic is the mean difference divided by its standard error: t = −4.39 with 5 degrees of freedom, giving p ≈ 0.007.
Step 3 — the plain-English verdict. p = 0.007 is well below 0.05, so the weight loss is statistically significant: this pattern is very unlikely to be chance. The 95% confidence interval for the true average loss is 0.62 to 2.38 kg — a modest but real effect.
Notice what the test did not say: it did not say the diet is clinically worthwhile, and it did not prove the diet caused the loss (there is no control group). Significance is about chance, not importance.
In your results section you would write: "Mean weight fell from 72.0 kg to 70.5 kg after one month (mean difference −1.5 kg; paired t = −4.39, df = 5, p = 0.007; 95% CI: −2.38 to −0.62 kg)."
Assumptions, in plain words
A t-test trusts its own answer only if three things are roughly true:
- Numeric data. Averages must make sense — a t-test on "male/female" codes is nonsense.
- Roughly symmetric data. With small samples, the values (or the paired differences) should look bell-shaped, not heavily skewed with wild outliers.
- Independent observations. Each person is counted once per group; for the paired test, each pair is independent of the others.
When to use Mann–Whitney U or Wilcoxon instead
If your data is badly skewed or has extreme outliers, the non-parametric cousins are safer — they compare ranks instead of means, so outliers cannot hijack the result:
- Two separate groups → Mann–Whitney U test
- Same people twice → Wilcoxon signed-rank test
For our diet example, the Wilcoxon test gives p ≈ 0.048 — the same conclusion as the t-test, which is reassuring. Running both and comparing is good practice.
Common mistakes
- Using an independent t-test on paired data (or vice versa). Before/after measurements on the same patients must be paired — this is the single most common error in student theses.
- T-testing a yes/no outcome. The t-test needs a numeric measurement. For two categorical variables, use the chi-square test instead.
- Running many t-tests instead of one ANOVA. Comparing three or more groups needs ANOVA, not a pile of t-tests — each extra test inflates your chance of a false "significant" finding.
- Reporting only the p-value. Always give the group means, the mean difference, and the confidence interval. A p of 0.04 with a 1 mmHg difference is significant but clinically meaningless.
Frequently asked questions
How do I choose between paired and independent?
Ask: are the two sets of numbers from the same people (paired) or different people (independent)? If anyone appears in both columns, it is paired.
What if my p-value is 0.06?
It is not significant at the 0.05 level — "almost significant" is not a statistical concept. Report it honestly together with the confidence interval.
When should I use a one-sample t-test?
When you compare one group's average against a fixed reference: a lab cutoff, a published norm, or a target value (for example, "is mean fasting glucose above 100 mg/dL?").
Do I need equal group sizes?
No — Welch's t-test handles unequal sizes and unequal spreads. It is the default you should reach for with two independent groups.
Try it free: run Welch's t-test, the paired t-test and their non-parametric partners on your own data in the FormStat app — no signup, works offline on your phone.
Run this analysis in seconds with the FormStat app. No signup, no laptop, works offline.