📊 FormStat
FormStat Guides · By Abdul Hannan

ANOVA Explained Simply: Comparing 3 or More Groups

You want to compare average test scores across three medical colleges — or blood pressure across four drug doses. The t-test only handles two groups, so what now? Meet ANOVA (Analysis of Variance): one test that asks whether any of the group averages differ, without the trap of running test after test.

Why not just run several t-tests?

With three groups (A, B, C) you would need three t-tests: A vs B, A vs C, B vs C. Here is the problem: each test at the 0.05 level carries a 5% chance of a false alarm. Run three, and your real chance of at least one false "significant" result climbs to about 14% (1 − 0.95³). With five groups you need ten tests, and the false-alarm rate passes 40%. ANOVA does the whole job in one test at a clean 5%.

What ANOVA actually compares

Despite the name, ANOVA compares means — it just does it by studying variances. The logic has two parts:

If the group averages are far apart relative to the natural wobble inside the groups, something real is going on. That ratio is the F statistic: F = (variation between groups) ÷ (variation within groups). F near 1 means "nothing to see here"; a large F means "at least one group stands out".

Worked example, step by step

Three colleges report student test scores (summary data — N, mean and SD per college):

CollegeStudents (N)Mean scoreSD
College 16355.110.93
College 21747.67.08
College 31549.410.20

The ANOVA table built from these summaries:

SourceSum of squaresd.f.Mean squareFp-value
Between groups967.82483.94.6060.0124
Within groups9665.492105.1––

Reading it: the between-group mean square (483.9) is 4.6 times the within-group mean square (105.1), so F = 4.606. The p-value is 0.0124 — below 0.05, so the differences between colleges are statistically significant. In plain words: at least one college's average genuinely differs from the others; the natural spread of scores cannot explain the gap. Note what ANOVA does not tell you: which pair differs — that needs follow-up pairwise comparisons.

Bartlett's check, in one paragraph

ANOVA assumes the groups have roughly similar spread. Bartlett's test checks exactly this: here it returns p = 0.141 — above 0.05, so there is no evidence the variances differ, and the ANOVA result stands on solid ground. If Bartlett's p had fallen below 0.05, the wise move would be the non-parametric alternative below.

One-way vs two-way ANOVA

The ANOVA in this guide is one-way: one grouping factor (college). Two-way ANOVA handles two factors at once — for example, college and gender — and can even detect an interaction (say, the gender gap differs between colleges). Two-way ANOVA needs raw data in a specific layout and is overkill for most student theses; if your objective names a single grouping variable, one-way is the right call, and it is what FormStat runs.

Kruskal–Wallis: the non-parametric alternative

When scores are skewed, groups are tiny, or Bartlett's test complains, use the Kruskal–Wallis test — ANOVA's rank-based cousin. It asks the same "does any group differ?" question without assuming bell-shaped data or equal variances. It is slightly less powerful than ANOVA when the data is well-behaved, which is why you check assumptions first instead of defaulting to it. If your data fails ANOVA's assumptions, this is your answer, not a fancier parametric fix.

Common mistakes

Frequently asked questions

Can I do ANOVA without the raw data?

Yes — ANOVA needs only N, mean and SD per group, which is exactly what published papers report. FormStat's "ANOVA from summary data" takes those three numbers per group and returns the full table, Bartlett's test and confidence intervals.

What does a non-significant ANOVA mean?

The data gives no convincing evidence of any difference between the group means. It does not prove the groups are identical — just that this sample cannot tell them apart.

How many groups can ANOVA handle?

Any number from three upward (with two groups it is equivalent to a t-test). Summary-data ANOVA accepts up to ten groups.

After a significant ANOVA, what comes next?

Pairwise comparisons to find which groups differ — and always report group means with confidence intervals, not just the p-value.

Should I report an effect size too?

Yes. The p-value says a difference exists; eta-squared says how big it is: between-group sum of squares ÷ total sum of squares. Here that is 967.8 ÷ 10633.2 ≈ 0.091 — about 9% of the variation in scores is explained by which college a student attends. Report it next to the p-value; examiners increasingly expect it.

Does ANOVA need equal group sizes?

No. Unequal groups (like our 63/17/15 example) are fine; the math weights each group by its size. Extremely tiny groups (fewer than 5–10 observations) are the real problem — their means are too unstable to trust.

Try it free: run one-way ANOVA and Kruskal–Wallis on raw data — or ANOVA from just N, mean and SD — in the FormStat app. No signup, works offline on your phone.

Try it yourself — free.
Run this analysis in seconds with the FormStat app. No signup, no laptop, works offline.

← All guides