Observed Versus Expected: The Logic Behind the Chi-Square Test
In a hurry? Skip straight to the numbers.
Open the Chi-Square Calculator →The companion calculator computes the chi-square statistic by comparing observed counts to expected counts, for both goodness-of-fit and independence tests. That comparison, what you actually observed versus what a model predicts, is one of the most general and intuitive ideas in statistical testing, and the chi-square statistic is its purest expression for categorical data. Understanding the logic of observed versus expected, how it tests whether categorical data fits a model or whether two variables are related, and what degrees of freedom capture turns a chi-square calculation into an appreciation of a foundational approach to analyzing counts.
The Observed-Versus-Expected Idea
The core logic of the chi-square test is beautifully simple: compare the counts you actually observed to the counts a model or hypothesis would predict, and judge whether the differences are small enough to be chance or large enough to signal something real. Every statistical model or hypothesis about categorical data implies expected counts, what you would see on average if the model were true, and the data provides observed counts, what you actually saw. If the model is correct, the observed counts should be close to the expected ones, with only the small discrepancies that random sampling produces. If the observed counts deviate too far from the expected, that is evidence the model is wrong. The chi-square statistic quantifies this by summing the squared differences between observed and expected, scaled by the expected counts, so larger deviations produce a larger statistic and stronger evidence against the model, as the calculator's formula embodies. Understanding the observed-versus-expected idea is the key to the chi-square test and to much of statistics: testing a hypothesis often means comparing what you see to what the hypothesis predicts, and asking whether the gap is bigger than chance would produce. The chi-square statistic is the measure of that gap for categorical counts, turning the intuitive comparison into a rigorous test.
Two Questions, One Logic
The chi-square test applies the same observed-versus-expected logic to two different questions, which is why the calculator offers both goodness-of-fit and independence tests.
| Test | Expected counts come from |
|---|---|
| Goodness of fit | A theoretical distribution or hypothesis |
| Test of independence | The assumption that two variables are unrelated |
In a goodness-of-fit test, the expected counts come from a hypothesized distribution, like equal counts for a fair die, and the test asks whether the observed counts fit that distribution, as the calculator's die-fairness example shows. In a test of independence, the expected counts are computed under the assumption that two categorical variables are unrelated, from the row and column totals, and the test asks whether the observed counts in the contingency table depart from what independence would predict, signaling that the variables are associated. Both use the identical logic: compute what you would expect under the model (a specific distribution, or independence), compare to what you observed, and measure the discrepancy with the chi-square statistic. The only difference is where the expected counts come from. Understanding that two questions share one logic reveals the versatility of the observed-versus-expected approach: whether testing a distribution or a relationship, the method is the same, compare observed to expected and judge the gap. This is why the chi-square test is so widely used for categorical data, from testing fairness to detecting associations, as the calculator's applications span, all through the single unifying idea of comparing counts to a model's predictions.
What Degrees of Freedom Capture
The chi-square test's result depends on the degrees of freedom, a concept that captures how much freedom the data had to vary given the constraints, and understanding it demystifies the test. Degrees of freedom count the number of independent pieces of information that go into the statistic, after accounting for the constraints imposed by the totals used to compute the expected counts. For a goodness-of-fit test, once the total count is fixed, the counts in all but one category are free to vary, and the last is determined, so the degrees of freedom are one less than the number of categories. For a contingency table, the degrees of freedom account for the row and column totals being fixed, leaving fewer cells free to vary independently. The degrees of freedom matter because the expected size of the chi-square statistic under a true model grows with them, so the same chi-square value means different things depending on how many degrees of freedom there are, which is why the p-value depends on both the statistic and the degrees of freedom, as the calculator reports. Understanding what degrees of freedom capture, the number of independent comparisons after constraints, explains why they appear in the test and why the critical chi-square values the calculator lists rise with degrees of freedom: more categories or a larger table means more room for random deviation, so a larger statistic is needed to signal a real effect. Degrees of freedom calibrate the test to the amount of data and structure involved.
Reading a Chi-Square Result
The practical upshot is that a chi-square test tells you whether the discrepancy between observed and expected is larger than random sampling would typically produce, and reading it correctly means understanding what it does and does not say. A large chi-square statistic and small p-value mean the observed counts deviate from the expected more than chance would likely produce, providing evidence against the model, whether that the distribution does not fit or that the variables are associated, as the calculator computes. A small statistic and large p-value mean the observed counts are consistent with the model, as in the calculator's fair-die example where the counts fit expectation and there is no evidence of bias. But the test has limits worth knowing: it detects that observed and expected differ, not why or how, and for the independence test, a significant result indicates association but not its strength or direction, which requires other measures. The test also assumes adequate expected counts and independent observations to be valid. Understanding how to read a chi-square result completes the picture: it answers whether the gap between observed and expected is beyond chance, providing evidence about a distribution or relationship, but it is one step in analysis, indicating that something differs from the model without fully characterizing it. The calculator computes the statistic, degrees of freedom, and p-value; understanding the observed-versus-expected logic is what reveals what that result means and what it leaves for further analysis.
Understanding the Chi-Square Test
Use the calculator to compute a chi-square statistic, and understand its logic: the test compares observed counts to the counts a model predicts, measuring whether the discrepancy is beyond chance, and the same observed-versus-expected logic tests both whether data fits a distribution and whether two variables are related, with degrees of freedom capturing the independent comparisons after constraints. The calculation gives the statistic and p-value; understanding observed versus expected is what reveals how comparing what you see to what a model predicts tests categorical data.
Ready to Put This Into Practice?
Now that you understand how it works, plug in your own numbers and get an instant, accurate result.
Use the Chi-Square Calculator Now →