Hypothesis Tests

McNemar Test Calculator

Tests marginal symmetry in paired binary data from the two discordant cell counts. The example keeps the method and inputs visible so the result can be checked independently.

Test inputs

Describe the observed sample

pairs
pairs
Calculated result

McNemar test

Result
exact binomial test of discordant pairs

    Two coherent versions of the analysis during independent review

    Create a second scenario that changes one uncertain input rather than mixing optimistic values from unrelated cases. Compare both the center and the uncertainty or test statistic. The worked values provide a baseline for the comparison.

    If the interpretation reverses under a small defensible change, report that sensitivity. It is more informative than presenting one apparently exact test result. The report should state this boundary plainly.

    Before accepting mcnemar test, compare the result with the scale of the raw measurement or event rate. A numerically small difference can matter on a tightly controlled scale, while a larger difference may be uninformative when ordinary variation is much wider. The substantive benchmark belongs beside the statistical calculation.

    The target parameter

    Tests marginal symmetry in paired binary data from the two discordant cell counts. The displayed result follows exact binomial test of discordant pairs, with every symbol tied to a labeled input. A changed sample requires the same check again.

    Seven versus 19 discordant pairs give an exact two-sided p-value of about 0.029. This worked condition is a reproducible arithmetic check, not evidence that the model fits every dataset. This point matters before the result enters another model.

    For mcnemar test, a useful audit begins with the numerator and denominator rather than the final display. Write the observed quantity, its reference value, and the uncertainty term on separate lines. That layout makes a misplaced square root, reversed group order, or percentage-scale error visible before rounding.

    Assumptions carried by the formula under the stated design

    Concordant pairs do not enter the test statistic, although they remain part of the study description. This prevents a plausible number from carrying the wrong meaning.

    The unit of analysis, sampling frame, dependence structure, and treatment of missing values remain outside the final number. Record those choices before interpreting this test. That is a design choice, not a display setting.

    Input definitions that matter in the worked condition

    Check that counts are whole observations, scales refer to the same measurement, and standard errors or deviations come from the population or sample named on the page. A percentage and a proportion differ by a factor of 100. That choice determines which comparison is defensible.

    If a critical value is entered, it must match the intended tail convention and reference degrees of freedom. Changing confidence level without changing that value creates a mislabeled result. The source record should resolve that question.

    A direct numerical cross-check before the result is reused

    Recalculate one intermediate quantity from exact binomial test of discordant pairs and then work backward from the displayed endpoint or statistic. This catches swapped groups, reversed quantiles, and copied denominators. A reverse calculation can expose an inconsistency here.

    Vary one credible input while holding the rest fixed. The direction and size of the change should agree with the formula before the result is carried into a report. This is where a group-order error is easiest to catch.

    From numerical result to conclusion during independent review

    The p-value measures compatibility between the observed statistic and the null model. It is not the probability that the null hypothesis is true. The numerical precision does not override that requirement.

    Practical importance requires the effect size, measurement scale, uncertainty, and consequences of a decision. A threshold crossing by itself does not supply that context. This condition can be checked without relying on the final display.

    Conditions that change the method

    Sparse cells, strong skew, influential observations, clustering, pairing, estimated nuisance parameters, or unequal variances can change the reference distribution. Concordant pairs do not enter the test statistic, although they remain part of the study description. This check belongs before rounding.

    Do not choose among methods by selecting the answer that looks most favorable. Choose from the data-generating design, then preserve the method name and convention. That step separates arithmetic from interpretation.

    Before reporting the result

    When should mcnemar test be repeated?

    Repeat it when an input, exclusion, group definition, confidence level, tail choice, or model assumption changes. For this page, the reported quantity is mcnemar test.

    With the source record open, how many digits should be reported?

    Retain guard digits during checking, then round to a level justified by the source measurement and the decision that follows. For this page, the reported quantity is mcnemar test.

    For the selected tail convention, can a missing value be entered as zero?

    Only when zero was observed. Missingness and a measured zero have different statistical meanings. For this page, the reported quantity is mcnemar test.

    Before drawing a conclusion, does a narrow interval prove the estimate is unbiased?

    No. Precision under a model does not repair selection, measurement, nonresponse, or specification bias. For this page, the reported quantity is mcnemar test.

    During an independent check, what belongs in a reproducible record?

    Save the input summaries or data, unit of analysis, formula convention, exclusions, unrounded output, and software or table method used. For this page, the reported quantity is mcnemar test.

    When inputs are revised, what does the reported p-value mean?

    It describes how unusual this statistic or a more extreme one would be under the stated null model; it is not the probability that the null is true. For this page, the reported quantity is mcnemar test.