Anaes · Measurement & monitoring physics
Hypothesis testing, p-values, confidence intervals and error
Also known as Hypothesis testing · p-values · Confidence intervals · Type I and Type II error · Statistical power · Sample size · Multiple testing · Effect size
Every anaesthetic trial asks whether a difference between two groups is real or the product of chance, and the framework that answers that question is hypothesis testing. The model rests on seven exam-critical ideas. First, the NULL HYPOTHESIS (H0) assumes no difference between groups and is presumed true until the evidence is strong enough to reject it; the ALTERNATIVE HYPOTHESIS (H1) asserts that a difference exists. Second, a TYPE I ERROR (alpha) is a FALSE POSITIVE — rejecting H0 when it is true — conventionally set at 0.05, and the p-value is the probability of the observed data or more extreme data IF H0 is true. Third, a TYPE II ERROR (beta) is a FALSE NEGATIVE — failing to reject H0 when it is false — conventionally set at 0.20. Fourth, POWER is 1 minus beta, the probability of correctly detecting a true difference, conventionally 0.80, and it rises with sample size, effect size and alpha and falls with variability; an underpowered study wastes resources and manufactures false negatives. Fifth, the p-value is widely misused — it is NOT the probability that H0 is true, a p less than 0.05 does not guarantee clinical importance, and multiple comparisons inflate the Type I error rate, corrected by the BONFERRONI method (divide alpha by the number of tests). Sixth, a CONFIDENCE INTERVAL is a range that contains the true population parameter with a specified probability (typically 95 percent), conveying both the effect size and its precision; if the CI crosses 1.0 for a ratio or 0 for a difference the result is not statistically significant. Seventh, the EFFECT SIZE (number needed to treat equals 1 divided by the absolute risk reduction; odds ratio and relative risk as measures of association) tells you HOW BIG a difference is, complementing the p-value which tells you only IF a difference exists; parametric tests (t-test, ANOVA) suit normal data, non-parametric tests (Mann-Whitney U, Kruskal-Wallis) suit non-normal data, and chi-squared suits categorical data; survival analysis uses Kaplan-Meier curves, the log-rank test and the hazard ratio for time-to-event outcomes such as time to recovery or time to first analgesic. Built on the cerebral autoregulation algorithm study (Albanese 2026), the emergence-hypertension risk-factor study (Jiao 2026), the hip-arthroplasty cohort (Ali 2026), the statistical-power commentary (Hanif 2026), the neurological-deterioration risk-prediction study (Chen 2026), the amyloid-beta memory study (Kawabe 2026), the retained-placenta outcomes study (Mukouyama 2026), and the evidence-hierarchies reflection (Tarrant 2026).
Practise this topic
On this page
Study tools
Your progress
Saved on this device.
Target exams
Red flags
- A TYPE I ERROR (alpha) is a FALSE POSITIVE — rejecting a true null hypothesis — conventionally 0.05. A TYPE II ERROR (beta) is a FALSE NEGATIVE — failing to reject a false null hypothesis — conventionally 0.20.
- POWER equals 1 minus beta, the probability of detecting a true difference, conventionally 0.80. Power rises with sample size and effect size and falls with variability.
- The p-value is the probability of the OBSERVED DATA (or more extreme) IF the null hypothesis is true — it is NOT the probability that the null hypothesis is true, and a p less than 0.05 does not mean the result is clinically important.
- A 95 percent CONFIDENCE INTERVAL is a range that would contain the true population parameter in 95 percent of repeated studies; it conveys both effect size and precision. If the CI crosses 1.0 for a ratio or 0 for a difference, the result is NOT statistically significant.
- NUMBER NEEDED TO TREAT equals 1 divided by the ABSOLUTE risk reduction. The p-value tells you IF a difference exists; the effect size (NNT, odds ratio, relative risk) tells you HOW BIG it is.
- Multiple comparisons inflate the Type I error rate. The BONFERRONI correction divides alpha by the number of tests — reducing false positives at the cost of power. A pre-specified single primary outcome avoids the problem.
Meet the patient
A journal club paper claims a new induction agent reduces postoperative nausea compared with propofol, p equals 0.04. The registrar asks whether you should switch. The confidence interval is wide, the sample is small, and the absolute difference in nausea rates is 2 per cent.[1][4]
Two questions decide whether this paper changes practice: is the difference real? (the p-value and confidence interval answer this) and is it big enough to matter? (the effect size answers this). A p-value of 0.04 in a small sample with a tiny absolute benefit is statistically significant but clinically trivial — and that distinction is the heart of this topic.[1][4]
References8ShowHide
- [1]Albanese A, et al. A Novel Algorithm for Continuous Real-Time Cerebral Autoregulation Assessment Based on Mean Arterial Pressure and Cerebral Oxygen Saturation Anesth Analg, 2026.PMID 42363900
- [2]Jiao SS, et al. Risk factors for unanticipated hypertension during emergence from general anesthesia in elderly surgical patients: A retrospective cohort analysis Medicine (Baltimore), 2026.PMID 42363520
- [3]Ali U, et al. Outcomes of lateral and posterior approaches in hip arthroplasty: A cohort from low-middle-income country J Pak Med Assoc, 2026.PMID 42363370
- [4]Hanif L, et al. Comment on Avoidant Personality Traits and Avoidant Coping in Cognitive-Behavioral Therapy vs. Short-Term Psychodynamic Psychotherapy for Adult Depression Personal Ment Health, 2026.PMID 42363611
- [5]Chen Y, et al. Risk prediction models for early neurological deterioration after intravenous thrombolysis in acute ischemic stroke patients: a systematic review and meta-analysis BMC Neurol, 2026.PMID 42363152
- [6]Kawabe N, et al. Logical memory is associated with amyloid-β positivity in patients with early Alzheimer's disease eligible for lecanemab J Alzheimers Dis, 2026.PMID 42363806
- [7]Mukouyama F, et al. Association of Retained Placenta After Spontaneous Separation With Conversion to Hysterectomy in Cases of Cesarean Sections With Suspected Placenta Accreta Spectrum J Obstet Gynaecol Res, 2026.PMID 42363659
- [8]Tarrant A, et al. What counts as knowing? Reflections on evidence hierarchies in qualitative longitudinal and participatory research with young fathers Evid Policy, 2026.PMID 42362383