Anaes · Measurement & monitoring physics
Diagnostic test performance: sensitivity, specificity, predictive values, likelihood ratios and ROC
Also known as Sensitivity and specificity · Predictive values · Likelihood ratios · ROC curves · Bayes' theorem · Screening versus diagnostic testing · STARD guidelines · Diagnostic test accuracy
Every anaesthetic decision rests on a test — a troponin, a D-dimer, a lactate, a FAST scan, a Mallampati score, an AI flag on a CT — and the framework that tells you whether to trust that test is diagnostic test performance. The model rests on seven exam-critical ideas. First, every test result can be laid out in a 2x2 CONTINGENCY TABLE of disease present or absent against test positive or negative, giving four cells — TRUE POSITIVES (TP), FALSE POSITIVES (FP), FALSE NEGATIVES (FN) and TRUE NEGATIVES (TN) — from which every metric is derived. Second, SENSITIVITY (Sn) is the TRUE POSITIVE RATE, TP divided by (TP plus FN), the probability the test is positive GIVEN disease; because it is computed only in those who have disease it is a fixed property of the test, independent of prevalence, and a highly sensitive test used to RULE OUT disease gives the mnemonic SnNout. Third, SPECIFICITY (Sp) is the TRUE NEGATIVE RATE, TN divided by (TN plus FP), the probability the test is negative GIVEN no disease, also prevalence-independent, and a highly specific test used to RULE IN disease gives the mnemonic SpPin. Fourth, the PREDICTIVE VALUES answer the clinical question: POSITIVE PREDICTIVE VALUE is TP divided by (TP plus FP), the probability of disease GIVEN a positive test, and NEGATIVE PREDICTIVE VALUE is TN divided by (TN plus FN), the probability of no disease GIVEN a negative test — and unlike sensitivity and specificity the predictive values DEPEND ON PREVALENCE, so a test that performs well in a high-risk ICU population may generate mostly false positives in a low-risk screening clinic. Fifth, BAYES' THEOREM links the pre-test probability (the prevalence) to the post-test probability through the LIKELIHOOD RATIO: post-test odds equals pre-test odds times the likelihood ratio. Sixth, the LIKELIHOOD RATIOS combine sensitivity and specificity into a single prevalence-independent number — LR+ equals sensitivity divided by (1 minus specificity) and LR- equals (1 minus sensitivity) divided by specificity; an LR+ greater than 10 or an LR- less than 0.1 is very useful, and because LRs do not depend on prevalence they travel with the test from the validation population to your patient. Seventh, the RECEIVER OPERATING CHARACTERISTIC (ROC) curve plots sensitivity against (1 minus specificity) for every possible threshold and the AREA UNDER THE CURVE (AUC) summarises overall discrimination — 0.5 is no better than chance, 1.0 is perfect, and an AUC greater than 0.9 is excellent; the chosen threshold trades sensitivity against specificity, so a SCREENING threshold is set for high sensitivity (do not miss cases — D-dimer for PE, lactate for sepsis, HIV screening) while a CONFIRMATORY threshold is set for high specificity (do not raise false alarms — biopsy, coronary angiography). New tests are evaluated against a REFERENCE STANDARD under the STARD reporting guidelines, guarding against SPECTRUM BIAS (the study population does not match the clinical one) and VERIFICATION BIAS (only positive tests get the gold standard). Built on the AI iridocorneal-angle classification study (Rubegni 2026), the gadoxetate-MRI diagnostic-value study (Sandrasegaran 2026), the cerebral-autoregulation algorithm (Albanese 2026), the deep-learning glaucoma referral study (Lima-Cabrita 2026), the AI-CT diagnostic-value study for MACE (Xu 2026), the bladder-endometriosis diagnosis study (Ozdemir 2026), the generative-AI-in-healthcare review (Li JZ 2026), and the regional-versus-general-anaesthesia meta-analysis (Li P 2026).
Practise this topic
On this page
Study tools
Your progress
Saved on this device.
Target exams
Red flags
- SENSITIVITY is the TRUE POSITIVE RATE, TP divided by (TP plus FN) — the probability the test is positive GIVEN disease. It is prevalence-INDEPENDENT. A highly sensitive test RULES OUT (SnNout). SPECIFICITY is the TRUE NEGATIVE RATE, TN divided by (TN plus FP) — probability test negative GIVEN no disease, prevalence-independent, and a highly specific test RULES IN (SpPin).
- PPV equals TP divided by (TP plus FP) and NPV equals TN divided by (TN plus FN). The predictive values DEPEND ON PREVALENCE — high prevalence RAISES the PPV and LOWERS the NPV. This is why a test validated in a high-risk ICU may fail in a low-risk screening clinic.
- LIKELIHOOD RATIOS are prevalence-INDEPENDENT. LR+ equals sensitivity divided by (1 minus specificity); LR- equals (1 minus sensitivity) divided by specificity. An LR+ greater than 10 or an LR- less than 0.1 is very useful. Post-test odds equals pre-test odds times the likelihood ratio (Bayes).
- The ROC curve plots SENSITIVITY against (1 minus SPECIFICITY) for every threshold. AREA UNDER THE CURVE equals 0.5 is no better than chance, 1.0 is perfect, greater than 0.9 is excellent. Moving the threshold trades sensitivity for specificity along the curve.
- SCREENING uses a HIGH-SENSITIVITY threshold in an ASYMPTOMATIC population to avoid missing cases (D-dimer for PE, lactate for sepsis) and accepts false positives. CONFIRMATION uses a HIGH-SPECIFICITY threshold in a SYMPTOMATIC patient to avoid false alarms (biopsy, angiography).
- Evaluate a new test against a reference standard under the STARD guidelines. Guard against SPECTRUM BIAS (the study population does not match your patient) and VERIFICATION BIAS (only positive tests receive the gold standard). Mallampati has low sensitivity AND low specificity, which is why it poorly predicts difficult laryngoscopy.
Meet the patient
A 35-year-old woman presents with pleuritic chest pain and a Wells score that puts her at moderate risk of pulmonary embolism. The D-dimer comes back elevated. Does she have a PE? The answer depends not on the D-dimer alone but on what a positive result means in a patient at her pre-test probability — and that is a question about sensitivity, specificity, and prevalence.[1]
The two questions every test result forces are: how good is the test at detecting disease when it is there? (sensitivity) and how likely is this positive result to be true in this patient? (predictive value). Get those two straight and the rest of the framework follows.[1]
References8ShowHide
- [1]Rubegni G, et al. Automated iridocorneal angle classification using a multimodal large language model Graefes Arch Clin Exp Ophthalmol, 2026.PMID 42363984
- [2]Sandrasegaran K, et al. Do hepatic and biliary functional data from gadoxetate-enhanced MRI add value in predicting outcomes in primary sclerosing cholangitis (PSC)? Eur Radiol, 2026.PMID 42363965
- [3]Albanese A, et al. A Novel Algorithm for Continuous Real-Time Cerebral Autoregulation Assessment Based on Mean Arterial Pressure and Cerebral Oxygen Saturation Anesth Analg, 2026.PMID 42363900
- [4]Lima-Cabrita A, et al. Deep learning in glaucoma referral: Performance assessment using a real-world setting Acta Ophthalmol, 2026.PMID 42363827
- [5]Xu T, et al. The diagnostic and predictive value of AI-combined multilayer spiral CT for MACE after emergency PCI in STEMI patients: A prospective cohort study Medicine (Baltimore), 2026.PMID 42363530
- [6]Özdemir BG, Özdemir H Bladder endometriosis: a current overview of pathogenesis, diagnosis, and treatment approaches Abdom Radiol (NY), 2026.PMID 42364038
- [7]Li JZ, et al. Generative AI in healthcare: redefining clinical practice through digital transformation Health Econ Rev, 2026.PMID 42364007
- [8]Li P, et al. Regional versus general anesthesia for femur and hip fracture surgery: A meta-analysis of postoperative outcomes and complications J Int Med Res, 2026.PMID 42363795