ROC Curve and AUC Calculator
Build a receiver operating characteristic (ROC) curve from the scores of subjects with and without a condition, with the AUC, its interval and the Youden-optimal threshold.
Formula
- area under the ROC curve: 0.5 is chance, 1 is perfect
- sensitivity + specificity − 1, maximised at the suggested threshold
How it works
A diagnostic test or a model that gives a score is turned into a yes-or-no call by a threshold. A ROC curve plots sensitivity against 1 − specificity for every possible threshold. The area under it, the AUC, equals the probability that a randomly chosen subject with the condition scores higher than one without, which is the Mann-Whitney statistic divided by the number of pairs.
The standard error of the AUC is found by the Hanley-McNeil method, from which the interval follows. The threshold that maximises the sum of sensitivity and specificity is shown, but it is optimistic, because it is chosen on the same data it is judged on. The best threshold in practice depends on the costs of missing a case against a false alarm, and on the prevalence of the condition.
Worked example
Six subjects with the condition (scores 0.9, 0.8, 0.75, 0.6, 0.55, 0.4) and seven without (0.5, 0.45, 0.3, 0.6, 0.2, 0.1, 0.35).
- In 37.5 of the 42 pairs the score of the subject with the condition is higher (one tie counted as ½): AUC = 0.893.
- Hanley-McNeil SE = 0.099, so the 95% interval is 0.699 to 1.000.
- The threshold 0.55 gives sensitivity 83.3% and specificity 85.7%.
AUC 0.893 (0.699 to 1.000): good separation, but estimated from only thirteen subjects.
These are the values the calculator opens with, so you can check its output against this example.
Assumptions
- The scores come from a defined reference standard for who has the condition.
- Independent subjects, each scored once.
- A higher score means the condition is more likely; the direction is selectable.
Common mistakes
- Reporting the optimal threshold as if it would perform the same on new data.
- Using AUC alone to judge a test, since it ignores calibration and prevalence.
- Evaluating with a small sample and overlooking the width of the interval.
Related tools
Related equipment
Service documentation, failure modes and parts for the instruments this calculation is used with.