The Observer XT 17 - Reliability Analysis - Reliability Analysis Statistics
Last updated: Jul 28, 2026
Reliability Analysis Statistics
Significance of Rho
To test the significance of ρ, a standard score t is calculated. A one-tailed test is carried out, and the probability is shown next to ρ.
The test of significance of r is based on the assumption that the distribution of the residual values (that is, the deviations of the column totals from the regression over the row totals) follows the normal distribution, and that the variability of the residual values is the same for all values of the independent variable. However, Monte Carlo simulations suggest that meeting those assumptions closely is not absolutely crucial if your sample size is not very small and when the departure from normality is not very large. If the number of rows and columns of your confusion matrix is 50 or more then serious biases are unlikely, and if it is over 100 then you do not need to be concerned with the normality assumptions.
Prevalence Index
Prevalence index is the degree to which a particular event occurs more in a group of subjects than another event. For example, when the number of agreements in Reliability Analysis in one behavior group is higher than in another group, the Prevalence index is high but the Kappa is low. The Prevalence index is therefore useful to explain odd Kappa values.
Please note that special columns in the Confusion Matrix, such as No Records, Window Error and Total are not taken into account in the calculation of the Prevalence index.
Confidence Interval
The 95% confidence interval low and high values are given in the Confusion Matrix. First, the standard error of Kappa is calculated, where Ao is the observed proportion of agreements, Ae is the proportion of agreements expected by chance, and n is the sum of the column totals and the number of columns. The confidence interval is then calculated from this standard error.
Extra Kappa Statistics
Three extra statistics are available if you selected the option Show Extra Kappa Statistics in the Reliability Analysis Settings. These are Minimum Kappa, Maximum Kappa and Average Kappa, and represent the summarization of Cohen's Kappa for all pairs.
Note that the Average Kappa is not the same as the Kappa obtained by summing up the agreements and disagreements from two or more pairs of observations (see Combine Multiple Pairs of Observations), although those two values are often correlated.