The Observer XT 17 - Reliability Analysis - Frequently Asked Questions About Reliability Analy
Last updated: Jul 28, 2026
Frequently Asked Questions About Reliability Analysis
Why Do I Get Different Kappa Values with Different Comparison Methods?
Different comparison methods have different ways to calculate the number of agreements (A) and disagreements (D). A and D are used to calculate Kappa. See:
- The Frequency/Sequence Method in Detail
- The Duration/Sequence Method in Detail
- The Frequency Method in Detail
- The Duration Method in Detail
- Statistics
Why Are the Combined Kappa and Average Kappa Not the Same?
- The Combined kappa is the kappa of the sum of all agreements and disagreements in the observation pairs.
- Average kappa is the average of all kappas of the observation pairs.
These values differ in the same way as the average of all values in a number of samples differs from the mean of the averages of the individual samples. When there is a bias in one observation pair, this has a larger effect on the Average kappa than on the Combined kappa.
Why Is Rho High While Kappa Is Low?
Cohen's Kappa and Pearson's Rho do not measure the same thing. Rho measures correlation between observations while Kappa measures agreement between observations. A high Rho could result even when agreement measured by Kappa is low.
Example: Events are coded by means of predefined numerical modifiers (10, 20, 30...). If one coder consistently codes 10 points higher than the other coder, the correlation between observations is high, while the agreement is low. So Rho is high, but Kappa is low (and will even be negative).
Why Does the Frequency Method Give Different Statistics Compared to the Frequency/Sequence Method When I Have Only Point Events?
With the frequency method, the total number of each event in both observations is compared. With the frequency/sequence method you investigate event pairs instead of the number of events.
Consider the example of two observations with the only difference that at the same time behavior B was scored in observation 1 and behavior A was scored in observation 2. The behaviors A and B are point events in the same start-stop group.
Frequency/Sequence Method
With the frequency/sequence method, the behaviors A and B are compared with each other, which leads to 1 disagreement.
Frequency Method
With the frequency method, the total number of occurrences of behavior A is different between the two observations, and the total number of occurrences of behavior B is different as well. This results in two disagreements. It also adds an extra count to the total number, which is the sum of all agreements and disagreements.
See:
- The Frequency/Sequence Method in Detail
- The Frequency Method in Detail
Why Is the Total Agreement Duration Longer Than My Observation Duration?
The total agreement duration can be longer than the observation duration in the following cases:
- Your observations contain overlapping behaviors. You defined a behavior group in which behaviors can overlap, or you have multiple behavior groups.
- You used the Duration/Sequence or the Duration comparison method.
Consider the example below with a behavior group in which behaviors A, B, and C may overlap. The three behaviors each overlap for 10 seconds in both observations. Both observations have a duration of 15 s.
The Duration/Sequence method calculates three agreements of 10 seconds, which is a total of 30 seconds. This is longer than the observation duration. Similarly, the Duration method calculates a total duration of 10 seconds for all three behaviors in both observations, also resulting in a total agreement of 30 seconds.
See also:
- The Duration/Sequence Method in Detail
- The Duration Method in Detail
- Effect of Behavior Group - Duration/Sequence Method
Why Is the Agreement Much Higher Than I Expect?
If you have many gaps between the events and select the option Analyze gaps between events, you may get an unrealistically high agreement. This is because gaps are treated as events, so the gaps in the observation pairs are compared with each other and counted as agreements. Therefore, if there are many gaps, they bias the statistics towards agreements. In this example the Duration/Sequence method was used.
To solve this, deselect Analyze gaps between events. The Reliability Analysis would then result in only the two disagreements.
Why Does Rho Change for Old Pairs When I Add New Pairs?
If you add more observation pairs to the list of pairs and run the analysis again, it could happen that the Rho value for the pre-existing pairs differs slightly from those obtained in the previous analysis. This happens because the confusion matrix, which is used to calculate the statistics, always contains all event types in the observations selected, regardless of which event types were scored in a specific observation. When you add more observations to the list of pairs, this could result in adding event types to the confusion matrix that were not scored in the pre-existing observations. Additional event types means extra rows and columns. Because of the way Rho is computed (see Pearson's Rho), this changes the values of Rho in the Statistics page. If you want to keep the confusion matrix fixed, select the Show Not Scored Elements option (see Choose Table Layout Options).
Why Is an Event Paired with Multiple Events in the Comparison List?
The comparison list is shown if you use the Frequency/Sequence comparison method. In some cases, an event is associated with two or more events (not necessarily of the same type) in the other observation. For example, 4.24 Walk in observation 1 is associated with 6.00 Walk (agreement) and 0.00 Walk in observation 2 (disagreement).
This is an inevitable consequence of the fact that one observation contains more events than the other. In such cases only one agreement is counted (usually the one found in run 1 of the algorithm; see The Frequency/Sequence Method in Detail). All other associations are scored as disagreements, no matter if the event type is the same.
Source: The Observer XT 17 Help - Reliability Analysis, Noldus Information Technology