View Categories

Paired or within person tests

When we have more than one data point per person, typically from the same measure on more than one occasion, testing for a difference between the scores is a “paired” or “within person” test. This is different from comparing scores between different people and you can assume that a lot of things that might cause scores to be different between people (gender, age, employment, history of previous problems) are staying the same when scores from one person are compared. This means that statistical analyses for these tests are different from those (“between groups” tests) to be used when comparing between groups.

The term “within subjects test” was widely used but is rightly deprecated, perpetuating the stereotype of the passive “subject” and the paradigm of the experiment rather than, more often the case in our field, survey data.

Details #

When we want to know whether, in general, clients’ scores on some change measure at the end of an episode is different from their starting score we have a paired test, when we want to test whether clients’ baseline scores are, in general, different depending on their employment status we have a between groups test. The typical Null Hypothesis Significance Test (NHST) for the first question is the paired t-test whether the the score change values across your, maybe 55 clients, taking into account the variance of those change values, is sufficiently different from zero that you reject the null hypothesis that the mean change across therapy in a theoretical, infinitely large, population from which those 55 clients came is zero.

For the second question, perhaps using data from the same dataset, we might have 35 clients who were in employment and 20 who weren’t and the paradigmatic test is the unpaired t-test which tests whether the mean baseline score from those in employment is sufficiently different from the mean for those not in employment, given the variances of those scores in each group is such that you can reject the null hypothesis that employment status has any relationship to baseline scores in, again, a theoretical, infinite, population of new clients.

That is the logic of Null Hypothesis Significance Testing. One key issue about paired versus unpaired tests is that for the first question the answer turns on just three numbers: the mean change, the variance of that change and the number of clients in the dataset. For the second question the answer turns on six numbers: the mean baseline score in each group, the variance of those scores in each group and the group sizes.

There are minutiae about using the t-test as the typical pedagogical test to choose and fortunately the research world is moving away from too automatic use of the NHST paradigm and more towards estimation and confidence intervals, reporting effect sizes such as Cohen’s d and Hedges’s g, and complementing all tests and estimation with simple descriptive and graphical methods. However, this distinction between paired and unpaired scenarios remains important and, when we move beyond having just two values per participant to say looking at sessional scores, this takes us into the whole, glorious, world of multi-level models!

Try also … #

Chapters #

Chapter 3 in the OMbook.

Online support #

In the first version of this entry I said: “I hope to put up some simulations illustrating the distinction between paired and between person tests, and, I hope, an app allowing people to put in change data and get a full set of change analyses.”
I haven’t done the simulations yet but I one of my shiny apps allows you to upload paired YP-CORE data and get analyses of your change (and other things). I must get on and add an app that does similar analyses for paired scores of any sort.

Dates #

Created 9.viii.21, expanded and more links added 28.ix.26.

Powered by BetterDocs