top of page

Wilcoxon and Mann-Whitney tests | AQA A-Level Psychology Revision

Updated: 5 days ago

For 7182 specification, first teach in September 2025


AQA A-Level Psychology | Free Revision Notes

Estimated study time: 55 minutes

This Wilcoxon and Mann-Whitney tests A-Level Psychology revision page explains how psychologists select an inferential test for a difference using ordinal data. You will learn that Wilcoxon is used when the two sets of scores are related, while Mann-Whitney is used when they are unrelated. Both tests involve ranked or ordered data, so the experimental design is the crucial distinction. This lesson builds on choosing an inferential test and levels of measurement.


Learning Objectives 🎯

By the end of this revision page, you should be able to:

  • Identify when the Wilcoxon test should be used.

  • Identify when the Mann-Whitney test should be used.

  • Distinguish related from unrelated data.

  • Recognise ordinal data within a research scenario.

  • Explain the similarities and differences between the two tests.

  • Select and justify the correct test from an unfamiliar psychological investigation.


Revision Notes 📚


Wilcoxon and Mann-Whitney tests in A-Level Psychology

Wilcoxon and Mann-Whitney are inferential statistical tests used to investigate a difference.

Both tests are appropriate when the results are measured at the ordinal level.

The difference between them concerns the relationship between the two sets of scores:

  • Wilcoxon is used with related data.

  • Mann-Whitney is used with unrelated data.

The central rules are:

$$\text{Difference}+\text{related data}+\text{ordinal data}=\text{Wilcoxon}$$

$$\text{Difference}+\text{unrelated data}+\text{ordinal data}=\text{Mann-Whitney}$$

The choice between the tests is therefore determined by the experimental design.


Both tests are tests of difference

Before selecting either Wilcoxon or Mann-Whitney, establish that the researcher is investigating a difference.

A test of difference compares the results produced by:

  • two conditions;

  • two groups;

  • two occasions;

  • two treatments.

Examples include investigating whether:

  • anxiety ratings differ before and after therapy;

  • two groups give different ratings of a treatment;

  • concentration scores differ between silence and noise;

  • participants rate two teaching methods differently.

Words suggesting a test of difference include:

  • difference;

  • higher;

  • lower;

  • more;

  • fewer;

  • increase;

  • decrease;

  • effect;

  • compare.


Example of a difference hypothesis

A psychologist might predict:

There will be a difference in participants’ anxiety ratings before and after a relaxation activity.

This is a test of difference because the researcher compares two sets of anxiety ratings.

The researcher must then determine:

  1. Whether the ratings are related or unrelated.

  2. Whether the data are ordinal.


When neither test is appropriate

Wilcoxon and Mann-Whitney are not suitable when the researcher is investigating:

  • a correlation between two co-variables;

  • an association between categorical variables;

  • related nominal signs;

  • interval data requiring a \(t\)-test.

For example:

There will be a correlation between stress rank and sleep-quality rank.

This is a test of correlation, so Spearman’s rho would be appropriate rather than Wilcoxon or Mann-Whitney.

Correlation tests are covered in Spearman’s rho and Pearson’s r.


Ordinal data


What are ordinal data?

Ordinal data can be placed into a meaningful order or rank, but the intervals between values cannot be assumed to be equal.

Examples include:

  • finishing positions;

  • ranks;

  • ordered categories;

  • Likert-type ratings;

  • scores converted into ranks;

  • ratings such as low, moderate and high.

Suppose participants rate their anxiety using the following scale:

Rating

Meaning

\(1\)

Very low anxiety

\(2\)

Low anxiety

\(3\)

Moderate anxiety

\(4\)

High anxiety

\(5\)

Very high anxiety

The ratings have a meaningful order:

$$1<2<3<4<5$$

However, the psychological difference between ratings:

$$1\text{ and }2$$

is not necessarily equal to the difference between:

$$4\text{ and }5$$

The data are therefore ordinal.


Ranked data

Participants may be placed in order according to their performance.

For example:

Participant

Rank

A

\(1\)

B

\(2\)

C

\(3\)

D

\(4\)

A rank of \(1\) may represent the highest score.

A rank of \(4\) may represent the lowest score.

The ranks show order, but they do not communicate equal differences between participants’ performances.


Scores converted into ranks

A researcher may originally collect numerical scores and then convert them into ranks for analysis.

For example:

Raw score

Rank

\(18\)

\(1\)

\(15\)

\(2\)

\(11\)

\(3\)

\(8\)

\(4\)

Once the analysis uses the ranked positions, the values entered into the test are ordinal.

Wilcoxon or Mann-Whitney may then be appropriate, depending on whether the two sets of data are related or unrelated.


Related and unrelated data


Why experimental design matters

Wilcoxon and Mann-Whitney have the same:

  • research purpose;

  • required level of measurement.

Both are used for:

$$\text{A test of difference using ordinal data}$$

The choice between them is made by identifying the experimental design.

The relevant designs are covered in experimental designs.


Related data

Related data contain scores that are meaningfully paired.

Related scores are produced by:

  • repeated measures designs;

  • matched pairs designs.

Each score in one condition has a corresponding score in the other condition.


Repeated measures designs

In a repeated measures design, the same participants take part in both conditions.

For example, participants rate their anxiety:

  • before a relaxation activity;

  • after the relaxation activity.

Each participant provides two scores.

Participant A’s before score is paired with Participant A’s after score.

Participant B’s before score is paired with Participant B’s after score.

The scores are related because they come from the same people.


Example of related ordinal data

Participant

Anxiety before relaxation

Anxiety after relaxation

A

\(5\)

\(3\)

B

\(4\)

\(2\)

C

\(3\)

\(3\)

D

\(5\)

\(4\)

E

\(4\)

\(1\)

The researcher is testing for a difference between the before and after ratings.

The ratings are ordinal.

The same participants provide both sets of scores, so the data are related.

The appropriate test is:

$$\boxed{\text{Wilcoxon}}$$


Matched pairs designs

In a matched pairs design, each participant in one condition is paired with a similar participant in the other condition.

Participants may be matched according to:

  • age;

  • ability;

  • initial anxiety score;

  • memory performance;

  • another relevant characteristic.

Each matched pair produces two related scores.

Although the scores come from different people, they are treated as related because the participants were deliberately matched.


Example of matched-pairs data

A psychologist compares two therapies.

Participants are matched according to their initial anxiety ratings.

One participant from each pair receives Therapy A, while the other receives Therapy B.

Their final satisfaction ratings are ordinal.

The study investigates a difference, the scores are related through matching, and the data are ordinal.

The appropriate test is:

$$\boxed{\text{Wilcoxon}}$$


Unrelated data

Unrelated data come from participants or cases that are not meaningfully paired.

An independent groups design normally produces unrelated data.

For example:

  • one group completes Condition A;

  • a different group completes Condition B.

No participant takes part in both conditions.

No participant in one condition has been deliberately matched with a particular participant in the other condition.


Example of unrelated ordinal data

Participant in Therapy A

Satisfaction rating

Participant in Therapy B

Satisfaction rating

A

\(5\)

F

\(3\)

B

\(4\)

G

\(2\)

C

\(4\)

H

\(4\)

D

\(5\)

I

\(3\)

E

\(3\)

J

\(1\)

The researcher compares the ratings of two separate groups.

The satisfaction scores are ordinal.

The groups are unrelated because different participants take part in each condition.

The appropriate test is:

$$\boxed{\text{Mann-Whitney}}$$


Related and unrelated designs compared

Feature

Related data

Unrelated data

Same participants complete both conditions

Yes

No

Participants are deliberately matched

Possibly

No

Different independent groups

No

Yes

Each score has a corresponding paired score

Yes

No

Ordinal difference test

Wilcoxon

Mann-Whitney


The Wilcoxon test


When should Wilcoxon be used?

The Wilcoxon test should be used when:

  1. The researcher is testing for a difference.

  2. The two sets of data are related.

  3. The data are measured at the ordinal level.

A useful selection rule is:

$$\text{Difference}+\text{related}+\text{ordinal}=\text{Wilcoxon}$$


Research designs suitable for Wilcoxon

Wilcoxon may be used with:

  • repeated measures designs;

  • matched pairs designs.

Both produce paired data.


Worked example: repeated measures

A psychologist investigates whether mindfulness training changes participants’ ratings of examination stress.

Each participant rates their stress:

  • before the training;

  • after the training.

The scale ranges from very low stress to very high stress.


Purpose

The researcher compares stress ratings before and after training:

$$\text{Test of difference}$$


Design

The same participants provide both ratings:

$$\text{Related data}$$


Level of measurement

The stress ratings are ordered:

$$\text{Ordinal data}$$


Test

$$\boxed{\text{Wilcoxon test}}$$


Writing a complete Wilcoxon justification

A complete answer might state:

The Wilcoxon test is appropriate because the researcher is investigating a difference, the same participants provide scores in both conditions so the data are related, and the rating-scale data are ordinal.

All three elements are needed for a full justification:

  • difference;

  • related;

  • ordinal.


Worked example: matched pairs

A psychologist compares the effectiveness of two revision techniques.

Participants are matched according to their previous examination grades. One member of each pair uses Technique A, while the other uses Technique B.

Afterwards, participants rate the usefulness of the technique from very unhelpful to very helpful.


Purpose

The researcher compares the ratings for two techniques:

$$\text{Test of difference}$$


Design

Participants are matched:

$$\text{Related data}$$


Level of measurement

The usefulness ratings are ordered:

$$\text{Ordinal data}$$


Test

$$\boxed{\text{Wilcoxon test}}$$


Wilcoxon is not selected only because the same participants are used

Repeated measures establishes that the data are related, but the level of measurement must still be identified.

For related data:

  • nominal signs require the sign test;

  • ordinal data require Wilcoxon;

  • interval data require the related \(t\)-test.

Therefore, repeated measures does not automatically mean Wilcoxon.


Wilcoxon and the sign test

Wilcoxon and the sign test can both be used with related data and tests of difference.

The distinction is the level of measurement.

Feature

Sign test

Wilcoxon

Purpose

Difference

Difference

Data relationship

Related

Related

Level of measurement

Nominal signs

Ordinal

Information used

Direction of difference

Ranked information

Example

Increase or decrease categories

Ordered anxiety ratings

The sign test is covered in the sign test.


Wilcoxon and the related \(t\)-test

Wilcoxon and the related \(t\)-test both examine differences using related scores.

The distinction is again the level of measurement.

Feature

Wilcoxon

Related \(t\)-test

Purpose

Difference

Difference

Data relationship

Related

Related

Level of measurement

Ordinal

Interval

Example

Ordered satisfaction ratings

Response times in seconds

The related \(t\)-test is covered in related and unrelated t-tests.


The Mann-Whitney test


When should Mann-Whitney be used?

The Mann-Whitney test should be used when:

  1. The researcher is testing for a difference.

  2. The two sets of data are unrelated.

  3. The data are measured at the ordinal level.

A useful selection rule is:

$$\text{Difference}+\text{unrelated}+\text{ordinal}=\text{Mann-Whitney}$$


Research designs suitable for Mann-Whitney

Mann-Whitney is used with an independent groups design.

Different participants take part in each condition or group.


Worked example: independent groups

A psychologist investigates whether two types of music produce different concentration ratings.

  • Group A listens to instrumental music.

  • Group B listens to music containing lyrics.

Participants rate their concentration from very poor to very good.


Purpose

The researcher compares the ratings produced by two conditions:

$$\text{Test of difference}$$


Design

Different participants take part in each condition:

$$\text{Unrelated data}$$


Level of measurement

The concentration ratings are ordered:

$$\text{Ordinal data}$$


Test

$$\boxed{\text{Mann-Whitney test}}$$


Writing a complete Mann-Whitney justification

A complete answer might state:

The Mann-Whitney test is appropriate because the researcher is investigating a difference, different participants take part in the two conditions so the data are unrelated, and the concentration ratings are ordinal.

Again, all three features should be included:

  • difference;

  • unrelated;

  • ordinal.


Worked example: comparing separate groups

A researcher compares the wellbeing ratings of:

  • students who regularly exercise;

  • students who do not regularly exercise.

Different participants are included in the two groups.

Wellbeing is rated as low, moderate or high.


Purpose

The researcher compares two groups:

$$\text{Test of difference}$$


Design

The groups contain different participants:

$$\text{Unrelated data}$$


Level of measurement

Low, moderate and high form an ordered scale:

$$\text{Ordinal data}$$


Test

$$\boxed{\text{Mann-Whitney test}}$$


Mann-Whitney is not selected only because different participants are used

Independent groups establish that the data are unrelated, but the level of measurement must also be identified.

For unrelated data:

  • ordinal data require Mann-Whitney;

  • interval data require the unrelated \(t\)-test.

Therefore, an independent groups design does not automatically mean Mann-Whitney.


Mann-Whitney and the unrelated \(t\)-test

Both tests compare two sets of unrelated data.

Feature

Mann-Whitney

Unrelated \(t\)-test

Purpose

Difference

Difference

Data relationship

Unrelated

Unrelated

Level of measurement

Ordinal

Interval

Example

Ordered anxiety ratings

Reaction times in milliseconds

The choice depends on whether the scores are ordinal or interval.


Comparing Wilcoxon and Mann-Whitney


Similarities

Wilcoxon and Mann-Whitney are similar because both:

  • are inferential statistical tests;

  • test for a difference;

  • are used with ordinal data;

  • involve ordered or ranked scores;

  • produce an observed test value;

  • require comparison with a critical value;

  • may be used to determine statistical significance;

  • lead to a decision about the null hypothesis.


Differences

The main difference is the relationship between the scores.

Feature

Wilcoxon

Mann-Whitney

Purpose

Difference

Difference

Level of measurement

Ordinal

Ordinal

Data relationship

Related

Unrelated

Suitable design

Repeated measures or matched pairs

Independent groups

Pairing between scores

Yes

No


The central distinction

Remember:

$$\text{Wilcoxon}=\text{Related ordinal difference}$$

$$\text{Mann-Whitney}=\text{Unrelated ordinal difference}$$

The letter \(W\) in Wilcoxon does not provide a reliable clue by itself. The safest method is always to work through the full selection process.


Selecting the correct test


The three-question method

Use these three questions whenever you are given a research scenario.


Question 1: Is the researcher testing for a difference?

If the researcher is comparing conditions, groups or occasions, the answer is likely to be yes.

If the researcher is examining a relationship between co-variables, neither Wilcoxon nor Mann-Whitney is suitable.


Question 2: Are the scores related or unrelated?

Related scores come from:

  • repeated measures;

  • matched pairs.

Unrelated scores come from:

  • independent groups.


Question 3: Are the data ordinal?

Look for:

  • ranks;

  • ordered ratings;

  • Likert-type responses;

  • categories such as low, moderate and high;

  • scores converted into ranks.

If the data are interval, select a suitable \(t\)-test instead.


Decision sequence

For a test of difference using ordinal data:


Same participants or matched pairs

$$\text{Related data}\rightarrow\text{Wilcoxon}$$


Separate independent groups

$$\text{Unrelated data}\rightarrow\text{Mann-Whitney}$$


Worked selection scenario 1

Participants rate their anxiety before and after a breathing exercise using an ordered scale.


Difference or correlation?

Two occasions are compared:

$$\text{Difference}$$


Related or unrelated?

The same participants provide both ratings:

$$\text{Related}$$


Level of measurement?

The ratings are ordered:

$$\text{Ordinal}$$


Correct test

$$\boxed{\text{Wilcoxon}}$$


Worked selection scenario 2

One group rates a cognitive therapy and another group rates a behavioural therapy.

The groups contain different participants.


Difference or correlation?

The researcher compares ratings of two therapies:

$$\text{Difference}$$


Related or unrelated?

Different participants provide the ratings:

$$\text{Unrelated}$$


Level of measurement?

The ratings are ordinal:

$$\text{Ordinal}$$


Correct test

$$\boxed{\text{Mann-Whitney}}$$


Worked selection scenario 3

Participants rank their levels of stress and rank the quality of their sleep.

The researcher investigates whether the two rankings are related.


Difference or correlation?

The researcher investigates a relationship between two co-variables:

$$\text{Correlation}$$


Correct test

$$\boxed{\text{Spearman's rho}}$$

Neither Wilcoxon nor Mann-Whitney is appropriate because the study does not test a difference.


Worked selection scenario 4

The same participants complete a response-time task before and after sleep deprivation.

Response time is measured in milliseconds.


Purpose

$$\text{Difference}$$


Design

$$\text{Related}$$


Level of measurement

$$\text{Interval}$$


Correct test

$$\boxed{\text{Related }t\text{-test}}$$

Wilcoxon is not appropriate because the data are interval rather than ordinal.


Worked selection scenario 5

Two separate groups complete a memory task.

The researcher records each participant’s numerical score using equal units.


Purpose

$$\text{Difference}$$


Design

$$\text{Unrelated}$$


Level of measurement

$$\text{Interval}$$


Correct test

$$\boxed{\text{Unrelated }t\text{-test}}$$

Mann-Whitney is not appropriate because the scores are interval.


Worked selection scenario 6

Participants are matched according to age and initial confidence rating.

One member of each pair completes a public-speaking workshop, while the other does not.

Afterwards, confidence is rated from very low to very high.


Purpose

The conditions are compared:

$$\text{Difference}$$


Design

Participants are matched:

$$\text{Related}$$


Level of measurement

The ratings are ordered:

$$\text{Ordinal}$$


Correct test

$$\boxed{\text{Wilcoxon}}$$

Do not select Mann-Whitney simply because different people take part in the two conditions. The matched-pairs design makes the scores related.


Interpreting statistical significance


Observed values

Wilcoxon and Mann-Whitney produce an observed value calculated from the data.

The exact symbol used depends on the test and the information provided in the question.

The observed value is compared with a critical value from the relevant statistical table.


Critical values

The critical value may depend on:

  • the number of scores or pairs;

  • the sizes of the groups;

  • the chosen significance level;

  • whether the hypothesis is directional or non-directional;

  • the inferential test being used.

A directional hypothesis requires a one-tailed critical value.

A non-directional hypothesis requires a two-tailed critical value.

The use of tables is covered in probability and significance.


Significance rule

For Wilcoxon and Mann-Whitney, the result is statistically significant when the observed value is equal to or smaller than the relevant critical value.

The general decision rule is:

$$\text{Observed value}\leq\text{Critical value}$$

This means that a smaller observed value provides stronger evidence against the null hypothesis.


Worked Wilcoxon significance example

A Wilcoxon test produces:

$$T_{\text{observed}}=8$$

The statistical table gives:

$$T_{\text{critical}}=10$$

Compare the values:

$$8\leq10$$

The result is statistically significant.

The null hypothesis should be rejected.

A suitable conclusion would be:

The observed value was smaller than the critical value, so the difference was statistically significant at the selected level. The null hypothesis was rejected.

Non-significant Wilcoxon example

Suppose:

$$T_{\text{observed}}=14$$

and:

$$T_{\text{critical}}=10$$

Compare the values:

$$14>10$$

The result is not statistically significant.

The null hypothesis should be retained.


Worked Mann-Whitney significance example

A Mann-Whitney test produces:

$$U_{\text{observed}}=9$$

The critical value is:

$$U_{\text{critical}}=9$$

Compare the values:

$$9\leq9$$

The result is statistically significant.

Equality meets the significance criterion.


Non-significant Mann-Whitney example

Suppose:

$$U_{\text{observed}}=15$$

and:

$$U_{\text{critical}}=11$$

Compare the values:

$$15>11$$

The result is not statistically significant.

The null hypothesis should be retained.


Writing a complete statistical conclusion

A complete conclusion should include:

  1. The observed value.

  2. The critical value.

  3. The comparison between the two.

  4. Whether the result is statistically significant.

  5. Whether the null hypothesis is rejected or retained.

  6. A statement referring to the study.

For example:

The observed Mann-Whitney value of \(9\) was equal to the critical value of \(9\). The result was statistically significant at \(p\leq0.05\), so the null hypothesis was rejected. There was a statistically significant difference in anxiety ratings between the two treatment groups.

Statistical significance does not explain the cause

A statistically significant Wilcoxon or Mann-Whitney result shows that a difference meets the selected probability criterion.

It does not automatically show:

  • why the difference occurred;

  • whether the effect is practically important;

  • whether the procedure was valid;

  • whether uncontrolled variables influenced the result.

The statistical conclusion should be interpreted alongside the quality of the research method.


Writing a strong test justification


Wilcoxon sentence structure

Wilcoxon is appropriate because the study tests for a difference, the same participants take part in both conditions so the scores are related, and the data are ordinal.

For a matched pairs design:

Wilcoxon is appropriate because the study tests for a difference, the participants are matched so the scores are related, and the ratings are ordinal.

Mann-Whitney sentence structure

Mann-Whitney is appropriate because the study tests for a difference, different participants take part in the two conditions so the data are unrelated, and the ratings are ordinal.

Why brief answers may lose marks

An answer such as:

Use Wilcoxon because the data are related.

does not identify:

  • the purpose of the analysis;

  • the level of measurement.

An answer such as:

Use Mann-Whitney because the data are ordinal.

does not identify:

  • that the study tests a difference;

  • that the scores are unrelated.

For a complete justification, include all three features.


Complete comparison with other inferential tests

Test

Purpose

Data relationship

Level of measurement

Sign test

Difference

Related

Nominal

Wilcoxon

Difference

Related

Ordinal

Mann-Whitney

Difference

Unrelated

Ordinal

Related \(t\)-test

Difference

Related

Interval

Unrelated \(t\)-test

Difference

Unrelated

Interval

Spearman’s rho

Correlation

Paired co-variable scores

Ordinal

Pearson’s \(r\)

Correlation

Paired co-variable scores

Interval

Chi-squared

Association

Categorical frequencies

Nominal


The ordinal-data decision

When the data are ordinal, ask what the researcher is investigating.


Correlation

$$\text{Spearman's rho}$$


Related difference

$$\text{Wilcoxon}$$


Unrelated difference

$$\text{Mann-Whitney}$$

Ordinal data alone are not enough to select the test.


The related-data decision

When the data are related and the researcher tests a difference, identify the level of measurement.


Nominal signs

$$\text{Sign test}$$


Ordinal scores

$$\text{Wilcoxon}$$


Interval scores

$$\text{Related }t\text{-test}$$


The unrelated-data decision

When the data are unrelated and the researcher tests a difference, identify the level of measurement.


Ordinal scores

$$\text{Mann-Whitney}$$


Interval scores

$$\text{Unrelated }t\text{-test}$$


Key Words 🔑

Key word

Student-friendly definition

How it may be used in an exam

Wilcoxon test

An inferential test used to investigate a difference between two sets of related ordinal data.

You may identify or justify it from a repeated measures or matched pairs scenario.

Mann-Whitney test

An inferential test used to investigate a difference between two sets of unrelated ordinal data.

You may identify or justify it from an independent groups scenario.

Test of difference

An analysis examining whether two conditions, groups or occasions produce different results.

Both Wilcoxon and Mann-Whitney are tests of difference.

Related data

Scores connected through the same participants or deliberately matched participants.

Related ordinal data require Wilcoxon.

Unrelated data

Scores produced by separate participants who are not meaningfully paired.

Unrelated ordinal data require Mann-Whitney.

Repeated measures

An experimental design in which the same participants take part in both conditions.

It produces related data.

Matched pairs

A design in which participants in different conditions are paired on relevant characteristics.

It also produces related data.

Independent groups

A design in which different participants take part in each condition.

It produces unrelated data.

Ordinal data

Data that can be placed in order but do not necessarily contain equal intervals.

Both tests require ordinal data.

Rank

A position within an ordered sequence.

Scores may be converted into ranks before analysis.

Observed value

The test statistic produced from the research data.

It is compared with the critical value.

Critical value

The threshold obtained from the relevant statistical table.

It is used to determine whether the result is significant.

Statistical significance

A judgement that a result meets the selected probability criterion.

For both tests, the observed value must be equal to or smaller than the critical value.

Null hypothesis

A prediction that there is no difference and that any apparent result occurred through chance.

It is rejected if the statistical result is significant.

Directional hypothesis

A hypothesis predicting the direction of a difference.

It requires a one-tailed critical value.

Non-directional hypothesis

A hypothesis predicting a difference without stating its direction.

It requires a two-tailed critical value.

Hints from the Examiner Reports 💡

No lesson-specific examiner guidance was identified in the provided reports.


Common Mistakes ⚠️


Mistake: Selecting Wilcoxon for any repeated measures study.

Why this is incorrect:

Repeated measures only establishes that the scores are related. The level of measurement must also be ordinal.

How to improve:

Check whether the data are nominal, ordinal or interval before selecting the test.


Mistake: Selecting Mann-Whitney for every independent groups study.

Why this is incorrect:

Independent groups establish that the data are unrelated, but Mann-Whitney specifically requires ordinal data.

How to improve:

If the unrelated data are interval, select the unrelated \(t\)-test.


Mistake: Reversing Wilcoxon and Mann-Whitney.

Why this is incorrect:

Wilcoxon requires related data, while Mann-Whitney requires unrelated data.

How to improve:

Remember:

$$\text{Wilcoxon}=\text{Related ordinal difference}$$

$$\text{Mann-Whitney}=\text{Unrelated ordinal difference}$$


Mistake: Treating matched-pairs scores as unrelated.

Why this is incorrect:

The participants have been deliberately paired, so each score has a corresponding matched score.

How to improve:

Classify both repeated measures and matched pairs as related designs.


Mistake: Selecting a correlation test for before-and-after scores.

Why this is incorrect:

A before-and-after investigation normally compares two occasions rather than examining a relationship between two co-variables.

How to improve:

Identify whether the aim predicts a difference or a correlation.


Mistake: Selecting Spearman’s rho whenever the data are ordinal.

Why this is incorrect:

Spearman’s rho is used for a correlation. Wilcoxon and Mann-Whitney are used for differences.

How to improve:

Identify the research purpose before using the measurement level.


Mistake: Describing Likert-type ratings as interval solely because numbers are used.

Why this is incorrect:

The values represent ordered responses, but equal psychological differences between adjacent ratings cannot automatically be assumed.

How to improve:

Identify ordered ratings as ordinal unless the scenario provides a justified equal-interval measure.


Mistake: Justifying Wilcoxon only by saying that the same participants were used.

Why this is incorrect:

A complete justification must also identify a test of difference and ordinal data.

How to improve:

Use all three features:

$$\text{Difference}+\text{related}+\text{ordinal}$$


Mistake: Justifying Mann-Whitney only by saying that the groups are different.

Why this is incorrect:

A complete justification must also identify the purpose and measurement level.

How to improve:

Use:

$$\text{Difference}+\text{unrelated}+\text{ordinal}$$


Mistake: Using the greater-than rule when interpreting significance.

Why this is incorrect:

For Wilcoxon and Mann-Whitney, a result is significant when the observed value is equal to or smaller than the critical value.

How to improve:

Remember:

$$\text{Observed}\leq\text{Critical}$$


Mistake: Treating equal observed and critical values as non-significant.

Why this is incorrect:

Equality means that the observed value has reached the critical threshold.

How to improve:

Include the equality symbol:

$$\leq$$


Mistake: Selecting a test from the data alone.

Why this is incorrect:

The same ordinal data could require Spearman’s rho, Wilcoxon or Mann-Whitney.

How to improve:

Always identify:

  1. Difference, correlation or association.

  2. Related or unrelated scores.

  3. Level of measurement.


Exam-Style Questions ✍️


Question 1

State when the Wilcoxon test should be used.[3 marks]


Question 2

State when the Mann-Whitney test should be used.[3 marks]


Question 3

Explain one similarity and one difference between Wilcoxon and Mann-Whitney.[4 marks]


Question 4

The same participants rate their anxiety before and after a mindfulness activity. Anxiety is rated using an ordered scale.

Identify an appropriate inferential statistical test. Explain your answer.[4 marks]


Question 5

One group rates the usefulness of Therapy A. A separate group rates the usefulness of Therapy B.

The ratings range from very unhelpful to very helpful.

Identify an appropriate inferential statistical test. Justify your answer.[4 marks]


Question 6

Participants are matched according to their initial memory scores. One member of each pair uses Revision Method A, while the other uses Revision Method B.

Participants then rank how useful they found the method.

Identify an appropriate inferential statistical test. Explain your answer.[4 marks]


Question 7

A psychologist investigates whether participants’ ranked stress scores are related to their ranked sleep-quality scores.

A student selects Wilcoxon because the data are ordinal.

Explain why Wilcoxon is not appropriate and identify the correct test.[4 marks]


Question 8

The same participants complete a reaction-time task before and after training. Reaction time is measured in milliseconds.

A researcher proposes using Wilcoxon.

Explain why Wilcoxon is not the most appropriate test and identify the correct test.[4 marks]


Question 9

Two independent groups complete a memory task. Participants receive ordinal performance ratings.

A student proposes using the unrelated \(t\)-test.

Explain why this test is inappropriate and identify the correct test.[4 marks]


Question 10

For each research scenario, identify the appropriate inferential test.

a) A difference between related ordinal ratings.[1 mark]

b) A difference between unrelated ordinal ratings.[1 mark]

c) A correlation between two ordinal co-variables.[1 mark]

d) A difference between related interval scores.[1 mark]

e) A difference between unrelated interval scores.[1 mark]


Question 11

A Wilcoxon test produces:

$$T_{\text{observed}}=7$$

The critical value is:

$$T_{\text{critical}}=9$$

For Wilcoxon, the observed value must be equal to or smaller than the critical value.

a) Determine whether the result is statistically significant.[1 mark]

b) Explain your answer.[2 marks]

c) State what should happen to the null hypothesis.[1 mark]


Question 12

A Mann-Whitney test produces:

$$U_{\text{observed}}=15$$

The critical value is:

$$U_{\text{critical}}=12$$

a) Determine whether the result is statistically significant.[1 mark]

b) Explain your answer using the observed and critical values.[2 marks]

c) State an appropriate conclusion about the null hypothesis.[1 mark]


Question 13

A psychologist compares the ordinal confidence ratings of two independent groups after they complete different public-speaking programmes.

The Mann-Whitney test produces:

$$U_{\text{observed}}=8$$

The critical value at:

$$p\leq0.05$$

is:

$$U_{\text{critical}}=8$$

Write a complete statistical conclusion. Your answer should:

  • compare the observed and critical values;

  • state whether the result is statistically significant;

  • state whether the null hypothesis should be rejected or retained;

  • refer to confidence ratings in the two programmes.

[4 marks]

Recent Posts

See All

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page