Introduction to statistical testing | AQA A-Level Psychology Revision
- Revision Notes
- Aug 5
- 15 min read
Updated: 6 days ago
For 7182 specification, first teach in September 2025
AQA A-Level Psychology | Free Revision Notes
Estimated study time: 55 minutes
This Introduction to statistical testing A-Level Psychology revision page explains why psychologists use inferential statistical tests and how they decide whether a result is statistically significant. You will learn to distinguish an observed value from a critical value and understand how the two are compared. These skills build on levels of measurement and prepare you for calculating the sign test, interpreting probability and selecting appropriate inferential tests.
Learning Objectives 🎯
By the end of this revision page, you should be able to:
Explain the purpose of inferential statistical testing.
Distinguish descriptive statistics from inferential statistics.
Define observed and critical values.
Explain how a critical value is found using a statistical table.
Determine whether a result is statistically significant.
Write an appropriate conclusion from the outcome of a statistical test.
Revision Notes 📚
Introduction to statistical testing in A-Level Psychology
Psychological researchers usually collect data from a sample rather than from every member of the target population.
The sample may appear to show:
a difference between two conditions;
a difference between two groups;
an association between categories;
a correlation between two co-variables.
However, an apparent pattern in a sample might have occurred because of chance variation.
Inferential statistical testing helps a researcher decide whether the result is sufficiently unlikely to have occurred by chance for it to be considered statistically significant.
The specification requires students to understand:
the introduction to statistical testing;
the use of statistical tables;
critical values;
the interpretation of statistical significance;
the factors affecting statistical test choice.
Detailed work on probability and significance follows in probability and significance.
Descriptive and inferential statistics
Psychologists use both descriptive statistics and inferential statistics, but they serve different purposes.
Descriptive statistics
Descriptive statistics summarise and describe the data collected in a study.
Examples include:
the mean;
the median;
the mode;
the range;
standard deviation;
percentages;
correlation coefficients.
For example, a researcher may find that the mean memory score was:
$$14.6$$
in one condition and:
$$11.2$$
in another condition.
The difference between the means is:
$$14.6-11.2=3.4$$
This describes the difference found in the sample. It does not establish whether the difference is statistically significant.
Inferential statistics
Inferential statistics allow researchers to analyse sample results and make a judgement about whether the pattern is likely to reflect something beyond chance variation.
An inferential test produces an observed value. This is compared with a critical value obtained from an appropriate statistical table.
The comparison helps the researcher determine whether the result is statistically significant.
Comparing descriptive and inferential statistics
Feature | Descriptive statistics | Inferential statistics |
Main purpose | Summarise the data collected | Assess whether a result is statistically significant |
Example | Mean, median, mode or range | Sign test, Spearman’s rho or Mann-Whitney |
Describes sample data | Yes | Uses sample data as part of a significance decision |
Assesses whether a result may be due to chance | No | Yes |
Produces an observed test value | No | Yes |
Descriptive statistics usually come before inferential testing. The researcher first organises and summarises the results, then applies an appropriate statistical test.
Why inferential testing is needed
Suppose a psychologist investigates whether background noise affects memory.
The mean scores are:
Condition | Mean number of words recalled |
Silence | \(15.4\) |
Background noise | \(12.7\) |
The difference between the means is:
$$15.4-12.7=2.7$$
The sample performed better in silence, but this difference alone does not show that the effect is statistically significant.
The difference might reflect:
a genuine pattern;
individual variation among participants;
an unusual sample;
random variation;
measurement variation.
Inferential testing provides a systematic method for judging whether the result is sufficiently unlikely to have occurred by chance.
Statistical significance
A result is statistically significant when it meets the criterion established for the selected statistical test and significance level.
This means that the result is judged unlikely to have occurred through chance variation alone.
Statistical significance does not necessarily mean that:
the effect is large;
the finding is important in everyday life;
the research procedure was valid;
the study was free from bias;
one variable caused another;
the result will always be replicated.
It is a statistical judgement about the likelihood of obtaining the result if there were no genuine effect, difference, association or correlation.
Statistical significance is not the same as a large result
A large-looking difference is not automatically statistically significant.
For example, two conditions might have means of:
$$18.2\text{ and }13.1$$
The difference is:
$$18.2-13.1=5.1$$
However, the scores might vary greatly within each condition. The apparent difference could therefore be less convincing than it first appears.
Similarly, a smaller difference may be statistically significant if the data show a consistent pattern.
The significance decision must be based on an appropriate inferential test rather than visual inspection alone.
The null hypothesis
Inferential testing assesses a result in relation to a null hypothesis.
A null hypothesis states that there is no difference, association or correlation, and that any apparent pattern is due to chance.
Examples include:
There will be no difference in the number of words recalled in silence and in background noise.
There will be no correlation between hours of sleep and memory-test score.
There will be no association between treatment type and whether participants improve.
The exact form depends on the research question.
The alternative hypothesis
The alternative hypothesis predicts that there is a difference, association or correlation.
It may be:
directional;
non-directional.
For example:
Participants will recall more words in silence than in background noise.
This is directional because it predicts which condition will produce the higher score.
Alternatively:
There will be a difference in the number of words recalled in silence and in background noise.
This is non-directional because it predicts a difference but does not specify its direction.
The distinction between directional and non-directional hypotheses can affect which critical value is selected from a statistical table. This is explored further in probability and significance.
The observed value
The observed value is the value calculated from the sample data using the selected inferential statistical test.
It may also be described as the calculated test statistic.
Examples of observed values might include:
$$S=2$$
for a sign test, or:
$$r_s=0.72$$
for Spearman’s rho.
The symbol and interpretation depend on the statistical test used.
The observed value is based on the actual results collected by the researcher.
The critical value
The critical value is the threshold that the observed value must meet for the result to be considered statistically significant.
It is found in a statistical table.
The correct critical value depends on information such as:
the statistical test being used;
the number of participants or scores;
the chosen significance level;
whether the hypothesis is directional or non-directional;
degrees of freedom, where required.
The critical value is not calculated directly from the raw results. It is looked up in the appropriate statistical table.
Observed and critical values compared
Value | Meaning | Where it comes from |
Observed value | The test statistic produced from the study’s data | Calculated using the selected statistical test |
Critical value | The threshold used to judge significance | Found in the appropriate statistical table |
A useful distinction is:
The observed value comes from the data. The critical value comes from the table.
Finding a critical value
Although statistical tables differ, the general process is similar.
Identify the correct statistical test.
Identify the number of participants, pairs of scores or degrees of freedom required by the table.
Identify the selected significance level.
Identify whether the hypothesis is directional or non-directional where required.
Locate the relevant row and column.
Read the critical value at their intersection.
Compare it with the observed value.
The table provided in an examination will contain the information needed for that particular question.
Statistical test choice comes first
A researcher cannot select a critical value before deciding which statistical test is appropriate.
The choice of test depends on factors including:
whether the study investigates a difference, association or correlation;
the level of measurement;
whether the scores are related or unrelated.
These factors are covered fully in choosing an inferential test.
Selecting the wrong test would lead to the use of the wrong statistical table and an invalid significance decision.
Comparing the observed and critical values
The observed value must be compared with the critical value according to the rule for the statistical test being used.
For some tests, the result is significant when:
$$\text{Observed value}\geq\text{Critical value}$$
For other tests, the result is significant when:
$$\text{Observed value}\leq\text{Critical value}$$
You must therefore know or be told the comparison rule for the specific test.
Do not assume that every test uses the same rule.
Tests where a larger observed value may be needed
For some inferential tests, a result is significant when the observed value is equal to or greater than the critical value.
The decision rule is:
$$\text{Observed value}\geq\text{Critical value}$$
Worked example
Suppose:
$$\text{Observed value}=16$$
and:
$$\text{Critical value}=12$$
Compare the values:
$$16\geq12$$
The observed value meets or exceeds the critical value.
Therefore:
$$\boxed{\text{The result is statistically significant}}$$
The null hypothesis would be rejected.
Tests where a smaller observed value may be needed
For some tests, including the sign test, a result is significant when the observed value is equal to or smaller than the critical value.
The decision rule is:
$$\text{Observed value}\leq\text{Critical value}$$
Worked example
Suppose a sign test produces:
$$S=2$$
The critical value is:
$$S_{\text{critical}}=3$$
Compare the values:
$$2\leq3$$
The observed value meets the significance criterion.
Therefore:
$$\boxed{\text{The result is statistically significant}}$$
The null hypothesis would be rejected.
Equal values count
If the observed value is exactly equal to the critical value, it has met the threshold for statistical significance.
For a test requiring the observed value to be equal to or greater than the critical value:
$$\text{Observed}=8$$
$$\text{Critical}=8$$
Therefore:
$$8\geq8$$
The result is significant.
For a test requiring the observed value to be equal to or smaller than the critical value:
$$\text{Observed}=4$$
$$\text{Critical}=4$$
Therefore:
$$4\leq4$$
The result is also significant.
A memory rule for test comparisons
For tests where a larger observed value is needed, students sometimes use the reminder:
Observed is greater than or equal to critical.
For tests where a smaller observed value is needed, including the sign test, the direction is reversed:
Observed is less than or equal to critical.
The safest method is always to learn the decision rule for the test rather than applying one comparison rule to every test.
When a result is significant
If the observed value meets the relevant critical-value criterion:
the result is statistically significant;
the null hypothesis is rejected;
the alternative hypothesis is supported;
the researcher concludes that the pattern is unlikely to be due to chance alone.
A suitable conclusion might be:
The observed value was equal to or more extreme than the critical value, so the result was statistically significant. The null hypothesis was rejected.
The conclusion should then refer directly to the variables or conditions in the study.
For example:
There was a statistically significant difference in memory scores between the silence and background-noise conditions.
When a result is not significant
If the observed value does not meet the critical-value criterion:
the result is not statistically significant;
the null hypothesis is retained;
there is insufficient evidence to support the alternative hypothesis;
the apparent pattern may have occurred through chance variation.
A suitable conclusion might be:
The observed value did not meet the critical value, so the result was not statistically significant. The null hypothesis was retained.
The conclusion should not claim that the researcher has proved that there is no effect.
A non-significant result means that the study did not provide sufficient evidence to reject the null hypothesis.
Rejecting and retaining the null hypothesis
Statistical outcome | Decision about the null hypothesis |
Significant result | Reject the null hypothesis |
Non-significant result | Retain the null hypothesis |
In examination answers, use the wording requested by the mark scheme or question.
It is safer to write retain the null hypothesis rather than claiming that the null hypothesis has been proved true.
Worked example: significant result
A researcher uses a statistical test for which the observed value must be greater than or equal to the critical value.
The values are:
$$\text{Observed value}=27$$
$$\text{Critical value}=21$$
Compare them:
$$27\geq21$$
The observed value exceeds the critical value.
The conclusion is:
The result is statistically significant. The null hypothesis should be rejected.
If the study compared two memory conditions, the contextual conclusion might be:
There was a statistically significant difference in memory performance between the two conditions.
Worked example: non-significant result
For the same type of test:
$$\text{Observed value}=18$$
$$\text{Critical value}=21$$
Compare them:
$$18<21$$
The observed value does not reach the critical value.
The conclusion is:
The result is not statistically significant. The null hypothesis should be retained.
Worked example: the sign-test rule
A sign test produces:
$$S=5$$
The critical value is:
$$S_{\text{critical}}=3$$
For the sign test, the observed value must be equal to or smaller than the critical value.
Compare the values:
$$5>3$$
The observed value is not equal to or smaller than the critical value.
Therefore:
$$\boxed{\text{The result is not statistically significant}}$$
The null hypothesis should be retained.
The calculation and interpretation of this test are covered in the sign test.
Statistical significance and chance
Inferential testing does not remove all uncertainty.
Instead, the researcher sets an acceptable probability of obtaining the result through chance if the null hypothesis is true.
A commonly used significance level is:
$$p\leq0.05$$
This means that the probability of obtaining the result through chance is judged to be no more than:
$$5\%$$
Another possible level is:
$$p\leq0.01$$
This sets a stricter criterion.
The interpretation of these values, including one-tailed and two-tailed tests, is covered in probability and significance.
Statistical and practical importance
A statistically significant result is not automatically important in practical terms.
For example, a study might find a statistically significant difference of:
$$0.2\text{ points}$$
between two conditions.
The result may be unlikely to have occurred by chance, but the difference may be too small to matter in a practical setting.
Conversely, a result that appears meaningful may fail to reach statistical significance if the study contains too few participants or highly variable scores.
Statistical significance should therefore be interpreted alongside:
the size of the difference or relationship;
the research design;
reliability;
validity;
the characteristics of the sample.
Statistical significance and causation
A significant correlation does not prove that one co-variable caused the other.
For example, a statistically significant correlation between stress and sleep means that the variables are related beyond the selected chance criterion.
It does not show whether:
stress affects sleep;
sleep affects stress;
another variable affects both.
The limits of correlational conclusions are covered in scattergrams and correlation coefficients.
Statistical significance and study quality
Inferential testing cannot correct weaknesses in the original research.
A statistically significant result may still come from a study affected by:
an unrepresentative sample;
poorly operationalised variables;
demand characteristics;
investigator effects;
low reliability;
low validity;
uncontrolled extraneous variables.
Statistical testing is one part of analysing psychological research. It does not replace the evaluation of the method.
A complete significance decision
A clear significance decision should include four parts:
State the observed value.
State the critical value.
Compare the values using the correct test rule.
State whether the result is significant and what happens to the null hypothesis.
Example structure
The observed value was \(14\) and the critical value was \(11\). As the observed value was greater than the critical value, the result was statistically significant. The null hypothesis was rejected.
For a lower-value test:
The observed value was \(2\) and the critical value was \(3\). As the observed value was equal to or less than the critical value, the result was statistically significant. The null hypothesis was rejected.
Writing a contextual conclusion
After making the statistical decision, refer back to the study.
Weak conclusion:
The result was significant.
Stronger conclusion:
There was a statistically significant difference in the number of words recalled in the silence and background-noise conditions.
For a correlation:
There was a statistically significant correlation between hours of sleep and memory-test score.
For an association:
There was a statistically significant association between treatment condition and whether participants improved.
Use the correct wording for the research purpose:
difference for comparisons;
correlation for two co-variables;
association for categorical variables.
A statistical-testing decision process
Use the following sequence when answering an inferential-testing question.
Step 1: Identify the research purpose
Is the researcher testing:
a difference;
a correlation;
an association?
Step 2: Identify the level of measurement
Are the data:
nominal;
ordinal;
interval?
Step 3: Identify whether scores are related or unrelated
For a test of difference, determine whether:
the same participants or matched participants provided both sets of scores;
different participants provided the scores.
Step 4: Select the statistical test
Use the purpose, design and level of measurement.
Step 5: Calculate or identify the observed value
Use the procedure for the selected test.
Step 6: Find the critical value
Use the correct statistical table, sample information, significance level and hypothesis direction.
Step 7: Compare the values
Apply the correct comparison rule for the test.
Step 8: Draw a conclusion
State:
whether the result is significant;
whether the null hypothesis is rejected or retained;
what the result means in the context of the study.
Key Words 🔑
Key word | Student-friendly definition | How it may be used in an exam |
Descriptive statistics | Numerical techniques used to organise and summarise collected data. | You may distinguish them from inferential statistical tests. |
Inferential statistics | Statistical procedures used to judge whether a sample result is unlikely to have occurred by chance. | You may explain their purpose or use them to determine significance. |
Statistical test | A procedure used to analyse data and produce a test statistic. | You may select a test using the research purpose, design and level of measurement. |
Statistical significance | A judgement that a result meets the selected statistical criterion and is unlikely to be due to chance alone. | You may determine significance by comparing observed and critical values. |
Observed value | The test statistic calculated from the study’s results. | You may compare it with a critical value. |
Critical value | The threshold obtained from a statistical table for determining significance. | You may select it using information about the test and study. |
Statistical table | A table containing critical values for a particular inferential test. | You may use an extract to identify the appropriate critical value. |
Null hypothesis | A statement predicting no difference, correlation or association, with any apparent pattern due to chance. | It is rejected when a result is statistically significant. |
Alternative hypothesis | A statement predicting a difference, correlation or association. | It receives support when the null hypothesis is rejected. |
Significance level | The probability threshold used to judge whether a result is statistically significant. | It helps determine which critical value is selected. |
Directional hypothesis | A hypothesis predicting the direction of a difference or relationship. | It may affect the column selected from a statistical table. |
Non-directional hypothesis | A hypothesis predicting a difference or relationship without stating its direction. | It may require a different critical value from a directional hypothesis. |
Chance variation | Random differences that may occur when a sample is selected or measured. | Inferential testing assesses whether a result is likely to exceed chance variation. |
Retain the null hypothesis | The decision made when the result is not statistically significant. | Use this conclusion when the observed value does not meet the critical-value rule. |
Reject the null hypothesis | The decision made when the result is statistically significant. | Use this conclusion when the observed value meets the critical-value rule. |
Hints from the Examiner Reports 💡
No lesson-specific examiner guidance was identified in the provided reports.
Common Mistakes ⚠️
Mistake: Assuming that a difference between two means is automatically significant.
Why this is incorrect:
A numerical difference describes the sample but does not show whether the result is unlikely to have occurred through chance.
How to improve:
Use an appropriate inferential statistical test before making a significance claim.
Mistake: Confusing the observed value with the critical value.
Why this is incorrect:
The observed value is calculated from the study’s results. The critical value is obtained from a statistical table.
How to improve:
Remember:
Observed comes from the observations. Critical comes from the critical-values table.
Mistake: Using the same comparison rule for every statistical test.
Why this is incorrect:
Some tests require the observed value to be equal to or greater than the critical value. Other tests require it to be equal to or smaller than the critical value.
How to improve:
Learn the rule for the specific test and check the instructions provided with the statistical table.
Mistake: Treating equal observed and critical values as non-significant.
Why this is incorrect:
The observed value has reached the threshold when it is exactly equal to the critical value.
How to improve:
Include the equality symbol in the decision rule:
$$\geq$$
or:
$$\leq$$
as appropriate.
Mistake: Saying that a significant result proves the alternative hypothesis.
Why this is incorrect:
Statistical testing involves probability and does not provide absolute proof.
How to improve:
Write that the null hypothesis is rejected and the alternative hypothesis is supported.
Mistake: Saying that a non-significant result proves there is no effect.
Why this is incorrect:
A non-significant result means that there is insufficient evidence to reject the null hypothesis.
How to improve:
Write that the null hypothesis is retained rather than proved.
Mistake: Selecting a critical value before selecting the statistical test.
Why this is incorrect:
Each inferential test has its own statistical table and decision procedure.
How to improve:
First identify the research purpose, design and level of measurement.
Mistake: Ignoring whether the hypothesis is directional or non-directional.
Why this is incorrect:
The direction of the hypothesis may affect which critical value is selected from the table.
How to improve:
Read the hypothesis carefully before using the statistical table.
Mistake: Writing only “the result is significant”.
Why this is incorrect:
The answer does not show how the decision was reached or what the result means.
How to improve:
Include:
the observed value;
the critical value;
the comparison;
the decision about the null hypothesis;
a contextual conclusion.
Mistake: Describing a statistically significant correlation as causal.
Why this is incorrect:
Statistical significance does not change a correlational design into an experiment.
How to improve:
Describe the variables as associated or correlated, not as causing one another.
Exam-Style Questions ✍️
Question 1
Define inferential statistical testing.[2 marks]
Question 2
Explain one difference between descriptive statistics and inferential statistics.[2 marks]
Question 3
Define each of the following:
a) Observed value.[1 mark]
b) Critical value.[1 mark]
Question 4
A researcher compares memory performance in two conditions. The mean scores are:
Condition | Mean memory score |
Condition A | \(16.8\) |
Condition B | \(13.4\) |
Explain why the difference between these means is not enough to conclude that the result is statistically significant.[3 marks]
Question 5
For a particular statistical test, the observed value must be equal to or greater than the critical value for significance.
The researcher obtains:
$$\text{Observed value}=19$$
The statistical table gives:
$$\text{Critical value}=16$$
a) Determine whether the result is statistically significant.[1 mark]
b) Explain your answer.[2 marks]
c) State what should happen to the null hypothesis.[1 mark]
Question 6
For another statistical test, the observed value must be equal to or smaller than the critical value for significance.
The researcher obtains:
$$\text{Observed value}=6$$
The critical value is:
$$\text{Critical value}=4$$
Determine whether the result is statistically significant. Explain your answer.[3 marks]
Question 7
A sign test produces the following values:
$$S_{\text{observed}}=3$$
$$S_{\text{critical}}=3$$
For the sign test, the observed value must be equal to or smaller than the critical value.
a) State whether the result is significant.[1 mark]
b) Explain the decision.[2 marks]
c) State whether the null hypothesis should be rejected or retained.[1 mark]
Question 8
A student writes:
“The observed value is the number found in the statistical table, while the critical value is calculated from the participants’ scores.”
Explain why the student is incorrect.[3 marks]
Question 9
A psychologist investigates whether there is a relationship between hours of sleep and memory-test score.
The inferential test shows a statistically significant correlation.
Explain two conclusions the psychologist can and cannot make from this result.[4 marks]
Question 10
A researcher compares anxiety ratings before and after a relaxation activity.
The observed value is:
$$2$$
The critical value is:
$$4$$
For the selected test, a result is significant when:
$$\text{Observed value}\leq\text{Critical value}$$
Write a complete conclusion for the study. Your answer should:
compare the observed and critical values;
state whether the result is statistically significant;
state whether the null hypothesis should be rejected or retained;
refer to the difference in anxiety ratings.
[4 marks]



Comments