Analysing and reporting an investigation | AQA A-Level Psychology Revision
- Revision Notes
- Aug 5
- 28 min read
Updated: 5 days ago
For 7182 specification, first teach in September 2025
AQA A-Level Psychology | Free Revision Notes
Estimated study time: 70 minutes
This Analysing and reporting an investigation A-Level Psychology revision page explains what happens after psychological data have been collected. You will learn how to select suitable descriptive statistics, choose an appropriate inferential test, present results clearly and draw justified conclusions. You will also examine the structure of a scientific report, including the abstract, introduction, method, results, discussion and references. Accurate analysis and transparent reporting allow other psychologists to evaluate, replicate and build on an investigation.
Learning Objectives 🎯
By the end of this revision page, you should be able to:
Select suitable descriptive statistics for psychological data.
Select an appropriate inferential statistical test.
Present quantitative data using suitable tables and graphs.
Organise qualitative data appropriately.
Interpret central tendency, dispersion, correlations and statistical significance.
Write a contextual statistical conclusion.
Distinguish the results section from the discussion section.
Describe the purpose of each section of a scientific report.
Report an investigation clearly enough to support replication.
Use references to acknowledge and identify source material.
Revision Notes 📚
Analysing and reporting psychological research
Once an investigation has been conducted, the researcher must make sense of the collected evidence.
The main stages are:
Check and organise the raw data.
Identify the type and level of data.
Select suitable descriptive statistics.
Present the findings in tables or graphs.
Select an appropriate inferential test.
Determine whether the result is statistically significant.
Interpret the result in the context of the investigation.
Evaluate the procedure and findings.
Write a structured scientific report.
Reference any source material.
The analysis must match the original aim, hypothesis, method and data.
These features should have been recorded carefully during conducting an investigation.
Preparing the data for analysis
Preserving raw data
Raw data are the original results collected from participants or observations.
Examples include:
each participant’s memory score;
each response time;
every questionnaire rating;
the frequency of each observed behaviour;
participants’ original answers to open questions;
paired scores from a correlational study.
Researchers should preserve the raw data before calculating summary values.
For example:
Participant | Silence score | Speech score |
P001 | \(16\) | \(12\) |
P002 | \(14\) | \(11\) |
P003 | \(18\) | \(13\) |
P004 | \(15\) | \(10\) |
If only the condition means were saved, individual differences and possible data-entry errors could no longer be examined.
Checking the data
Before beginning analysis, the researcher should check for:
incomplete records;
missing responses;
duplicate cases;
incorrect participant codes;
values entered in the wrong condition;
impossible values;
inconsistent units;
transcription errors;
lost pairings in related or correlational data.
For example, a questionnaire using a scale from \(1\) to \(5\) should not contain:
$$12$$
unless the score has a different meaning.
Missing data
Missing data should not automatically be replaced with zero.
A zero may be a genuine score.
For example:
a participant may recall zero words;
an observed behaviour may occur zero times.
A missing value means that no valid result was obtained.
Possible reasons include:
the participant skipped a question;
equipment failed;
the participant withdrew;
the recording was incomplete.
The reason should be documented where possible.
Preserving related scores
In repeated measures, matched pairs and correlational studies, the link between scores must be preserved.
For example:
Participant | Hours slept | Memory score |
P001 | \(6\) | \(12\) |
P002 | \(8\) | \(17\) |
P003 | \(5\) | \(9\) |
If the memory scores are rearranged independently of the sleep scores, the correlation will be invalid because the original pairs have been lost.
Quantitative and qualitative analysis
Quantitative data
Quantitative data are numerical.
Examples include:
frequencies;
ratings;
response times;
memory scores;
percentages;
ranks;
correlation coefficients.
Quantitative data may be analysed using:
descriptive statistics;
graphs and tables;
inferential statistical tests.
Qualitative data
Qualitative data are expressed in words rather than numerical measurements.
Examples include:
interview responses;
answers to open questions;
written accounts;
descriptions of personal experience;
material used in a content analysis.
The distinction between the data types is covered in quantitative and qualitative data.
Organising qualitative data
Qualitative material may be organised using coding categories.
For example, answers to the question:
What makes revision difficult?
might be coded into categories such as:
lack of time;
distractions;
lack of motivation;
uncertainty about what to revise.
The categories should be:
clearly defined;
relevant to the research question;
sufficiently distinct;
applied consistently.
The researcher may then count how often each category occurs.
This converts aspects of the qualitative material into quantitative frequency data while the original responses remain available.
Avoiding selective interpretation
A researcher should not select only statements that support the hypothesis.
Analysis should represent:
common responses;
contrasting responses;
unexpected findings;
relevant patterns across the material.
The categories should not be altered simply to make the evidence appear clearer or more supportive.
Descriptive statistics
What are descriptive statistics?
Descriptive statistics organise and summarise the data collected in an investigation.
They describe the sample but do not determine whether a result is statistically significant.
Descriptive statistics include:
measures of central tendency;
measures of dispersion;
percentages;
frequencies;
correlation coefficients.
Selecting descriptive statistics
The researcher should consider:
the level of measurement;
the shape of the distribution;
whether extreme scores are present;
what information needs to be communicated.
A useful starting point is:
Level of measurement | Suitable measure of central tendency | Suitable descriptive presentation |
Nominal | Mode | Frequencies or percentages |
Ordinal | Median or mode | Ordered frequencies, median and range |
Interval | Mean, median or mode as appropriate | Mean and standard deviation or range |
The final choice should fit the actual data rather than being selected mechanically.
Measures of central tendency
Mean
The mean is calculated by adding all the scores and dividing by the number of scores.
$$\bar{x}=\frac{\sum x}{N}$$
The mean uses every score and can provide a sensitive summary of interval data.
However, it may be affected by extreme values.
Worked mean example
Scores in a memory condition are:
$$10,\ 12,\ 13,\ 15,\ 15$$
Add the scores:
$$10+12+13+15+15=65$$
Divide by the number of scores:
$$\bar{x}=\frac{65}{5}$$
$$\boxed{\bar{x}=13}$$
The mean memory score is \(13\).
Median
The median is the middle score after the values have been arranged in order.
For:
$$5,\ 7,\ 9,\ 12,\ 16$$
the median is:
$$\boxed{9}$$
The median is useful for ordinal data and may be less affected by extreme scores than the mean.
Mode
The mode is the most frequently occurring score or category.
For:
$$2,\ 3,\ 3,\ 4,\ 5$$
the mode is:
$$\boxed{3}$$
For nominal data, the mode identifies the most common category.
Comparing measures of central tendency
Measure | Main feature | Possible limitation |
Mean | Uses every numerical score | Affected by extreme scores |
Median | Uses the middle position | Does not use the precise value of every score |
Mode | Identifies the most frequent value or category | There may be more than one mode or no clear mode |
The calculation and selection of averages are covered in measures of central tendency.
Measures of dispersion
What is dispersion?
Dispersion refers to the spread or variability of scores.
Two conditions may have the same mean but very different levels of consistency.
Consider:
Condition A
$$9,\ 10,\ 10,\ 10,\ 11$$
Condition B
$$2,\ 6,\ 10,\ 14,\ 18$$
Both conditions have a mean of:
$$10$$
However, Condition B is much more widely spread.
Range
The range is calculated using:
$$\text{Range}=\text{Highest score}-\text{Lowest score}$$
For Condition B:
$$18-2=16$$
The range is:
$$\boxed{16}$$
The range is simple to calculate but uses only the highest and lowest scores.
Standard deviation
The standard deviation describes the typical spread of scores around the mean.
A smaller standard deviation indicates that scores are more closely clustered around the mean.
A larger standard deviation indicates greater variability.
For example:
Condition | Mean | Standard deviation |
Silence | \(15.8\) | \(2.1\) |
Speech | \(12.4\) | \(4.9\) |
The silence condition has:
the higher mean;
the smaller standard deviation.
Participants therefore recalled more words on average in silence, and their scores were more consistent.
Measures of spread are covered in measures of dispersion.
Reporting central tendency with dispersion
A measure of central tendency should often be accompanied by a measure of dispersion.
For example:
Participants in the silent condition recalled a mean of \(15.8\) words, with a standard deviation of \(2.1\).
This provides information about both:
average performance;
variation among participants.
Percentages and frequencies
Frequencies
A frequency shows how often a score, category or behaviour occurred.
For example:
Revision method | Frequency |
Flashcards | \(18\) |
Practice questions | \(27\) |
Mind maps | \(15\) |
The modal category is:
Practice questions
because it has the highest frequency.
Percentages
Percentages allow groups of different sizes to be compared.
Use:
$$\text{Percentage}=\frac{\text{Frequency}}{\text{Total}}\times100$$
Suppose \(18\) out of \(60\) students chose flashcards:
$$\frac{18}{60}\times100=30\%$$
The percentage is:
$$\boxed{30\%}$$
Percentage points
Suppose:
\(60\%\) of Group A select a response;
\(45\%\) of Group B select the response.
The difference is:
$$60-45=15$$
Group A is \(15\) percentage points higher.
This is not automatically the same as saying that Group A is \(15\%\) higher.
Correlation coefficients
A correlation coefficient describes the direction and strength of a relationship.
It lies between:
$$-1\leq r\leq+1$$
or:
$$-1\leq r_s\leq+1$$
depending on the statistical test.
Direction
A positive coefficient indicates a positive relationship.
For example:
$$r=+0.78$$
A negative coefficient indicates a negative relationship.
For example:
$$r_s=-0.81$$
Strength
The closer the coefficient is to:
$$+1$$
or:
$$-1$$
the stronger the relationship.
The closer the coefficient is to:
$$0$$
the weaker the relationship.
The sign shows direction, not strength.
For example:
$$r=-0.86$$
is stronger than:
$$r=+0.32$$
Descriptive and inferential interpretations
A correlation coefficient describes the relationship found in the sample.
The researcher must then use the appropriate inferential procedure to determine whether the correlation is statistically significant.
A large-looking coefficient is not automatically significant because significance also depends on features such as the number of paired scores and the selected significance level.
Presenting quantitative findings
Selecting an appropriate display
Results may be presented using:
tables;
bar charts;
histograms;
scattergrams.
The display must fit the data and research question.
Data or purpose | Suitable presentation |
Exact results or summary statistics | Table |
Comparison of separate categories or conditions | Bar chart |
Continuous scores grouped into intervals | Histogram |
Relationship between two co-variables | Scattergram |
Presentation techniques are covered in tables and graphs.
Tables
Features of an effective results table
A results table should include:
an informative title;
clearly labelled rows and columns;
units;
consistent decimal places;
the relevant descriptive statistics;
totals where appropriate.
Example results table
Table 1. Mean number of words recalled and standard deviation in each sound condition
Condition | Mean words recalled | Standard deviation |
Silence | \(15.8\) | \(2.1\) |
Background speech | \(12.4\) | \(4.9\) |
The title explains:
the measured variable;
the conditions;
the statistics displayed.
Avoiding duplication
A report should not include several displays that communicate exactly the same information without a clear purpose.
For example, a table and bar chart containing only the same two means may be unnecessary unless both formats contribute to understanding.
Bar charts
A bar chart is suitable for separate categories or experimental conditions.
The chart should include:
an informative title;
the conditions or categories on the horizontal axis;
the dependent variable on the vertical axis;
units;
an equal-interval scale;
equal-width bars;
gaps between the bars.
The gaps show that the conditions are separate categories.
Describing a bar chart
A good interpretation uses numerical evidence.
Weak interpretation:
Participants did better in silence.
Improved interpretation:
Participants recalled a mean of \(15.8\) words in silence and \(12.4\) words with background speech, a difference of \(3.4\) words.
The graph alone does not show that the difference is statistically significant.
Histograms
A histogram is used for continuous data arranged into class intervals.
Unlike a bar chart:
the bars touch;
the horizontal axis forms a continuous numerical scale;
the intervals must remain in numerical order.
A histogram may help identify:
a normal distribution;
positive skew;
negative skew;
the interval containing the most scores;
the spread of the distribution.
Scattergrams
A scattergram displays paired scores for two co-variables.
Each participant contributes one point.
For example:
Participant | Hours slept | Memory score |
P001 | \(5\) | \(9\) |
P002 | \(6\) | \(11\) |
P003 | \(7\) | \(14\) |
P004 | \(8\) | \(16\) |
The plotted coordinates are:
$$(5,9)$$
$$(6,11)$$
$$(7,14)$$
$$(8,16)$$
The points should not be joined.
The researcher should describe:
direction;
approximate strength;
any unusual values.
A scattergram does not establish causation.
Selecting an inferential technique
What is inferential testing?
An inferential statistical test helps the researcher decide whether an obtained result is statistically significant.
It allows a probability-based decision about the null hypothesis.
The researcher must select the test before interpreting its result.
Factors affecting test choice
The correct test depends on:
Whether the study tests a difference, correlation or association.
Whether data are related or unrelated, where relevant.
The level of measurement.
The full selection process is covered in choosing an inferential test.
Inferential test-selection table
Statistical test | Purpose | Data relationship | Level of measurement |
Sign test | Difference | Related | Nominal signs |
Wilcoxon | Difference | Related | Ordinal |
Mann-Whitney | Difference | Unrelated | Ordinal |
Related t-test | Difference | Related | Interval |
Unrelated t-test | Difference | Unrelated | Interval |
Spearman’s rho | Correlation | Paired co-variable scores | Ordinal |
Pearson’s \(r\) | Correlation | Paired co-variable scores | Interval |
Chi-squared | Association | Categorical frequencies | Nominal |
Selecting a test of difference
Related nominal data
$$\text{Sign test}$$
Related ordinal data
$$\text{Wilcoxon}$$
Unrelated ordinal data
$$\text{Mann-Whitney}$$
Related interval data
$$\text{Related }t\text{-test}$$
Unrelated interval data
$$\text{Unrelated }t\text{-test}$$
Selecting a correlation test
Ordinal co-variable data
$$\text{Spearman's rho}$$
Interval co-variable data
$$\text{Pearson's }r$$
Selecting a test of association
For an association between two nominal categorical variables using frequencies:
$$\text{Chi-squared}$$
Writing a test justification
A complete justification should identify every relevant feature.
Difference example
A related t-test is appropriate because the researcher is testing for a difference, the same participants complete both conditions so the data are related, and the number of words recalled is measured at the interval level.
Correlation example
Spearman’s rho is appropriate because the researcher is testing for a correlation between two co-variables and both variables are measured using ordinal ratings.
Association example
Chi-squared is appropriate because the researcher is testing for an association between two categorical variables and the results are nominal frequencies.
Statistical significance
Significance level
A significance level is the probability criterion used to determine whether the result is statistically significant.
Common criteria include:
$$p\leq0.05$$
and:
$$p\leq0.01$$
The \(0.01\) criterion is more stringent because stronger evidence is required before the null hypothesis is rejected.
Using a statistical table
The researcher may need to identify:
the relevant sample information;
degrees of freedom;
the significance level;
whether the hypothesis is directional or non-directional;
the appropriate row and column in the table.
The observed test statistic is then compared with the critical value.
The use of tables is covered in probability and significance.
Comparing observed and critical values
The significance rule depends on the statistical test.
Observed value must be equal to or smaller than the critical value
This rule applies to:
sign test;
Wilcoxon;
Mann-Whitney.
$$\text{Observed value}\leq\text{Critical value}$$
Observed value must be equal to or greater than the critical value
This rule applies to:
Spearman’s rho;
Pearson’s \(r\);
related t-test;
unrelated t-test;
Chi-squared.
$$\text{Observed value}\geq\text{Critical value}$$
Where a correlation coefficient is negative, its magnitude is compared with the critical value.
For example:
$$r_s=-0.72$$
has the magnitude:
$$|-0.72|=0.72$$
Equality reaches the critical value
If the observed value equals the critical value, the result has reached the significance threshold.
For example:
$$U_{\text{observed}}=9$$
and:
$$U_{\text{critical}}=9$$
Because:
$$9\leq9$$
the result is statistically significant.
Rejecting or retaining the null hypothesis
Significant result
When the result meets the critical criterion:
the result is statistically significant;
the null hypothesis is rejected;
the alternative hypothesis receives support.
The conclusion may be expressed as:
$$p\leq0.05$$
or another relevant significance level.
Non-significant result
When the result does not meet the critical criterion:
the result is not statistically significant;
the null hypothesis is retained;
there is insufficient evidence to support the alternative hypothesis.
At the \(0.05\) level, this may be written as:
$$p>0.05$$
Retain rather than prove
A non-significant result does not prove that the null hypothesis is true.
It means that the investigation did not provide sufficient evidence to reject it.
A suitable statement is:
The null hypothesis was retained.
Avoid:
The null hypothesis was proved.
Writing a statistical conclusion
A complete statistical conclusion should include:
The name of the statistical test where appropriate.
The observed value.
The critical value.
The comparison rule.
Whether the result is significant.
The probability level.
Whether the null hypothesis is rejected or retained.
A conclusion written in the context of the study.
Worked significant result
A related t-test produces:
$$t_{\text{observed}}=2.84$$
The critical value is:
$$t_{\text{critical}}=2.26$$
Compare the values:
$$2.84\geq2.26$$
A complete conclusion is:
The observed t-value of \(2.84\) was greater than the critical value of \(2.26\). The result was statistically significant at \(p\leq0.05\), so the null hypothesis was rejected. There was a statistically significant difference in the number of words recalled in the silent and background-speech conditions.
Worked non-significant result
A Mann-Whitney test produces:
$$U_{\text{observed}}=18$$
The critical value is:
$$U_{\text{critical}}=13$$
Compare:
$$18>13$$
The result does not meet the required rule:
$$U_{\text{observed}}\leq U_{\text{critical}}$$
A complete conclusion is:
The observed Mann-Whitney value of \(18\) was greater than the critical value of \(13\). The result was not statistically significant at the \(0.05\) level, so \(p>0.05\) and the null hypothesis was retained. There was insufficient evidence of a difference in the ordinal anxiety ratings of the two groups.
Worked correlational conclusion
A Spearman’s rho test produces:
$$r_{s,\text{observed}}=-0.74$$
The critical value is:
$$r_{s,\text{critical}}=0.62$$
Use the magnitude:
$$|-0.74|=0.74$$
Compare:
$$0.74\geq0.62$$
A suitable conclusion is:
The magnitude of the observed Spearman’s rho value was greater than the critical value. The result was statistically significant at \(p\leq0.05\), so the null hypothesis was rejected. There was a statistically significant negative correlation between examination-stress ratings and sleep-quality ratings.
This does not show that examination stress caused poor sleep quality.
Interpreting findings
Interpretation requires more than repeating values
A researcher should explain what the findings mean in relation to:
the aim;
the hypothesis;
the psychological variables;
the method;
possible alternative explanations.
For example:
Participants recalled more words in silence than with background speech.
This is a descriptive finding.
If the inferential result is significant, the researcher may add:
The statistical analysis suggests that the difference was unlikely to have occurred through chance alone at the selected significance level.
Using numerical evidence
Interpretations should be supported with values.
Weak statement:
Silence was much better.
Improved statement:
Participants recalled a mean of \(15.8\) words in silence and \(12.4\) with background speech, a difference of \(3.4\) words.
Interpreting dispersion
The researcher should consider whether performance was consistent.
For example:
The standard deviation was lower in silence, suggesting that participants’ recall scores were more closely clustered around the mean in this condition.
A high mean accompanied by a high standard deviation may indicate that participants did not all respond similarly.
Interpreting a non-significant result
A non-significant result should not be described as showing that absolutely no effect or relationship exists.
It means that sufficient statistical evidence was not obtained.
Possible explanations include:
the null hypothesis is true;
the study failed to detect a genuine result;
the sample was unsuitable;
measurement lacked sensitivity;
uncontrolled variation obscured the effect.
These are possible interpretations, not automatic conclusions.
Type I and Type II errors
A significant result may involve a Type I error if the null hypothesis is actually true.
A non-significant result may involve a Type II error if the null hypothesis is actually false.
The researcher normally cannot know with certainty whether either error has occurred.
Avoiding causal conclusions
A significant correlation does not establish causation.
The relationship may involve:
the first co-variable influencing the second;
the second influencing the first;
a third variable influencing both.
The conclusion should use language such as:
correlated with;
related to;
associated with.
Descriptive and inferential findings compared
Descriptive finding | Inferential finding |
Summarises the sample | Supports a probability-based decision |
Includes means, medians and percentages | Uses a statistical test |
Shows the size or pattern of an obtained result | Determines whether the result reaches a significance criterion |
Does not by itself establish significance | Leads to rejection or retention of the null hypothesis |
Example: Condition A had a higher mean | Example: The difference was statistically significant |
Both forms should be used appropriately.
A report that includes only inferential statistics may fail to show the actual size and direction of the result.
A report that includes only descriptive statistics cannot determine statistical significance.
Scientific reporting
Why psychologists write scientific reports
A scientific report communicates:
why the investigation was conducted;
how it was carried out;
what was found;
what the findings mean;
how the work relates to existing psychological knowledge.
Clear reporting allows other researchers to:
evaluate the study;
inspect the evidence;
identify strengths and limitations;
replicate the procedure;
compare findings;
build on the research.
Sections of a scientific report
The required sections are:
Abstract.
Introduction.
Method.
Results.
Discussion.
References.
These sections are introduced in reporting psychological investigations.
The abstract
Purpose of the abstract
The abstract is a concise summary of the complete investigation.
It allows a reader to understand the main features without reading the entire report.
It is placed near the beginning of the report but is usually easiest to write after the remaining sections have been completed.
What should an abstract contain?
An abstract should briefly identify:
the aim;
the method;
the sample or important design features;
the main findings;
the main conclusion.
Example abstract
This study investigated whether background speech affected immediate word recall in sixth-form students. Thirty participants completed a word-recall task in silent and background-speech conditions using a repeated measures design. Participants recalled more words in silence, and a related t-test showed that the difference was statistically significant. The findings suggest that background speech was associated with reduced immediate recall in this task.
The abstract does not include every procedural detail or a lengthy evaluation.
Common abstract problem
A weak abstract may describe only the topic:
This report is about memory and background sound.
This does not summarise:
what was done;
what was found;
what conclusion was reached.
The introduction
Purpose of the introduction
The introduction explains the background and rationale for the investigation.
It moves from the wider psychological topic towards the specific study.
What should an introduction contain?
An introduction may include:
relevant psychological concepts;
relevant theories or previous research;
the reason for conducting the investigation;
the research aim;
the hypothesis or hypotheses.
The material should be directly relevant to the investigation.
Logical structure
A useful structure is:
Introduce the psychological topic.
Explain relevant theory or previous findings.
Identify the question that the present investigation addresses.
State the aim.
State the hypothesis.
Writing the rationale
The rationale explains why the study is worthwhile.
For example:
Previous psychological research suggests that competing verbal information may interfere with immediate recall. The present investigation examines whether background speech reduces the number of words recalled in a controlled classroom task.
The rationale should connect existing knowledge with the new investigation.
Hypotheses in the introduction
The report should state the operationalised research hypothesis.
For example:
Sixth-form students will recall significantly fewer words from a list of \(20\) words when a recorded conversation is played than when the task is completed in silence.
A null hypothesis may also be reported where appropriate.
The method
Purpose of the method
The method explains exactly how the investigation was carried out.
It should contain enough information for another researcher to replicate the study.
The method is written as a report of what was done rather than a future plan.
Method content
The method should explain:
the research method;
the experimental design where relevant;
the participants and sampling method;
the materials or apparatus;
the variables and their operationalisation;
the procedure;
control techniques;
ethical safeguards.
Design
The design information may identify:
laboratory, field, natural or quasi-experiment;
repeated measures, independent groups or matched pairs;
observation or self-report format;
independent and dependent variables;
relevant controls.
Example:
A laboratory experiment using a repeated measures design was conducted. The independent variable was whether the recall task was completed in silence or while a recorded conversation was played. The dependent variable was the number of words correctly recalled from a list of \(20\) words.
Participants
The participant description should include relevant details such as:
sample size;
target population;
age range where relevant;
sampling method;
allocation to conditions;
inclusion criteria where necessary.
Example:
Thirty sixth-form students aged 16 to 18 were recruited using opportunity sampling from one college. All participants completed both experimental conditions.
The report should not include participant names.
Materials or apparatus
Materials should be described precisely enough to support replication.
Examples include:
word lists;
questionnaires;
recordings;
timers;
computers;
observation schedules;
standardised instructions;
consent and debrief documents.
Example:
Two matched lists of \(20\) common words were used. A recording of two adults speaking was played at the same volume through headphones in the background-speech condition.
Procedure
The procedure should describe the stages in the correct order.
It may include:
recruitment;
consent;
participant instructions;
allocation or counterbalancing;
task sequence;
timing;
recording and scoring;
withdrawal procedures;
debriefing.
Example:
Participants studied each word list for two minutes. After a one-minute distractor task, they had two minutes to write as many words as they could remember. The order of the silent and speech conditions was counterbalanced.
Controls
The method should report important controls.
For example:
standardised instructions;
equal task duration;
matched materials;
consistent equipment;
random allocation;
counterbalancing;
objective scoring.
Controls should be described specifically rather than saying only:
Variables were controlled.
Ethics
The method should explain how relevant ethical issues were addressed.
For example:
informed consent;
right to withdraw;
confidentiality;
protection from harm;
debriefing.
The report should describe what was actually done.
The results section
Purpose of the results section
The results section presents the findings of the investigation.
It should report the evidence clearly without giving a lengthy explanation of why the result occurred.
What should the results section contain?
The results section may include:
descriptive statistics;
tables;
graphs;
inferential test results;
observed and critical values;
the probability level;
the decision about the null hypothesis;
a concise statement of the result.
Descriptive results
For example:
Participants recalled a mean of \(15.8\) words in silence and \(12.4\) words in the background-speech condition. The standard deviations were \(2.1\) and \(4.9\) respectively.
Inferential results
For example:
A related t-test produced an observed value of \(2.84\), which exceeded the critical value of \(2.26\) at \(p\leq0.05\). The result was statistically significant and the null hypothesis was rejected.
Tables and figures
Each table or graph should:
have a number where appropriate;
have an informative title;
be referred to in the text;
include clear labels and units;
avoid unnecessary duplication.
For example:
As shown in Table 1, the silent condition produced a higher mean recall score and a smaller standard deviation.
Results without discussion
The results section should not contain a detailed explanation such as:
Participants performed worse because the speech overloaded their working memory.
That interpretation belongs in the discussion.
The results section should focus on what was found.
The discussion
Purpose of the discussion
The discussion explains and evaluates the findings.
It moves beyond the numerical result to consider what the evidence means.
What should the discussion contain?
A discussion may include:
a summary of the main finding;
whether the hypothesis was supported;
interpretation of the result;
links to relevant psychological explanations or previous research;
consideration of reliability and validity;
methodological strengths and limitations;
ethical or practical issues;
alternative explanations;
suggested improvements;
implications;
suggestions for further research;
an overall conclusion.
Linking to the hypothesis
The discussion should state whether the findings support the research hypothesis.
For example:
The findings supported the directional hypothesis because participants recalled fewer words in the background-speech condition, and the difference was statistically significant.
Avoid saying that the hypothesis was proved.
Explaining the findings
The researcher may consider why the result occurred.
The explanation should be:
psychologically relevant;
linked to the evidence;
appropriately cautious;
consistent with the research method.
Evaluating reliability
The discussion may consider:
whether the procedure was standardised;
whether measurement was consistent;
whether observers agreed;
whether the result might be replicated;
whether equipment performed reliably.
For example:
Reliability was supported by the use of identical written instructions and automated timing in both conditions.
Evaluating validity
The researcher may consider:
whether the operational measure captured the intended construct;
whether extraneous variables were controlled;
whether participants behaved naturally;
whether the findings apply to everyday settings;
whether demand characteristics affected behaviour.
For example:
The use of isolated word lists may have reduced ecological validity because everyday memory often involves meaningful material rather than unrelated words.
Evaluating the sample
The discussion should consider:
sample size;
sampling method;
representativeness;
generalisation;
possible sampling bias.
For example:
The opportunity sample came from one sixth-form college, so the findings may not generalise to students in other settings or to other age groups.
Alternative explanations
A researcher should consider whether factors other than the intended variable could explain the result.
Examples include:
unequal materials;
order effects;
participant variables;
investigator effects;
situational variables;
response bias;
demand characteristics.
Improvements
An improvement should address a specific limitation.
Weak improvement:
Use a better sample.
Improved version:
Recruit a stratified sample from several colleges so the proportions of relevant student groups are represented and the findings are less dependent on one institution.
Further research
Further research may:
replicate the study with a different population;
use a more realistic task;
investigate another condition;
improve the measure;
examine whether the finding occurs in a different setting.
Suggestions should follow logically from the findings or limitations.
References
Purpose of referencing
Referencing identifies the sources used in the report.
Sources may include:
journal articles;
books;
reports;
psychological measures;
other published material.
Referencing allows readers to:
identify the source of an idea or finding;
locate the original work;
check how evidence has been represented;
distinguish the researcher’s contribution from previous work;
avoid presenting another person’s work as their own.
Citations and reference list
A citation identifies a source within the report.
A reference list provides the full details needed to locate each cited source.
Every cited source should appear in the reference list.
Sources not used in the report should not normally be included simply to make the list appear longer.
Consistency
A report should use a consistent referencing format.
The exact layout may depend on the required referencing system, but information commonly includes:
author;
year;
title;
publication details.
Referencing research materials
Researchers should also acknowledge established measures, scales or published materials where these have been used.
This helps readers understand:
where the material originated;
whether the measure has an established history;
how to locate further information.
Results and discussion compared
Results section | Discussion section |
Presents what was found | Explains what the findings mean |
Includes statistics, tables and graphs | Links findings to theory and previous research |
Reports significance | Evaluates reliability and validity |
States whether the null is rejected or retained | Considers limitations and alternative explanations |
Uses concise factual language | Develops interpretation and evaluation |
Does not give lengthy explanations | Suggests improvements and further research |
Example results statement
Participants recalled a mean of \(15.8\) words in silence and \(12.4\) words with background speech. The difference was statistically significant at \(p\leq0.05\).
Example discussion statement
The findings suggest that background speech reduced performance on this immediate recall task. However, the artificial word-list procedure may limit ecological validity.
Writing objectively
Objective reporting
A scientific report should use clear, evidence-based language.
Weak statement:
The speech condition was obviously terrible and ruined everyone’s memory.
Improved statement:
Participants recalled fewer words on average in the background-speech condition.
The improved version:
uses measured evidence;
avoids emotional language;
does not exaggerate the conclusion.
Separating evidence from interpretation
The report should distinguish:
what the data show;
how the researcher interprets the data.
For example:
The mean was lower in the speech condition.
is a direct description.
Competing verbal information may have interfered with recall.
is an interpretation.
Both may be useful, but they belong in different parts of the report.
Avoiding unsupported certainty
Scientific conclusions should reflect the limits of the evidence.
Appropriate wording includes:
suggests;
supports;
was associated with;
may indicate;
is consistent with;
provides evidence for.
Avoid:
proves;
definitely causes;
applies to everyone;
can never be wrong.
Complete worked analysis
Research aim
To investigate whether background speech affects the number of words recalled by sixth-form students.
Design and data
The study uses:
a repeated measures design;
related interval scores;
number of words recalled as the dependent variable.
Descriptive results
Condition | Mean words recalled | Standard deviation |
Silence | \(15.8\) | \(2.1\) |
Background speech | \(12.4\) | \(4.9\) |
Descriptive interpretation
Participants recalled:
$$15.8-12.4=3.4$$
more words on average in silence.
The smaller standard deviation in silence suggests that scores were more consistent in that condition.
Selecting an inferential test
The study investigates:
$$\text{A difference}$$
The data are:
$$\text{Related}$$
The level of measurement is:
$$\text{Interval}$$
The appropriate test is:
$$\boxed{\text{Related }t\text{-test}}$$
Inferential result
Suppose:
$$t_{\text{observed}}=2.84$$
and:
$$t_{\text{critical}}=2.26$$
Because:
$$2.84\geq2.26$$
the result is statistically significant.
Statistical conclusion
The observed t-value of \(2.84\) was greater than the critical value of \(2.26\). The result was statistically significant at \(p\leq0.05\), so the null hypothesis was rejected. There was a statistically significant difference in the number of words recalled in the silent and background-speech conditions.
Discussion interpretation
The findings support the directional hypothesis because participants recalled fewer words with background speech. The result suggests that speech may interfere with performance on this immediate recall task. However, the use of an opportunity sample from one college limits population validity, and recalling isolated words may not represent everyday memory.
Complete worked correlational analysis
Aim
To investigate whether examination-stress ratings are related to sleep-quality ratings.
Data
Both variables are measured using ordered rating scales.
The study therefore uses:
two co-variables;
paired scores;
ordinal data.
Suitable descriptive presentation
A scattergram should be used to show:
examination-stress rating on one axis;
sleep-quality rating on the other;
one point for each participant.
Selecting an inferential test
The purpose is:
$$\text{Correlation}$$
The level of measurement is:
$$\text{Ordinal}$$
The appropriate test is:
$$\boxed{\text{Spearman's rho}}$$
Inferential result
Suppose:
$$r_{s,\text{observed}}=-0.68$$
and:
$$r_{s,\text{critical}}=0.57$$
Use the magnitude:
$$|-0.68|=0.68$$
Because:
$$0.68\geq0.57$$
the result is statistically significant.
Conclusion
There was a statistically significant negative correlation between examination-stress ratings and sleep-quality ratings at \(p\leq0.05\). Higher stress ratings tended to be associated with lower sleep-quality ratings. The null hypothesis was rejected.
The researcher cannot conclude that stress caused poor sleep quality.
Complete worked Chi-squared analysis
Aim
To investigate whether revision method is associated with whether students pass or fail a test.
Data
Both variables are categorical.
The results consist of frequencies.
Suitable presentation
A contingency table should show the frequencies for each combination of:
revision method;
pass or fail outcome.
Selecting the test
The study investigates:
$$\text{An association}$$
The data are:
$$\text{Nominal frequencies}$$
The appropriate test is:
$$\boxed{\text{Chi-squared}}$$
Degrees of freedom
Suppose the table has:
three revision-method rows;
two outcome columns.
Then:
$$df=(r-1)(c-1)$$
$$df=(3-1)(2-1)$$
$$df=2\times1$$
$$\boxed{df=2}$$
Interpretation
The researcher uses:
\(df=2\);
the selected significance level;
the Chi-squared table.
If:
$$\chi^2_{\text{observed}}\geq\chi^2_{\text{critical}}$$
the association is statistically significant.
The conclusion should refer to an association, not a correlation or causal effect.
Scientific report checklist
Abstract
Is the aim stated?
Is the method summarised?
Are the main findings included?
Is the overall conclusion clear?
Is the section concise?
Introduction
Is the psychological background relevant?
Is previous theory or research explained accurately?
Is the rationale clear?
Is the aim stated?
Is the hypothesis operationalised?
Method
Is the research method identified?
Is the design described?
Is the sample explained?
Are materials described?
Is the procedure replicable?
Are variables operationalised?
Are controls reported?
Are ethical safeguards included?
Results
Are descriptive statistics suitable?
Are tables and graphs clearly labelled?
Is an appropriate inferential test used?
Are observed and critical values reported?
Is the probability level stated?
Is the null hypothesis rejected or retained correctly?
Are interpretations kept brief?
Discussion
Is the main finding explained?
Is the hypothesis addressed?
Are relevant psychological explanations considered?
Are reliability and validity evaluated?
Are sample limitations discussed?
Are alternative explanations considered?
Are improvements specific?
Are future research suggestions relevant?
Is the conclusion appropriately cautious?
References
Is every cited source included?
Can the sources be located?
Is the format consistent?
Have sources been acknowledged clearly?
Key Words 🔑
Key word | Student-friendly definition | How it may be used in an exam |
Raw data | Original participant responses or measurements before summarising. | You may organise raw results into tables or calculate descriptive statistics. |
Descriptive statistic | A value used to organise or summarise collected data. | Examples include the mean, median, range and percentage. |
Inferential test | A statistical procedure used to determine whether a result is statistically significant. | You may need to select and justify an appropriate test. |
Central tendency | A value representing the centre or typical score in a distribution. | The mean, median and mode are measures of central tendency. |
Dispersion | The spread or variability of scores. | The range and standard deviation are measures of dispersion. |
Mean | The total of the scores divided by the number of scores. | It may be reported with the standard deviation for interval data. |
Median | The middle score after values are placed in order. | It is often suitable for ordinal data or skewed scores. |
Mode | The most frequently occurring score or category. | It can be used with nominal data. |
Range | The difference between the highest and lowest scores. | It provides a simple measure of dispersion. |
Standard deviation | A measure showing how scores are spread around the mean. | A smaller value indicates greater consistency. |
Frequency | The number of times a score, response or category occurs. | Frequencies may be converted into percentages or displayed in a table. |
Correlation coefficient | A value showing the direction and strength of a relationship. | It must be interpreted alongside statistical significance. |
Observed value | The test statistic calculated from the research results. | It is compared with the critical value. |
Critical value | The threshold obtained from a statistical table. | It determines whether the observed result reaches significance. |
Statistical significance | A judgement that a result meets the selected probability criterion. | A significant result leads to rejection of the null hypothesis. |
Null hypothesis | A prediction that there is no genuine difference, correlation or association. | It is rejected or retained following inferential testing. |
Abstract | A concise summary of the complete investigation. | It includes the aim, method, main findings and conclusion. |
Introduction | The section explaining the psychological background, rationale, aim and hypothesis. | It establishes why the study was conducted. |
Method | The section describing how the investigation was carried out. | It should provide enough detail for replication. |
Results | The section presenting descriptive and inferential findings. | It contains tables, graphs and statistical conclusions. |
Discussion | The section interpreting and evaluating the findings. | It considers explanations, limitations, improvements and implications. |
Reference | Full publication information identifying a source used in the report. | References allow readers to locate original material. |
Replication | Repeating an investigation using the same procedure. | A detailed method supports replication. |
Objective reporting | Communicating findings using evidence rather than personal opinion. | Scientific reports should avoid exaggerated or emotive claims. |
Hints from the Examiner Reports 💡
No lesson-specific examiner guidance was identified in the provided reports.
Common Mistakes ⚠️
Mistake: Calculating statistics before checking the raw data.
Why this is incorrect:
Incorrect codes, missing values or lost pairings may make the later analysis invalid.
How to improve:
Check completeness, ranges, units and participant pairings first.
Mistake: Replacing missing data with zero.
Why this is incorrect:
Zero may be a genuine result, while missing data mean that no valid result was collected.
How to improve:
Use a clear missing-data code and document the reason.
Mistake: Selecting the mean simply because the data contain numbers.
Why this is incorrect:
Numbers may represent nominal codes or ordinal ratings.
How to improve:
Identify the level of measurement and the shape of the distribution before selecting a measure.
Mistake: Reporting a measure of central tendency without dispersion.
Why this is incorrect:
Two groups may have the same average but very different variability.
How to improve:
Report an appropriate measure of spread alongside the average where relevant.
Mistake: Comparing frequencies from groups with different sample sizes.
Why this is incorrect:
The larger group may produce a greater frequency even when a smaller proportion shows the behaviour.
How to improve:
Convert frequencies into percentages before comparing unequal groups.
Mistake: Describing a visible graph difference as statistically significant.
Why this is incorrect:
A graph provides descriptive evidence only.
How to improve:
Use the term statistically significant only when supported by an inferential test.
Mistake: Selecting a statistical test from the level of measurement alone.
Why this is incorrect:
Ordinal data could require Spearman’s rho, Wilcoxon or Mann-Whitney.
How to improve:
Identify the research purpose and design before considering the level of measurement.
Mistake: Justifying an inferential test only by naming the experimental design.
Why this is incorrect:
The purpose and level of measurement are also required.
How to improve:
State whether the study tests a difference, correlation or association, whether scores are related, and the level of data.
Mistake: Using the wrong observed-critical comparison rule.
Why this is incorrect:
Some tests require the observed value to be smaller, while others require it to be larger.
How to improve:
Learn the decision rule for the selected test.
Mistake: Treating equality as non-significant.
Why this is incorrect:
An observed value equal to the critical value has reached the significance threshold.
How to improve:
Use the correct \(\leq\) or \(\geq\) symbol.
Mistake: Stating that a non-significant result proves the null hypothesis.
Why this is incorrect:
The study may simply have failed to detect a genuine result.
How to improve:
State that the null hypothesis was retained because there was insufficient evidence to reject it.
Mistake: Ignoring the direction of a correlation.
Why this is incorrect:
A negative coefficient and a positive coefficient describe different patterns.
How to improve:
Report both the strength and direction of the relationship.
Mistake: Claiming causation from a significant correlation.
Why this is incorrect:
A correlation cannot establish the direction of cause and effect or rule out third variables.
How to improve:
Use the terms correlation, relationship or association.
Mistake: Putting explanations in the results section.
Why this is incorrect:
The results section reports what was found. Detailed explanation belongs in the discussion.
How to improve:
Keep the results factual and move interpretation to the discussion.
Mistake: Repeating every numerical result in the discussion.
Why this is incorrect:
The discussion should interpret and evaluate rather than duplicate the results section.
How to improve:
Summarise the main result and concentrate on what it means.
Mistake: Writing a method as a list of vague instructions.
Why this is incorrect:
The method should report what was done and contain enough detail for replication.
How to improve:
Include participants, design, materials, operationalised variables, controls and procedure.
Mistake: Including participant names in the report.
Why this is incorrect:
This could breach confidentiality.
How to improve:
Use participant codes and report group-level results.
Mistake: Writing an abstract that contains only background information.
Why this is incorrect:
The abstract should summarise the entire investigation.
How to improve:
Include the aim, method, main finding and conclusion.
Mistake: Describing a limitation without suggesting a relevant improvement.
Why this is incorrect:
A useful evaluation should explain how the weakness could be addressed.
How to improve:
Link each improvement directly to the identified limitation.
Mistake: Using references inconsistently.
Why this is incorrect:
Readers may be unable to identify or locate the source.
How to improve:
Use one consistent referencing format and include every cited source in the reference list.
Exam-Style Questions ✍️
Question 1
Explain one difference between descriptive and inferential statistics.[2 marks]
Question 2
A questionnaire produces nominal categories showing participants’ preferred revision method.
Identify:
a) an appropriate measure of central tendency;[1 mark]
b) an appropriate graphical display.[1 mark]
Question 3
The results of a memory investigation are shown below.
Condition | Mean words recalled | Standard deviation |
Silence | \(17.4\) | \(2.2\) |
Speech | \(13.1\) | \(5.6\) |
Interpret the findings. Refer to both central tendency and dispersion.[4 marks]
Question 4
A researcher compares ordinal anxiety ratings before and after therapy. The same participants provide both ratings.
Identify an appropriate inferential test and justify your answer.[4 marks]
Question 5
Two independent groups complete a memory task. The dependent variable is the number of words correctly recalled.
Identify an appropriate inferential test and explain why it should be used.[4 marks]
Question 6
A related t-test produces:
$$t_{\text{observed}}=2.71$$
The critical value is:
$$t_{\text{critical}}=2.18$$
The researcher used a significance level of:
$$p\leq0.05$$
Write a complete statistical conclusion.[4 marks]
Question 7
A Mann-Whitney test produces:
$$U_{\text{observed}}=17$$
The critical value is:
$$U_{\text{critical}}=12$$
For Mann-Whitney, the observed value must be equal to or smaller than the critical value.
Explain:
a) whether the result is statistically significant;[2 marks]
b) what should happen to the null hypothesis;[1 mark]
c) the appropriate probability statement at the \(0.05\) level.[1 mark]
Question 8
Explain two differences between the results and discussion sections of a scientific report.[4 marks]
Question 9
Describe the purpose of each of the following sections:
a) abstract;[2 marks]
b) introduction;[2 marks]
c) method;[2 marks]
d) results;[2 marks]
e) discussion;[2 marks]
f) references.[2 marks]
Question 10
A psychologist investigates whether background speech affects immediate word recall.
The study uses a repeated measures design with \(24\) sixth-form students. Participants recall a mean of \(16.2\) words in silence and \(12.7\) words with speech. The related t-test result is statistically significant at \(p\leq0.05\).
Write:
a) a suitable results paragraph;[4 marks]
b) a suitable discussion paragraph that interprets and evaluates the findings.[6 marks]



Comments