Reliability | AQA A-Level Psychology Revision
- Revision Notes
- Aug 4
- 17 min read
Updated: 6 days ago
For 7182 specification, first teach in September 2025
AQA A-Level Psychology | Free Revision Notes
Estimated study time: 55 minutes
Reliability concerns the consistency of psychological research. This Reliability A-Level Psychology revision page explains how reliability applies across experiments, observations, questionnaires, interviews and other research methods. You will learn how test-retest and inter-observer reliability are measured, why unreliable procedures weaken confidence in findings and how psychologists can improve consistency. Reliability is a required part of the scientific processes section of AQA A-Level Psychology research methods and is closely connected to standardisation, operationalisation and replication.
Learning Objectives 🎯
By the end of this revision page, you should be able to:
Define reliability in psychological research.
Explain how reliability applies across different research methods.
Explain how test-retest reliability is assessed.
Explain how inter-observer reliability is assessed.
Identify possible causes of unreliable measurements or procedures.
Suggest and justify practical ways of improving reliability.
Revision Notes 📚
What is reliability?
Reliability is the consistency of a research procedure or measurement.
A reliable procedure should produce similar results when it is used consistently under similar circumstances.
Reliability may concern whether:
The same measure produces consistent results on different occasions.
Different observers record the same behaviour consistently.
Instructions and procedures are applied in the same way.
Responses are scored using consistent criteria.
Equipment produces dependable measurements.
Another researcher could follow the procedure in the same way.
Reliability is relevant across all psychological research methods, not only experiments.
Researchers must first define exactly what they intend to measure. This connects reliability with creating precise and measurable variables.
Why does reliability matter?
Psychologists use research findings to draw conclusions about behaviour and mental processes. If the procedure or measurement is inconsistent, differences in the results may have been caused by the way the investigation was conducted rather than by genuine differences between participants or conditions.
For example, imagine that two researchers score the same set of responses differently. The final results would depend partly on which researcher completed the scoring.
This would make it difficult to know whether the recorded differences reflect:
Genuine participant behaviour.
Changes in the variable being studied.
Inconsistent instructions.
Different interpretations by researchers.
Unreliable equipment or scoring.
Reliable procedures increase confidence that results are not simply the product of inconsistent measurement.
Reliability and replicability
Reliability and replicability are connected but are not identical.
Reliability concerns whether a procedure or measure is consistent.
Replicability concerns whether a study is described clearly enough for another researcher to repeat it.
A carefully standardised and clearly reported procedure is easier to replicate. Replication can then help psychologists investigate whether a finding occurs consistently.
Replicability is one of the characteristics considered in the scientific features of psychological research.
Reliability across psychological research methods
Reliability can be considered in:
Experiments.
Observations.
Questionnaires.
Interviews.
Correlations.
Content analyses.
Case studies.
The exact source of unreliability depends on the method being used.
Research method | Possible reliability problem |
Experiment | Participants receive different instructions or testing conditions |
Observation | Observers interpret the same behaviour differently |
Questionnaire | Questions are ambiguous or produce inconsistent responses |
Interview | Interviewers ask questions or use prompts differently |
Correlation | Co-variables are measured inconsistently |
Content analysis | Researchers apply coding categories differently |
Case study | Information is collected or interpreted without a clear procedure |
Reliability in experiments
An experiment may be unreliable if its procedure is not applied consistently.
Possible problems include:
Participants receive different instructions.
Different amounts of time are allowed.
Equipment is used differently.
Some participants receive additional help.
Testing conditions vary unnecessarily.
The dependent variable is scored inconsistently.
The procedure is not described clearly enough to repeat.
For example, a psychologist investigates whether background noise affects memory. Some participants are given two minutes to recall words, while others are given three minutes.
The memory measure is not being applied consistently. Differences in recall scores may be caused by the time available rather than by the independent variable.
Improving reliability in experiments
Researchers may improve reliability by:
Preparing standardised instructions.
Using the same equipment.
Allowing the same amount of time.
Keeping the testing environment consistent.
Following a written procedure.
Operationalising the dependent variable clearly.
Using objective scoring criteria.
Training investigators.
Piloting the procedure.
Recording enough detail for replication.
These improvements use the procedures covered in randomisation and standardisation within research.
Reliability in observations
Reliability is particularly important in observations because researchers must decide how behaviour should be categorised and recorded.
Observers may disagree about:
Whether a behaviour occurred.
Which category should contain it.
When the behaviour began or ended.
Whether an action meets the category definition.
How many times the behaviour occurred.
For example, one observer may record an interaction as aggressive while another records it as playful.
If the category is simply labelled “aggression”, personal interpretation may influence the recording.
A clearer category might be:
Pushes, hits or kicks another person.
This describes observable actions and reduces the amount of judgement required.
Improving reliability in observations
Researchers may:
Construct clear behavioural categories.
Ensure categories do not overlap unnecessarily.
Define what counts as one occurrence.
Use an observation schedule.
Train observers using examples.
Allow observers to practise before the main study.
Use more than one observer.
Compare the observers’ records.
Revise unclear categories after a pilot study.
Record behaviour independently before comparing results.
The construction of clear categories is covered in planning how observable behaviour will be recorded.
Reliability in questionnaires
Questionnaires may produce unreliable data if participants interpret questions differently or give inconsistent responses.
Reliability problems may be caused by:
Ambiguous wording.
Undefined terms such as “often” or “regularly”.
Questions that contain more than one idea.
Unclear response options.
Different instructions.
Changes in how the questionnaire is administered.
Questions that depend heavily on a temporary mood or situation.
Inconsistent scoring.
For example:
Do you regularly feel tired and unable to concentrate?
This question includes two different experiences. A participant may regularly feel tired but have no difficulty concentrating.
It is therefore unclear how the participant should answer.
Improving reliability in questionnaires
Researchers may improve consistency by:
Using clear and unambiguous wording.
Asking one thing in each question.
Defining time periods precisely.
Providing suitable response options.
Giving all participants the same instructions.
Presenting questions in a consistent format.
Using clear scoring rules.
Piloting the questionnaire.
Using test-retest reliability to examine the consistency of responses.
These decisions are relevant when constructing effective self-report questions.
Reliability in interviews
Interviews may be unreliable if interviewers do not conduct them consistently.
Possible sources of unreliability include:
Asking different questions.
Changing the order of questions.
Using different follow-up prompts.
Giving some participants more time.
Explaining questions differently.
Recording answers inconsistently.
Interpreting similar responses differently.
A structured interview is generally easier to standardise because the interviewer uses prepared questions in a predetermined order.
An unstructured interview allows the interviewer to respond flexibly to each participant. This flexibility may produce detailed information, but it can also make the procedure more difficult to repeat consistently.
Improving reliability in interviews
Researchers may:
Prepare a standardised list of questions.
Use the same question order.
Develop agreed prompts.
Train interviewers.
Record interviews consistently.
Prepare clear coding or scoring criteria.
Ensure interviewers do not paraphrase questions differently.
Pilot the interview schedule.
The balance between structure and flexibility is explored in structured and unstructured interviewing.
Reliability in correlations
A correlation investigates the relationship between two co-variables.
The measurements used for both co-variables should be reliable. If one variable is measured inconsistently, the apparent relationship may not represent the genuine pattern between them.
For example, a psychologist investigates the relationship between sleep and concentration.
If sleep is measured using a different time period for different participants, or concentration is tested under inconsistent conditions, the resulting correlation may be misleading.
Researchers should:
Operationalise each co-variable clearly.
Use the same measurement procedure for all participants.
Apply consistent scoring.
Keep relevant testing conditions standardised.
Check whether measures produce dependable results.
Reliability in content analysis
In a content analysis, researchers place material into coding categories.
Reliability may be low if:
Categories are vague.
Categories overlap.
Researchers interpret content differently.
Coding rules are incomplete.
Researchers change how categories are applied.
For example, researchers analysing television programmes might create a category labelled “negative behaviour”. This is subjective and could be interpreted in several ways.
The category should be replaced with specific, observable examples that can be identified consistently.
Researchers can improve reliability by:
Producing clear coding categories.
Defining examples and boundaries.
Training coders.
Having more than one coder analyse the same material independently.
Comparing their records.
Revising categories where agreement is low.
Reliability in case studies
Case studies involve detailed investigations of an individual, group, organisation or event.
Reliability can be difficult to achieve because:
The case may be unique.
Information may come from several sources.
Researchers may interpret detailed material differently.
Flexible procedures may develop as the case progresses.
The exact circumstances may not be repeatable.
Researchers can still improve reliability by:
Describing the procedure clearly.
Using consistent interview or observation schedules where appropriate.
Recording how evidence was collected.
Applying clear coding procedures.
Reporting the method in sufficient detail.
Avoiding unsupported personal interpretation.
However, improving reliability does not make a unique case completely repeatable.
Test-retest reliability
Test-retest reliability measures the consistency of a psychological test or measure over time.
The same measure is given to the same participants on two separate occasions. The two sets of scores are then compared.
A strong positive relationship between the scores suggests that the measure has good test-retest reliability.
The test-retest procedure
A researcher would:
Give a measure to a group of participants.
Record each participant’s score.
Wait for an appropriate period.
Give the same measure to the same participants again.
Record the second set of scores.
Compare the scores from the two occasions.
Judge whether participants obtained reasonably consistent results.
The procedure can be used with measures such as questionnaires and psychological tests.
Example of test-retest reliability
A psychologist develops a questionnaire measuring examination confidence.
The questionnaire is completed by the same 30 students on two occasions, two weeks apart.
If students who receive high scores on the first occasion also tend to receive high scores on the second occasion, the measure appears consistent over time.
If scores change greatly without a clear reason, the questionnaire may have low test-retest reliability.
Interpreting test-retest reliability
The researcher is interested in the consistency of the pattern of scores.
High test-retest reliability suggests that:
The measure is being interpreted consistently.
The scoring system is stable.
The measure is not producing highly variable results without explanation.
Low test-retest reliability may suggest that:
Questions are unclear.
Scoring rules are inconsistent.
Participants interpret items differently on each occasion.
The measure is strongly influenced by temporary circumstances.
The procedure changed between occasions.
Selecting an appropriate time interval
The interval between the two tests must be chosen carefully.
If the interval is very short:
Participants may remember their previous responses.
Familiarity may influence their second performance.
Practice effects may increase scores.
If the interval is very long:
The characteristic being measured may genuinely change.
Important experiences may occur between tests.
Changes may not reflect unreliability.
The researcher should select an interval that reduces memory of the first test but is not so long that genuine change is likely to dominate the results.
Limitations of test-retest reliability
Participants may remember their responses
If participants recall their earlier answers, they may deliberately repeat them. This could make the measure appear more reliable than it really is.
Practice effects may occur
Participants may improve because they have completed the test before.
This is particularly relevant when the measure contains tasks with correct answers.
Genuine change may occur
A participant’s attitudes, knowledge or behaviour may change between the two occasions.
A difference in scores would then not necessarily show that the measure is unreliable.
Participants may not be available twice
Some participants may not return for the second measurement, creating a practical difficulty.
Situational conditions may differ
Tiredness, mood, noise or other circumstances may change between occasions and affect responses.
Improving test-retest reliability
If a measure produces inconsistent results, researchers should examine the likely source of the problem.
They may:
Rewrite ambiguous questions.
Clarify response options.
Standardise the instructions.
Apply the same testing conditions.
Use the same scoring criteria.
Choose a more suitable time interval.
Remove items that participants interpret inconsistently.
Pilot the revised measure.
Repeat the test-retest assessment after changes.
Researchers should not simply repeat the measure until a preferred result appears. They should make justified improvements to the measure or procedure.
Inter-observer reliability
Inter-observer reliability is the extent to which different observers produce consistent records of the same behaviour.
It is also relevant where different researchers code or classify the same material.
High agreement suggests that:
The categories are clearly defined.
Observers understand the recording system.
Personal interpretation has been reduced.
The data do not depend heavily on which observer collected them.
Low agreement suggests that the procedure or categories may need improvement.
The inter-observer reliability procedure
Researchers would:
Create operationalised behavioural or coding categories.
Train two or more observers to use them.
Ask the observers to record the same behaviour independently.
Compare the records produced.
Identify categories or events where observers disagreed.
Revise the categories or training if agreement is unsatisfactory.
Repeat the comparison where necessary.
Observers should initially record independently. If they discuss each event while it is occurring, one observer may influence the other.
Example of inter-observer reliability
Two observers record behaviour during a group activity.
They use the categories:
Shares a task-related object.
Gives task-related information.
Interrupts another participant.
Leaves the work area.
If both observers record similar frequencies for the same behaviours, the observation has greater inter-observer reliability.
If one observer records 15 interruptions while the other records only four, they may be applying the category differently.
Causes of low inter-observer reliability
Low agreement may be caused by:
Vague categories.
Overlapping categories.
Different interpretations of the same behaviour.
Poor observer training.
Events occurring too quickly.
Too many categories.
Different positions or views of the behaviour.
Unclear decisions about when an event begins and ends.
Inconsistent recording procedures.
The researcher should identify the particular problem rather than assuming that disagreement is caused by carelessness.
Improving inter-observer reliability
Researchers may:
Operationalise each category precisely.
Remove or revise overlapping categories.
Provide examples of behaviour that belongs in each category.
Explain what should not be included.
Train observers using the same practice material.
Reduce the number of categories.
Improve the observation schedule.
Position observers so that they can see the same behaviour.
Record behaviour independently.
Compare records after a pilot observation.
Repeat training where disagreement remains high.
Test-retest and inter-observer reliability compared
Feature | Test-retest reliability | Inter-observer reliability |
What is being assessed? | Consistency of a measure over time | Consistency between observers or coders |
Who is involved? | The same participants complete the measure twice | Two or more observers record the same behaviour |
What is compared? | Scores from two occasions | Records produced by different observers |
Common use | Questionnaires and psychological tests | Observations and content analysis |
Main possible problem | Memory, practice or genuine change over time | Subjective or unclear categories |
Main improvement | Clarify the measure and standardise testing | Clarify categories and train observers |
Choosing the appropriate type of reliability
Ask what kind of consistency is being questioned.
Are scores inconsistent across time?
Use test-retest reliability.
Example:
A questionnaire produces very different scores when the same participants complete it again two weeks later.
Do researchers disagree about what they observed?
Use inter-observer reliability.
Example:
Two observers classify the same interaction differently.
The method selected should match the problem in the scenario.
Reliability and standardisation
Standardisation is one of the main ways of improving reliability.
A standardised procedure may specify:
Exact instructions.
Timing.
Equipment.
Materials.
Order of activities.
Researcher responses.
Recording rules.
Scoring criteria.
When these aspects are consistent, the research procedure is less likely to vary between participants, investigators or replications.
Standardisation is especially useful in experiments, questionnaires and structured interviews.
Reliability and operationalisation
Clear operationalisation makes measurements easier to apply consistently.
Compare the following dependent variables:
“How good the participant’s memory was.”
“The number of words correctly recalled from a list of 20 within two minutes.”
The second measure is more precise. Researchers know exactly what to count and how long to measure performance.
Similarly, compare these behavioural categories:
“Acts aggressively.”
“Hits, kicks or pushes another person.”
The second category requires less personal interpretation.
Reliability and investigator effects
Investigator effects may reduce reliability if researchers:
Give different instructions.
Provide different amounts of help.
Use different scoring rules.
Interpret responses according to their expectations.
Apply categories inconsistently.
Standardised procedures, training and objective measurement can reduce these differences.
This reinforces the importance of reducing researcher influence on participants and data.
Reliability and validity
Reliability and validity are related but different.
Reliability concerns consistency.
Validity concerns whether the research measures what it intends to measure and supports an appropriate conclusion.
A measure can be reliable without being valid.
For example, a questionnaire might consistently produce the same score but fail to measure the intended psychological concept.
A measure that produces inconsistent results is unlikely to provide a dependable measure of the intended concept. However, consistency alone does not prove validity.
The distinction is developed further in assessing whether psychological research measures what it intends to measure.
Example of a reliable but invalid measure
A researcher wants to measure concentration and records the number of times participants look at a worksheet.
Different observers may reliably agree about how many times each participant looks at the sheet.
However, looking at a worksheet may not accurately represent concentration. A participant could look at the page while thinking about something else.
The measure may therefore be reliable but have questionable validity.
Reliability A-Level Psychology revision: identifying problems
Consider this investigation:
Two interviewers ask students about examination stress. One interviewer asks every prepared question in order. The other changes the wording, skips questions and provides additional prompts.
The interviews may have low reliability because the students do not experience the same procedure.
Possible improvements include:
Using the same questions.
Keeping the order consistent.
Preparing agreed prompts.
Training both interviewers.
Recording and coding responses consistently.
Applying reliability to a questionnaire scenario
Students complete a questionnaire containing the question, “Do you often find revision difficult and stressful?”
The question may reduce reliability because:
“Often” is not defined.
It asks about difficulty and stress in the same item.
Participants may interpret it differently.
The same participant may focus on different parts of the question on different occasions.
The researcher should divide it into separate questions and define a clear time period or response scale.
Applying reliability to an observation scenario
Two observers record “disruptive behaviour” during lessons. One records talking, leaving a seat and looking away from work. The other records only shouting and physical interference.
The category is not operationalised consistently.
The researcher should:
Define which observable actions count.
Provide examples.
Use the same recording schedule.
Train both observers.
compare their records during a pilot observation.
Conducting a reliability check
A researcher could use this process:
Identify the measurement or procedure requiring assessment.
Decide whether consistency over time or between observers is relevant.
Select test-retest or inter-observer reliability.
Apply the same measure or categories consistently.
Compare the resulting scores or records.
Identify sources of disagreement or inconsistency.
Revise questions, categories, instructions or scoring.
Train researchers where necessary.
Repeat the check.
Document the final procedure clearly.
Piloting provides an opportunity to test these procedures before the full investigation. This connects with using a trial investigation to identify weaknesses.
Improving reliability across methods
Method | Reliability improvement |
Laboratory or field experiment | Standardise instructions, timing, equipment and scoring |
Natural or quasi-experiment | Measure outcomes consistently and record the procedure clearly |
Observation | Use operationalised categories, observer training and inter-observer checks |
Questionnaire | Use clear questions, consistent response options and test-retest checks |
Structured interview | Ask the same prepared questions using agreed prompts |
Unstructured interview | Record procedures carefully and use clear coding criteria |
Correlation | Measure each co-variable consistently for every participant |
Content analysis | Use clear coding categories and compare independent coders |
Case study | Record how information was collected and apply consistent interpretation procedures |
Limits of improving reliability
Researchers should avoid assuming that greater standardisation is always easy or appropriate.
For example:
An unstructured interview is intended to respond flexibly to participants.
A naturalistic observation may contain unpredictable events.
A case study may develop in response to new information.
Repeating a test may change participant performance.
Human behaviour may genuinely change over time.
Researchers must improve consistency without changing the intended nature of the method.
A more reliable procedure may also be less natural if control and standardisation make the situation artificial. Reliability should therefore be considered alongside validity and the purpose of the investigation.
Answering reliability questions in examinations
Questions may ask you to:
Define reliability.
Identify a suitable way of measuring reliability.
Explain test-retest reliability.
Explain inter-observer reliability.
Apply reliability to a research scenario.
Suggest how a procedure could be made more reliable.
Distinguish reliability from validity.
Answering an “explain” question
A strong explanation includes the procedure and what the result would show.
Weak answer:
Give the test twice.
Stronger answer:
The same participants should complete the questionnaire on two separate occasions. The two sets of scores should then be compared. Similar scores would indicate that the questionnaire has good test-retest reliability.
Answering an “apply” question
Use the details in the scenario.
Weak answer:
Use inter-observer reliability.
Stronger answer:
A second observer should independently record the same playground behaviour using the same categories. The two records can then be compared to determine whether the observers classified the behaviour consistently.
Answering an “improve” question
Identify the source of unreliability and give a targeted solution.
Weak answer:
Make the observation more reliable.
Stronger answer:
The category “friendly behaviour” should be replaced with specific actions, such as smiling at another person or offering an object. This would reduce subjective interpretation and increase agreement between observers.
Planning reliable research
When developing a complete psychological investigation, researchers should:
Operationalise variables and behaviours clearly.
Prepare standardised instructions.
Use consistent materials and equipment.
Develop objective scoring rules.
Train investigators and observers.
Pilot the procedure.
Assess test-retest or inter-observer reliability where appropriate.
Revise unclear measures or categories.
Follow the agreed procedure during data collection.
Report the procedure clearly enough for replication.
Key Words 🔑
Key word | Student-friendly definition | How it may be used in an exam |
Reliability | The consistency of a research procedure or measurement. | Define reliability or explain whether findings are dependable. |
Test-retest reliability | The consistency of a measure when the same participants complete it on two occasions. | Explain how the reliability of a questionnaire or test could be assessed. |
Inter-observer reliability | The extent to which different observers produce consistent records of the same behaviour. | Explain how the reliability of an observation could be assessed. |
Observer | A researcher who watches and records behaviour. | Explain why clear categories and training are required. |
Standardisation | Keeping instructions, materials and procedures consistent. | Suggest how experimental, questionnaire or interview reliability could be improved. |
Operationalisation | Defining a variable or behaviour precisely so that it can be measured. | Replace vague measures or behavioural categories. |
Behavioural category | A clearly defined type of observable behaviour used in an observation. | Explain how observer agreement could be improved. |
Observation schedule | A prepared record sheet containing behavioural categories. | Suggest a consistent method for recording behaviour. |
Scoring criteria | Prepared rules explaining how responses will be awarded scores. | Reduce differences between researchers when marking responses. |
Replicability | The extent to which a study is described clearly enough to be repeated. | Explain why standardised procedures should be reported accurately. |
Objective measurement | A measurement based on clear observable or numerical rules rather than personal judgement. | Suggest how investigator interpretation can be reduced. |
Validity | The extent to which research measures what it intends to measure and supports appropriate conclusions. | Distinguish accuracy of measurement from consistency. |
Practice effect | Improvement caused by previous experience of a task. | Explain a possible limitation of test-retest reliability. |
Pilot study | A small-scale trial used to identify problems before the main investigation. | Explain how unclear procedures or categories can be improved. |
Common Mistakes ⚠️
Mistake: Defining reliability as whether a study measures what it intends to measure.
Why this is incorrect:This describes validity. Reliability concerns consistency.
How to improve:Use the question “Would the procedure or measure produce consistent results?”
Mistake: Saying that reliability applies only to experiments.
Why this is incorrect:Reliability is relevant to observations, questionnaires, interviews, correlations, content analyses and case studies as well as experiments.
How to improve:Identify the particular source of inconsistency within the method described.
Mistake: Describing test-retest reliability as using different participants.
Why this is incorrect:The same participants complete the same measure on two occasions.
How to improve:State clearly that the two sets of scores come from the same people.
Mistake: Saying that test-retest reliability involves repeating an entire experiment.
Why this is incorrect:Test-retest reliability assesses the consistency of a test or measure over time.
How to improve:Focus on administering the same measurement twice and comparing the scores.
Mistake: Ignoring practice or memory effects in test-retest reliability.
Why this is incorrect:Participants may remember answers or improve through previous experience, making scores appear more consistent.
How to improve:Explain why the interval between tests must be chosen carefully.
Mistake: Saying inter-observer reliability means one observer watches participants twice.
Why this is incorrect:Inter-observer reliability compares records produced by different observers.
How to improve:State that two or more observers record the same behaviour independently.
Mistake: Claiming that observers must produce identical records.
Why this is incorrect:Human observation may involve some disagreement. Reliability concerns whether the level of agreement is sufficiently consistent.
How to improve:Explain that researchers compare records and investigate important disagreements.
Mistake: Suggesting clearer categories without applying the improvement.
Why this is incorrect:The answer does not explain how the vague category should change.
How to improve:Replace the category with precise observable actions from the scenario.
Mistake: Claiming that a reliable measure must be valid.
Why this is incorrect:A procedure can consistently measure the wrong concept.
How to improve:State that reliability is necessary for dependable measurement but does not prove validity.
Mistake: Giving a generic improvement such as “repeat the study”.
Why this is incorrect:Repeating an unreliable procedure does not correct the source of inconsistency.
How to improve:Identify and improve the instructions, categories, measurement, scoring or researcher training.
Exam-Style Questions ✍️
Question 1
Define reliability in psychological research. (2 marks)
Question 2
Name the type of reliability assessed when the same questionnaire is given to the same participants on two separate occasions. (1 mark)
Question 3
Explain how a researcher could assess the test-retest reliability of a questionnaire measuring examination confidence. (3 marks)
Question 4
Two observers independently record aggressive behaviour in a playground.
Explain how the researchers could assess inter-observer reliability. (3 marks)
Question 5
A questionnaire includes the item:
“Do you often feel tired and unable to concentrate?”
Explain one reason why this question may reduce reliability and suggest an improvement. (4 marks)
Question 6
Two interviewers investigate students’ revision habits. One interviewer follows a prepared list of questions. The other changes the wording, asks additional questions and allows some participants more time to answer.
Explain two ways in which the reliability of the interviews could be improved. (4 marks)
Question 7
Two observers record behaviour during a group task.
Behavioural category | Observer A | Observer B |
Shares an object | 12 | 11 |
Gives task-related information | 8 | 9 |
Behaves helpfully | 15 | 6 |
Leaves the work area | 4 | 4 |
Identify the category showing the greatest disagreement and explain one likely reason for this disagreement. (3 marks)
Question 8
A researcher uses the category “shows interest” during an observation of classroom behaviour.
Explain why this category may produce low inter-observer reliability. Construct a more clearly operationalised behavioural category. (4 marks)
Question 9
Compare test-retest reliability and inter-observer reliability. In your answer, explain how each is assessed and identify one limitation of each procedure. (6 marks)
Question 10
A psychologist designs a study investigating whether background noise affects performance on a concentration task.
Explain how the psychologist could improve the reliability of the investigation. Refer to:
Instructions
Testing conditions
Measurement of concentration
Scoring
Investigator training
Piloting
(8 marks)



Comments