top of page

Reliability | AQA A-Level Psychology Revision

Updated: 6 days ago

For 7182 specification, first teach in September 2025


AQA A-Level Psychology | Free Revision Notes

Estimated study time: 55 minutes

Reliability concerns the consistency of psychological research. This Reliability A-Level Psychology revision page explains how reliability applies across experiments, observations, questionnaires, interviews and other research methods. You will learn how test-retest and inter-observer reliability are measured, why unreliable procedures weaken confidence in findings and how psychologists can improve consistency. Reliability is a required part of the scientific processes section of AQA A-Level Psychology research methods and is closely connected to standardisation, operationalisation and replication.


Learning Objectives 🎯

By the end of this revision page, you should be able to:

  • Define reliability in psychological research.

  • Explain how reliability applies across different research methods.

  • Explain how test-retest reliability is assessed.

  • Explain how inter-observer reliability is assessed.

  • Identify possible causes of unreliable measurements or procedures.

  • Suggest and justify practical ways of improving reliability.


Revision Notes 📚


What is reliability?

Reliability is the consistency of a research procedure or measurement.

A reliable procedure should produce similar results when it is used consistently under similar circumstances.

Reliability may concern whether:

  • The same measure produces consistent results on different occasions.

  • Different observers record the same behaviour consistently.

  • Instructions and procedures are applied in the same way.

  • Responses are scored using consistent criteria.

  • Equipment produces dependable measurements.

  • Another researcher could follow the procedure in the same way.

Reliability is relevant across all psychological research methods, not only experiments.

Researchers must first define exactly what they intend to measure. This connects reliability with creating precise and measurable variables.


Why does reliability matter?

Psychologists use research findings to draw conclusions about behaviour and mental processes. If the procedure or measurement is inconsistent, differences in the results may have been caused by the way the investigation was conducted rather than by genuine differences between participants or conditions.

For example, imagine that two researchers score the same set of responses differently. The final results would depend partly on which researcher completed the scoring.

This would make it difficult to know whether the recorded differences reflect:

  • Genuine participant behaviour.

  • Changes in the variable being studied.

  • Inconsistent instructions.

  • Different interpretations by researchers.

  • Unreliable equipment or scoring.

Reliable procedures increase confidence that results are not simply the product of inconsistent measurement.


Reliability and replicability

Reliability and replicability are connected but are not identical.

  • Reliability concerns whether a procedure or measure is consistent.

  • Replicability concerns whether a study is described clearly enough for another researcher to repeat it.

A carefully standardised and clearly reported procedure is easier to replicate. Replication can then help psychologists investigate whether a finding occurs consistently.

Replicability is one of the characteristics considered in the scientific features of psychological research.


Reliability across psychological research methods

Reliability can be considered in:

  • Experiments.

  • Observations.

  • Questionnaires.

  • Interviews.

  • Correlations.

  • Content analyses.

  • Case studies.

The exact source of unreliability depends on the method being used.

Research method

Possible reliability problem

Experiment

Participants receive different instructions or testing conditions

Observation

Observers interpret the same behaviour differently

Questionnaire

Questions are ambiguous or produce inconsistent responses

Interview

Interviewers ask questions or use prompts differently

Correlation

Co-variables are measured inconsistently

Content analysis

Researchers apply coding categories differently

Case study

Information is collected or interpreted without a clear procedure


Reliability in experiments

An experiment may be unreliable if its procedure is not applied consistently.

Possible problems include:

  • Participants receive different instructions.

  • Different amounts of time are allowed.

  • Equipment is used differently.

  • Some participants receive additional help.

  • Testing conditions vary unnecessarily.

  • The dependent variable is scored inconsistently.

  • The procedure is not described clearly enough to repeat.

For example, a psychologist investigates whether background noise affects memory. Some participants are given two minutes to recall words, while others are given three minutes.

The memory measure is not being applied consistently. Differences in recall scores may be caused by the time available rather than by the independent variable.


Improving reliability in experiments

Researchers may improve reliability by:

  • Preparing standardised instructions.

  • Using the same equipment.

  • Allowing the same amount of time.

  • Keeping the testing environment consistent.

  • Following a written procedure.

  • Operationalising the dependent variable clearly.

  • Using objective scoring criteria.

  • Training investigators.

  • Piloting the procedure.

  • Recording enough detail for replication.

These improvements use the procedures covered in randomisation and standardisation within research.


Reliability in observations

Reliability is particularly important in observations because researchers must decide how behaviour should be categorised and recorded.

Observers may disagree about:

  • Whether a behaviour occurred.

  • Which category should contain it.

  • When the behaviour began or ended.

  • Whether an action meets the category definition.

  • How many times the behaviour occurred.

For example, one observer may record an interaction as aggressive while another records it as playful.

If the category is simply labelled “aggression”, personal interpretation may influence the recording.

A clearer category might be:

Pushes, hits or kicks another person.

This describes observable actions and reduces the amount of judgement required.


Improving reliability in observations

Researchers may:

  • Construct clear behavioural categories.

  • Ensure categories do not overlap unnecessarily.

  • Define what counts as one occurrence.

  • Use an observation schedule.

  • Train observers using examples.

  • Allow observers to practise before the main study.

  • Use more than one observer.

  • Compare the observers’ records.

  • Revise unclear categories after a pilot study.

  • Record behaviour independently before comparing results.

The construction of clear categories is covered in planning how observable behaviour will be recorded.


Reliability in questionnaires

Questionnaires may produce unreliable data if participants interpret questions differently or give inconsistent responses.

Reliability problems may be caused by:

  • Ambiguous wording.

  • Undefined terms such as “often” or “regularly”.

  • Questions that contain more than one idea.

  • Unclear response options.

  • Different instructions.

  • Changes in how the questionnaire is administered.

  • Questions that depend heavily on a temporary mood or situation.

  • Inconsistent scoring.

For example:

Do you regularly feel tired and unable to concentrate?

This question includes two different experiences. A participant may regularly feel tired but have no difficulty concentrating.

It is therefore unclear how the participant should answer.


Improving reliability in questionnaires

Researchers may improve consistency by:

  • Using clear and unambiguous wording.

  • Asking one thing in each question.

  • Defining time periods precisely.

  • Providing suitable response options.

  • Giving all participants the same instructions.

  • Presenting questions in a consistent format.

  • Using clear scoring rules.

  • Piloting the questionnaire.

  • Using test-retest reliability to examine the consistency of responses.

These decisions are relevant when constructing effective self-report questions.


Reliability in interviews

Interviews may be unreliable if interviewers do not conduct them consistently.

Possible sources of unreliability include:

  • Asking different questions.

  • Changing the order of questions.

  • Using different follow-up prompts.

  • Giving some participants more time.

  • Explaining questions differently.

  • Recording answers inconsistently.

  • Interpreting similar responses differently.

A structured interview is generally easier to standardise because the interviewer uses prepared questions in a predetermined order.

An unstructured interview allows the interviewer to respond flexibly to each participant. This flexibility may produce detailed information, but it can also make the procedure more difficult to repeat consistently.


Improving reliability in interviews

Researchers may:

  • Prepare a standardised list of questions.

  • Use the same question order.

  • Develop agreed prompts.

  • Train interviewers.

  • Record interviews consistently.

  • Prepare clear coding or scoring criteria.

  • Ensure interviewers do not paraphrase questions differently.

  • Pilot the interview schedule.

The balance between structure and flexibility is explored in structured and unstructured interviewing.


Reliability in correlations

A correlation investigates the relationship between two co-variables.

The measurements used for both co-variables should be reliable. If one variable is measured inconsistently, the apparent relationship may not represent the genuine pattern between them.

For example, a psychologist investigates the relationship between sleep and concentration.

If sleep is measured using a different time period for different participants, or concentration is tested under inconsistent conditions, the resulting correlation may be misleading.

Researchers should:

  • Operationalise each co-variable clearly.

  • Use the same measurement procedure for all participants.

  • Apply consistent scoring.

  • Keep relevant testing conditions standardised.

  • Check whether measures produce dependable results.


Reliability in content analysis

In a content analysis, researchers place material into coding categories.

Reliability may be low if:

  • Categories are vague.

  • Categories overlap.

  • Researchers interpret content differently.

  • Coding rules are incomplete.

  • Researchers change how categories are applied.

For example, researchers analysing television programmes might create a category labelled “negative behaviour”. This is subjective and could be interpreted in several ways.

The category should be replaced with specific, observable examples that can be identified consistently.

Researchers can improve reliability by:

  • Producing clear coding categories.

  • Defining examples and boundaries.

  • Training coders.

  • Having more than one coder analyse the same material independently.

  • Comparing their records.

  • Revising categories where agreement is low.


Reliability in case studies

Case studies involve detailed investigations of an individual, group, organisation or event.

Reliability can be difficult to achieve because:

  • The case may be unique.

  • Information may come from several sources.

  • Researchers may interpret detailed material differently.

  • Flexible procedures may develop as the case progresses.

  • The exact circumstances may not be repeatable.

Researchers can still improve reliability by:

  • Describing the procedure clearly.

  • Using consistent interview or observation schedules where appropriate.

  • Recording how evidence was collected.

  • Applying clear coding procedures.

  • Reporting the method in sufficient detail.

  • Avoiding unsupported personal interpretation.

However, improving reliability does not make a unique case completely repeatable.


Test-retest reliability

Test-retest reliability measures the consistency of a psychological test or measure over time.

The same measure is given to the same participants on two separate occasions. The two sets of scores are then compared.

A strong positive relationship between the scores suggests that the measure has good test-retest reliability.


The test-retest procedure

A researcher would:

  1. Give a measure to a group of participants.

  2. Record each participant’s score.

  3. Wait for an appropriate period.

  4. Give the same measure to the same participants again.

  5. Record the second set of scores.

  6. Compare the scores from the two occasions.

  7. Judge whether participants obtained reasonably consistent results.

The procedure can be used with measures such as questionnaires and psychological tests.


Example of test-retest reliability

A psychologist develops a questionnaire measuring examination confidence.

The questionnaire is completed by the same 30 students on two occasions, two weeks apart.

If students who receive high scores on the first occasion also tend to receive high scores on the second occasion, the measure appears consistent over time.

If scores change greatly without a clear reason, the questionnaire may have low test-retest reliability.


Interpreting test-retest reliability

The researcher is interested in the consistency of the pattern of scores.

High test-retest reliability suggests that:

  • The measure is being interpreted consistently.

  • The scoring system is stable.

  • The measure is not producing highly variable results without explanation.

Low test-retest reliability may suggest that:

  • Questions are unclear.

  • Scoring rules are inconsistent.

  • Participants interpret items differently on each occasion.

  • The measure is strongly influenced by temporary circumstances.

  • The procedure changed between occasions.


Selecting an appropriate time interval

The interval between the two tests must be chosen carefully.

If the interval is very short:

  • Participants may remember their previous responses.

  • Familiarity may influence their second performance.

  • Practice effects may increase scores.

If the interval is very long:

  • The characteristic being measured may genuinely change.

  • Important experiences may occur between tests.

  • Changes may not reflect unreliability.

The researcher should select an interval that reduces memory of the first test but is not so long that genuine change is likely to dominate the results.


Limitations of test-retest reliability


Participants may remember their responses

If participants recall their earlier answers, they may deliberately repeat them. This could make the measure appear more reliable than it really is.


Practice effects may occur

Participants may improve because they have completed the test before.

This is particularly relevant when the measure contains tasks with correct answers.


Genuine change may occur

A participant’s attitudes, knowledge or behaviour may change between the two occasions.

A difference in scores would then not necessarily show that the measure is unreliable.


Participants may not be available twice

Some participants may not return for the second measurement, creating a practical difficulty.


Situational conditions may differ

Tiredness, mood, noise or other circumstances may change between occasions and affect responses.


Improving test-retest reliability

If a measure produces inconsistent results, researchers should examine the likely source of the problem.

They may:

  • Rewrite ambiguous questions.

  • Clarify response options.

  • Standardise the instructions.

  • Apply the same testing conditions.

  • Use the same scoring criteria.

  • Choose a more suitable time interval.

  • Remove items that participants interpret inconsistently.

  • Pilot the revised measure.

  • Repeat the test-retest assessment after changes.

Researchers should not simply repeat the measure until a preferred result appears. They should make justified improvements to the measure or procedure.


Inter-observer reliability

Inter-observer reliability is the extent to which different observers produce consistent records of the same behaviour.

It is also relevant where different researchers code or classify the same material.

High agreement suggests that:

  • The categories are clearly defined.

  • Observers understand the recording system.

  • Personal interpretation has been reduced.

  • The data do not depend heavily on which observer collected them.

Low agreement suggests that the procedure or categories may need improvement.


The inter-observer reliability procedure

Researchers would:

  1. Create operationalised behavioural or coding categories.

  2. Train two or more observers to use them.

  3. Ask the observers to record the same behaviour independently.

  4. Compare the records produced.

  5. Identify categories or events where observers disagreed.

  6. Revise the categories or training if agreement is unsatisfactory.

  7. Repeat the comparison where necessary.

Observers should initially record independently. If they discuss each event while it is occurring, one observer may influence the other.


Example of inter-observer reliability

Two observers record behaviour during a group activity.

They use the categories:

  • Shares a task-related object.

  • Gives task-related information.

  • Interrupts another participant.

  • Leaves the work area.

If both observers record similar frequencies for the same behaviours, the observation has greater inter-observer reliability.

If one observer records 15 interruptions while the other records only four, they may be applying the category differently.


Causes of low inter-observer reliability

Low agreement may be caused by:

  • Vague categories.

  • Overlapping categories.

  • Different interpretations of the same behaviour.

  • Poor observer training.

  • Events occurring too quickly.

  • Too many categories.

  • Different positions or views of the behaviour.

  • Unclear decisions about when an event begins and ends.

  • Inconsistent recording procedures.

The researcher should identify the particular problem rather than assuming that disagreement is caused by carelessness.


Improving inter-observer reliability

Researchers may:

  • Operationalise each category precisely.

  • Remove or revise overlapping categories.

  • Provide examples of behaviour that belongs in each category.

  • Explain what should not be included.

  • Train observers using the same practice material.

  • Reduce the number of categories.

  • Improve the observation schedule.

  • Position observers so that they can see the same behaviour.

  • Record behaviour independently.

  • Compare records after a pilot observation.

  • Repeat training where disagreement remains high.


Test-retest and inter-observer reliability compared

Feature

Test-retest reliability

Inter-observer reliability

What is being assessed?

Consistency of a measure over time

Consistency between observers or coders

Who is involved?

The same participants complete the measure twice

Two or more observers record the same behaviour

What is compared?

Scores from two occasions

Records produced by different observers

Common use

Questionnaires and psychological tests

Observations and content analysis

Main possible problem

Memory, practice or genuine change over time

Subjective or unclear categories

Main improvement

Clarify the measure and standardise testing

Clarify categories and train observers


Choosing the appropriate type of reliability

Ask what kind of consistency is being questioned.


Are scores inconsistent across time?

Use test-retest reliability.

Example:

A questionnaire produces very different scores when the same participants complete it again two weeks later.

Do researchers disagree about what they observed?

Use inter-observer reliability.

Example:

Two observers classify the same interaction differently.

The method selected should match the problem in the scenario.


Reliability and standardisation

Standardisation is one of the main ways of improving reliability.

A standardised procedure may specify:

  • Exact instructions.

  • Timing.

  • Equipment.

  • Materials.

  • Order of activities.

  • Researcher responses.

  • Recording rules.

  • Scoring criteria.

When these aspects are consistent, the research procedure is less likely to vary between participants, investigators or replications.

Standardisation is especially useful in experiments, questionnaires and structured interviews.


Reliability and operationalisation

Clear operationalisation makes measurements easier to apply consistently.

Compare the following dependent variables:

  • “How good the participant’s memory was.”

  • “The number of words correctly recalled from a list of 20 within two minutes.”

The second measure is more precise. Researchers know exactly what to count and how long to measure performance.

Similarly, compare these behavioural categories:

  • “Acts aggressively.”

  • “Hits, kicks or pushes another person.”

The second category requires less personal interpretation.


Reliability and investigator effects

Investigator effects may reduce reliability if researchers:

  • Give different instructions.

  • Provide different amounts of help.

  • Use different scoring rules.

  • Interpret responses according to their expectations.

  • Apply categories inconsistently.

Standardised procedures, training and objective measurement can reduce these differences.


Reliability and validity

Reliability and validity are related but different.

  • Reliability concerns consistency.

  • Validity concerns whether the research measures what it intends to measure and supports an appropriate conclusion.

A measure can be reliable without being valid.

For example, a questionnaire might consistently produce the same score but fail to measure the intended psychological concept.

A measure that produces inconsistent results is unlikely to provide a dependable measure of the intended concept. However, consistency alone does not prove validity.


Example of a reliable but invalid measure

A researcher wants to measure concentration and records the number of times participants look at a worksheet.

Different observers may reliably agree about how many times each participant looks at the sheet.

However, looking at a worksheet may not accurately represent concentration. A participant could look at the page while thinking about something else.

The measure may therefore be reliable but have questionable validity.


Reliability A-Level Psychology revision: identifying problems

Consider this investigation:

Two interviewers ask students about examination stress. One interviewer asks every prepared question in order. The other changes the wording, skips questions and provides additional prompts.

The interviews may have low reliability because the students do not experience the same procedure.

Possible improvements include:

  • Using the same questions.

  • Keeping the order consistent.

  • Preparing agreed prompts.

  • Training both interviewers.

  • Recording and coding responses consistently.


Applying reliability to a questionnaire scenario

Students complete a questionnaire containing the question, “Do you often find revision difficult and stressful?”

The question may reduce reliability because:

  • “Often” is not defined.

  • It asks about difficulty and stress in the same item.

  • Participants may interpret it differently.

  • The same participant may focus on different parts of the question on different occasions.

The researcher should divide it into separate questions and define a clear time period or response scale.


Applying reliability to an observation scenario

Two observers record “disruptive behaviour” during lessons. One records talking, leaving a seat and looking away from work. The other records only shouting and physical interference.

The category is not operationalised consistently.

The researcher should:

  • Define which observable actions count.

  • Provide examples.

  • Use the same recording schedule.

  • Train both observers.

  • compare their records during a pilot observation.


Conducting a reliability check

A researcher could use this process:

  1. Identify the measurement or procedure requiring assessment.

  2. Decide whether consistency over time or between observers is relevant.

  3. Select test-retest or inter-observer reliability.

  4. Apply the same measure or categories consistently.

  5. Compare the resulting scores or records.

  6. Identify sources of disagreement or inconsistency.

  7. Revise questions, categories, instructions or scoring.

  8. Train researchers where necessary.

  9. Repeat the check.

  10. Document the final procedure clearly.

Piloting provides an opportunity to test these procedures before the full investigation. This connects with using a trial investigation to identify weaknesses.


Improving reliability across methods

Method

Reliability improvement

Laboratory or field experiment

Standardise instructions, timing, equipment and scoring

Natural or quasi-experiment

Measure outcomes consistently and record the procedure clearly

Observation

Use operationalised categories, observer training and inter-observer checks

Questionnaire

Use clear questions, consistent response options and test-retest checks

Structured interview

Ask the same prepared questions using agreed prompts

Unstructured interview

Record procedures carefully and use clear coding criteria

Correlation

Measure each co-variable consistently for every participant

Content analysis

Use clear coding categories and compare independent coders

Case study

Record how information was collected and apply consistent interpretation procedures


Limits of improving reliability

Researchers should avoid assuming that greater standardisation is always easy or appropriate.

For example:

  • An unstructured interview is intended to respond flexibly to participants.

  • A naturalistic observation may contain unpredictable events.

  • A case study may develop in response to new information.

  • Repeating a test may change participant performance.

  • Human behaviour may genuinely change over time.

Researchers must improve consistency without changing the intended nature of the method.

A more reliable procedure may also be less natural if control and standardisation make the situation artificial. Reliability should therefore be considered alongside validity and the purpose of the investigation.


Answering reliability questions in examinations

Questions may ask you to:

  • Define reliability.

  • Identify a suitable way of measuring reliability.

  • Explain test-retest reliability.

  • Explain inter-observer reliability.

  • Apply reliability to a research scenario.

  • Suggest how a procedure could be made more reliable.

  • Distinguish reliability from validity.


Answering an “explain” question

A strong explanation includes the procedure and what the result would show.

Weak answer:

Give the test twice.

Stronger answer:

The same participants should complete the questionnaire on two separate occasions. The two sets of scores should then be compared. Similar scores would indicate that the questionnaire has good test-retest reliability.

Answering an “apply” question

Use the details in the scenario.

Weak answer:

Use inter-observer reliability.

Stronger answer:

A second observer should independently record the same playground behaviour using the same categories. The two records can then be compared to determine whether the observers classified the behaviour consistently.

Answering an “improve” question

Identify the source of unreliability and give a targeted solution.

Weak answer:

Make the observation more reliable.

Stronger answer:

The category “friendly behaviour” should be replaced with specific actions, such as smiling at another person or offering an object. This would reduce subjective interpretation and increase agreement between observers.

Planning reliable research

  1. Operationalise variables and behaviours clearly.

  2. Prepare standardised instructions.

  3. Use consistent materials and equipment.

  4. Develop objective scoring rules.

  5. Train investigators and observers.

  6. Pilot the procedure.

  7. Assess test-retest or inter-observer reliability where appropriate.

  8. Revise unclear measures or categories.

  9. Follow the agreed procedure during data collection.

  10. Report the procedure clearly enough for replication.


Key Words 🔑

Key word

Student-friendly definition

How it may be used in an exam

Reliability

The consistency of a research procedure or measurement.

Define reliability or explain whether findings are dependable.

Test-retest reliability

The consistency of a measure when the same participants complete it on two occasions.

Explain how the reliability of a questionnaire or test could be assessed.

Inter-observer reliability

The extent to which different observers produce consistent records of the same behaviour.

Explain how the reliability of an observation could be assessed.

Observer

A researcher who watches and records behaviour.

Explain why clear categories and training are required.

Standardisation

Keeping instructions, materials and procedures consistent.

Suggest how experimental, questionnaire or interview reliability could be improved.

Operationalisation

Defining a variable or behaviour precisely so that it can be measured.

Replace vague measures or behavioural categories.

Behavioural category

A clearly defined type of observable behaviour used in an observation.

Explain how observer agreement could be improved.

Observation schedule

A prepared record sheet containing behavioural categories.

Suggest a consistent method for recording behaviour.

Scoring criteria

Prepared rules explaining how responses will be awarded scores.

Reduce differences between researchers when marking responses.

Replicability

The extent to which a study is described clearly enough to be repeated.

Explain why standardised procedures should be reported accurately.

Objective measurement

A measurement based on clear observable or numerical rules rather than personal judgement.

Suggest how investigator interpretation can be reduced.

Validity

The extent to which research measures what it intends to measure and supports appropriate conclusions.

Distinguish accuracy of measurement from consistency.

Practice effect

Improvement caused by previous experience of a task.

Explain a possible limitation of test-retest reliability.

Pilot study

A small-scale trial used to identify problems before the main investigation.

Explain how unclear procedures or categories can be improved.


Common Mistakes ⚠️


Mistake: Defining reliability as whether a study measures what it intends to measure.

Why this is incorrect:This describes validity. Reliability concerns consistency.

How to improve:Use the question “Would the procedure or measure produce consistent results?”


Mistake: Saying that reliability applies only to experiments.

Why this is incorrect:Reliability is relevant to observations, questionnaires, interviews, correlations, content analyses and case studies as well as experiments.

How to improve:Identify the particular source of inconsistency within the method described.


Mistake: Describing test-retest reliability as using different participants.

Why this is incorrect:The same participants complete the same measure on two occasions.

How to improve:State clearly that the two sets of scores come from the same people.


Mistake: Saying that test-retest reliability involves repeating an entire experiment.

Why this is incorrect:Test-retest reliability assesses the consistency of a test or measure over time.

How to improve:Focus on administering the same measurement twice and comparing the scores.


Mistake: Ignoring practice or memory effects in test-retest reliability.

Why this is incorrect:Participants may remember answers or improve through previous experience, making scores appear more consistent.

How to improve:Explain why the interval between tests must be chosen carefully.


Mistake: Saying inter-observer reliability means one observer watches participants twice.

Why this is incorrect:Inter-observer reliability compares records produced by different observers.

How to improve:State that two or more observers record the same behaviour independently.


Mistake: Claiming that observers must produce identical records.

Why this is incorrect:Human observation may involve some disagreement. Reliability concerns whether the level of agreement is sufficiently consistent.

How to improve:Explain that researchers compare records and investigate important disagreements.


Mistake: Suggesting clearer categories without applying the improvement.

Why this is incorrect:The answer does not explain how the vague category should change.

How to improve:Replace the category with precise observable actions from the scenario.


Mistake: Claiming that a reliable measure must be valid.

Why this is incorrect:A procedure can consistently measure the wrong concept.

How to improve:State that reliability is necessary for dependable measurement but does not prove validity.


Mistake: Giving a generic improvement such as “repeat the study”.

Why this is incorrect:Repeating an unreliable procedure does not correct the source of inconsistency.

How to improve:Identify and improve the instructions, categories, measurement, scoring or researcher training.


Exam-Style Questions ✍️


Question 1

Define reliability in psychological research. (2 marks)



Question 2

Name the type of reliability assessed when the same questionnaire is given to the same participants on two separate occasions. (1 mark)



Question 3

Explain how a researcher could assess the test-retest reliability of a questionnaire measuring examination confidence. (3 marks)



Question 4

Two observers independently record aggressive behaviour in a playground.

Explain how the researchers could assess inter-observer reliability. (3 marks)



Question 5

A questionnaire includes the item:

“Do you often feel tired and unable to concentrate?”

Explain one reason why this question may reduce reliability and suggest an improvement. (4 marks)



Question 6

Two interviewers investigate students’ revision habits. One interviewer follows a prepared list of questions. The other changes the wording, asks additional questions and allows some participants more time to answer.

Explain two ways in which the reliability of the interviews could be improved. (4 marks)



Question 7

Two observers record behaviour during a group task.

Behavioural category

Observer A

Observer B

Shares an object

12

11

Gives task-related information

8

9

Behaves helpfully

15

6

Leaves the work area

4

4

Identify the category showing the greatest disagreement and explain one likely reason for this disagreement. (3 marks)



Question 8

A researcher uses the category “shows interest” during an observation of classroom behaviour.

Explain why this category may produce low inter-observer reliability. Construct a more clearly operationalised behavioural category. (4 marks)



Question 9

Compare test-retest reliability and inter-observer reliability. In your answer, explain how each is assessed and identify one limitation of each procedure. (6 marks)



Question 10

A psychologist designs a study investigating whether background noise affects performance on a concentration task.

Explain how the psychologist could improve the reliability of the investigation. Refer to:

  • Instructions

  • Testing conditions

  • Measurement of concentration

  • Scoring

  • Investigator training

  • Piloting

(8 marks)

Recent Posts

See All

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page