Content analysis | AQA A-Level Psychology Revision
- Revision Notes
- Aug 4
- 27 min read
Updated: 6 days ago
For 7182 specification, first teach in September 2025
AQA A-Level Psychology | Free Revision Notes
Estimated study time: 50 minutes
These Content analysis A-Level Psychology revision notes explain how psychologists systematically examine communication and convert its content into analysable data. You will learn how researchers select material, create coding categories, record instances of content and check whether coding is reliable. You will also evaluate the method’s validity, objectivity, ethics and practical usefulness. AQA requires knowledge of content analysis as a research method and coding as a data-analysis technique.
Learning Objectives 🎯
By the end of this revision page, you should be able to:
Define content analysis.
Explain how researchers select and examine material.
Explain how coding categories are constructed and used.
Apply a coding system to unfamiliar material.
Explain how reliability and validity may be assessed or improved.
Evaluate content analysis as a psychological research method.
Revision Notes 📚
Content analysis A-Level Psychology revision overview
Content analysis is a research method used to examine communication systematically.
The material analysed may include:
Written text.
Spoken language.
Images.
Advertisements.
Television programmes.
Films.
Newspaper reports.
Social-media posts.
Interview transcripts.
Diaries.
Letters.
Recorded conversations.
Public information materials.
Researchers identify features of the material and record them using a coding system.
The central process is:
Select material → construct coding categories → apply the categories → record frequencies or patterns → interpret the findings
Specification boundary
The AQA specification requires knowledge of:
Content analysis.
Coding in content analysis.
Strengths and limitations of research methods.
Reliability and validity across methods.
The specification does not prescribe a single coding system or named content-analysis study.
Researchers must therefore design categories that suit the particular aim and material being investigated.
What is content?
In content analysis, content means the material being examined.
This might be communication produced by:
Individuals.
Groups.
Organisations.
News publishers.
Advertisers.
Online communities.
Researchers and participants.
For example, a psychologist might analyse:
Newspaper descriptions of mental health.
Advertisements aimed at teenagers.
Transcripts of interviews about examination stress.
Children’s television programmes.
Posts made within an online support forum.
Conversations recorded during a group task.
The researcher does not necessarily communicate directly with the people who produced the material.
What is analysis?
Analysis involves breaking the material into identifiable parts and looking for patterns.
The researcher may investigate:
How often a particular idea occurs.
Whether language is positive or negative.
Which people or groups are represented.
How frequently certain behaviours appear.
Whether content changes over time.
Differences between two sources.
Recurring themes within participant accounts.
A researcher should begin with a clear aim.
For example:
To investigate how examination stress is presented in online articles aimed at students.
The coding categories should then measure features relevant to that aim.
Content analysis as a research method
Content analysis can be used as the main method in an investigation.
For example, a researcher could:
Select 100 newspaper articles.
Develop categories describing how mental health is presented.
Code each article.
Compare the frequency of each category.
Draw conclusions about patterns within the selected articles.
Content analysis may also be used to analyse data collected through another method.
For example:
Open responses from written self-report questions
Transcripts from structured or unstructured discussions
Notes collected during an observation.
Written records within an in-depth investigation of an individual or group
Content analysis can therefore be both:
A method for investigating existing material.
A technique for organising qualitative data collected through another method.
Primary and secondary material
Content analysis may use material collected specifically for the investigation.
This would be primary data.
For example:
Interviews conducted by the researcher.
Diaries written for the study.
Conversations recorded during a research task.
It may also use material that already existed.
This would be secondary data.
For example:
Archived newspapers.
Published advertisements.
Existing television programmes.
Historical letters.
Public social-media posts.
The distinction is explored fully in Primary and secondary data.
Content analysis and observation
Content analysis and observation both involve recording identifiable features.
However, they are not identical.
Content analysis | Observation |
Examines recorded communication or material | Watches and records behaviour |
May analyse text, images or recordings | Usually focuses on observable actions |
Material may already exist | Behaviour is usually recorded during a study |
Uses coding categories | Uses behavioural categories |
May investigate content from an earlier time | Usually studies behaviour occurring during the observation |
A video recording could be used in either method.
The classification depends on the research aim.
For example:
Counting aggressive actions in a playground video resembles observation.
Analysing how aggression is represented in a television programme is content analysis.
The construction of observable categories is developed further in behavioural categories, event sampling and time sampling.
Coding in content analysis
What is coding?
Coding is the process of placing parts of the content into defined categories.
The researcher:
Identifies a relevant unit of content.
Decides which category applies.
Records the category.
Repeats the procedure across the selected material.
Coding converts complex material into a form that can be organised and analysed.
For example, a researcher studying articles about examination stress might code references to:
Academic workload.
Time pressure.
Physical symptoms.
Emotional responses.
Support from teachers.
Support from family.
Coping strategies.
What is a coding category?
A coding category is a clearly defined type of content that the researcher records.
A coding category might represent:
A word.
A phrase.
An idea.
A behaviour.
A person.
An image.
A topic.
A type of emotional expression.
For example, the category physical symptom might include references to:
Headaches.
Tiredness.
Nausea.
Shaking.
Sleep difficulty.
The category should be defined precisely enough for different coders to apply it consistently.
Coding unit
A coding unit is the piece of material to which the coding category is applied.
Depending on the research aim, the coding unit might be:
One word.
One sentence.
One paragraph.
One image.
One scene.
One social-media post.
One complete article.
One participant response.
The researcher should decide on the unit before coding begins.
For example:
Each complete social-media post will be treated as one coding unit.
This creates a consistent basis for analysis.
Why coding units matter
Suppose one researcher counts individual words while another counts complete paragraphs.
They may produce very different results from the same material.
Clear coding units help ensure that:
Every coder examines the same amount of material.
Frequencies have a consistent meaning.
Results can be compared.
The procedure can be replicated.
The coding unit must suit the research question.
A single word may be appropriate for investigating the frequency of particular terms, but not for understanding a complex emotional account.
Constructing coding categories
Coding categories should be created from the research aim.
Suppose the aim is:
To investigate how friendship is represented in programmes aimed at young people.
Possible categories might include:
Cooperation.
Conflict.
Emotional support.
Exclusion.
Loyalty.
Deception.
The researcher must then define each category.
For example:
Category | Operational definition |
Cooperation | Two or more characters working together towards the same goal |
Conflict | A verbal or physical disagreement between characters |
Emotional support | One character comforting, reassuring or encouraging another |
Exclusion | A character being deliberately prevented from joining an activity |
Loyalty | A character defending or continuing to support a friend |
Deception | A character intentionally giving another character false information |
These definitions make the categories observable and recordable.
Operationalising coding categories
Operationalisation means defining a category in a specific way so that researchers know exactly when it applies.
A vague category might be:
Positive friendship.
Different coders may interpret this differently.
A clearer category might be:
A character offers help, reassurance or practical support to someone identified as a friend.
The second definition tells the coder which content should be recorded.
Clear operationalisation improves:
Consistency.
Replication.
Objectivity.
Inter-coder reliability.
The wider process is covered in Variables and operationalisation.
Categories should be clear
A category should avoid vague terms such as:
Unpleasant.
Normal.
Positive.
Bad behaviour.
Friendly.
Appropriate.
These words require considerable personal interpretation.
A clearer category describes the exact content to be identified.
For example:
Instead of:
Negative language.
Use:
An insult, threat or direct criticism aimed at another person.
Categories should be distinct
Coding categories should be sufficiently different from one another.
Suppose a researcher uses:
Academic worry.
School worry.
Examination worry.
The categories overlap substantially.
Coders may disagree about where to place a statement such as:
“I am worried I will fail my psychology examination.”
More distinct categories could include:
Fear of failure.
Time pressure.
Workload.
Concern about others’ expectations.
Categories should cover relevant content
The coding system should include the important types of content identified in the material.
If relevant content does not fit any category, the coding system may be incomplete.
Researchers might include an other category during an early trial.
However, frequent use of “other” suggests that the main categories need to be revised.
Categories should suit the aim
A coding system is not automatically useful simply because it is detailed.
Suppose the aim concerns the representation of gender in advertising.
Categories about:
Background colour.
Music volume.
Length of advertisement.
may not address the aim unless the researcher can explain their relevance.
Every category should contribute to answering the research question.
Categories may be developed before analysis
Researchers may create categories before examining all the selected material.
This can be useful when:
The research aim is specific.
Earlier research suggests relevant categories.
The material has a predictable structure.
Quantitative comparison is required.
Using categories planned in advance may improve consistency.
However, important forms of content might be overlooked if the researcher’s expectations are too narrow.
Categories may be refined after examining a sample
Researchers may first examine a small sample of the material.
This can help them identify:
Common themes.
Ambiguous content.
Missing categories.
Overlapping categories.
Difficult coding units.
The coding system can then be revised before the main analysis.
This resembles a small-scale trial of a research procedure.
Coding frequency
A researcher may count how frequently each category occurs.
For example:
Category | Frequency across 50 articles |
Academic workload | 31 |
Time pressure | 27 |
Physical symptoms | 18 |
Family expectations | 12 |
Teacher support | 9 |
Coping strategies | 22 |
The frequencies can be used to compare categories.
The researcher might conclude that academic workload appeared more frequently than family expectations within the selected articles.
They should not automatically conclude that academic workload is the most important cause of stress in the wider population.
The analysis concerns the content selected, not necessarily reality outside that material.
Recording presence or absence
Researchers may record whether a category is present within each item.
For example:
Article | Workload | Time pressure | Physical symptoms | Coping strategy |
1 | Present | Present | Absent | Present |
2 | Present | Absent | Present | Absent |
3 | Absent | Present | Absent | Present |
This method prevents one long article from dominating the findings simply because it repeats the same idea many times.
The best recording approach depends on the research aim.
Recording intensity or ratings
Researchers may also rate features of content.
For example, the emotional tone of an article might be rated using a scale:
1: Very negative
2: Negative
3: Neutral
4: Positive
5: Very positive
However, rating scales require judgement.
Researchers need:
Clear descriptors.
Examples for each rating.
Coder training.
Checks of agreement.
A numerical scale does not automatically make a judgement objective.
A content-analysis procedure
Step 1: State the research aim
The aim should identify:
The material.
The feature being investigated.
Any comparison being made.
For example:
To investigate how coping with examination stress is represented in online articles aimed at A-Level students.
Step 2: Select the source material
The researcher should decide:
Which materials will be included.
Which dates will be covered.
Which publishers or sources will be used.
How many items will be analysed.
Why the selected material is appropriate.
For example:
Fifty student-focused articles published during the six months before the A-Level examination period will be selected.
Step 3: Decide how the material will be sampled
Researchers may be unable to analyse every relevant item.
They might select:
Every article within a defined period.
A random sample of eligible items.
Every fifth item.
Equal numbers from different sources.
Material meeting specific inclusion criteria.
The sampling procedure affects the representativeness of the findings.
The wider implications are covered in Populations and samples.
Step 4: Choose the coding unit
The researcher decides what will count as one unit.
Examples include:
One article.
One paragraph.
One social-media post.
One scene.
One participant response.
The unit should be used consistently.
Step 5: Develop coding categories
The categories should:
Address the research aim.
Be clearly defined.
Avoid unnecessary overlap.
Cover relevant content.
Be practical to apply.
Step 6: Create a coding frame
A coding frame is the organised set of categories and instructions used by coders.
It may include:
Category names.
Operational definitions.
Examples.
Exclusions.
Recording rules.
Guidance for ambiguous cases.
For example:
Code | Category | Include | Do not include |
A | Workload | References to amount of schoolwork or revision | General dislike of school |
B | Time pressure | References to deadlines or insufficient time | General busyness without a deadline |
C | Physical symptoms | Headaches, tiredness, nausea, sleep difficulty | Emotional worry without a physical symptom |
D | Coping strategy | Actions intended to reduce or manage stress | Descriptions of stress without a response |
Step 7: Pilot the coding frame
A small amount of material should be coded before the main analysis.
The pilot may reveal:
Categories that are unclear.
Content that fits several categories.
Important missing categories.
Inconsistent coding decisions.
Instructions that are difficult to apply.
The researcher can revise the coding frame before analysing the full sample.
Step 8: Train coders
Coders should understand:
The research aim.
The coding units.
The category definitions.
The recording procedure.
How to treat repeated content.
How to manage ambiguous examples.
Training should include practice material and discussion of disagreements.
Step 9: Code the material
Each unit is examined and the relevant categories are recorded.
Where more than one coder is used, they should initially code independently.
Independent coding allows agreement to be checked.
If coders discuss every item before recording it, the researcher cannot determine whether they would have reached the same decision independently.
Step 10: Check coding reliability
The researchers compare the coders’ decisions.
High agreement suggests that:
The categories are clear.
Coders have been trained adequately.
Personal interpretation has been reduced.
The coding procedure is replicable.
Low agreement suggests that:
Definitions are ambiguous.
Categories overlap.
Coders require further training.
Some content is difficult to classify.
Step 11: Summarise the findings
Researchers may produce:
Frequencies.
Percentages.
Tables.
Bar charts.
Written summaries.
Comparisons between sources.
Descriptions of recurring themes.
The selected presentation should suit the form of data collected.
Step 12: Interpret cautiously
The researcher should distinguish between:
What appeared in the selected content.
What can be concluded about the wider population or society.
For example:
Coping strategies appeared in 22 of the 50 selected articles.
This does not prove:
Most students use effective coping strategies.
The content may not accurately reflect students’ real behaviour.
Worked coding example
Research aim
To investigate how students describe difficulties with independent revision.
Source material
The researcher analyses 40 open questionnaire responses.
Coding unit
Each complete questionnaire response is treated as one unit.
Coding categories
Category | Definition |
Time management | Difficulty planning revision time or meeting deadlines |
Distraction | Attention being diverted by phones, people or entertainment |
Lack of understanding | Difficulty understanding the subject content |
Lack of motivation | Difficulty beginning or continuing revision |
Resource difficulty | Lack of suitable notes, equipment or study space |
Example response
“I keep checking my phone and then realise I have wasted most of the evening. I also find it hard to start when I do not understand the topic.”
This response could be coded as:
Distraction.
Lack of understanding.
Lack of motivation, only if the coding definition includes difficulty starting.
The coder should not add a category simply because it seems plausible.
Every decision must follow the operational definitions.
Example frequency table
Suppose the 40 responses produce the following results:
Category | Number of responses containing the category |
Time management | 21 |
Distraction | 28 |
Lack of understanding | 16 |
Lack of motivation | 24 |
Resource difficulty | 8 |
The researcher can report that distraction was the most frequently coded difficulty in this sample.
They cannot conclude that distraction definitely causes revision difficulties or that it is the most important issue for every student.
Quantitative content analysis
Producing quantitative data
Content analysis can produce quantitative data by recording:
Frequencies.
Totals.
Percentages.
Ratings.
Presence or absence.
For example:
28 of 40 responses contained a reference to distraction.
The researcher could calculate:
2840×100=70%\frac{28}{40}\times100=70\%4028×100=70%
Therefore, 70% of the analysed responses contained a reference coded as distraction.
Strengths of quantitative coding
Quantitative coding allows researchers to:
Summarise large amounts of material.
Make comparisons.
Identify common patterns.
Present findings clearly.
Apply statistical analysis where appropriate.
Replicate the procedure.
However, a frequency does not explain why content appeared or what it meant to the person producing it.
Limitation of counting frequency
The most frequent category is not automatically:
The most important.
The most psychologically significant.
The most accurate.
The cause of another pattern.
Representative of wider society.
One idea may appear frequently because:
A source repeats it.
Articles copy similar language.
The coding category is broad.
The selected sample favours a particular perspective.
Frequency should be interpreted within context.
Qualitative content analysis
Examining meaning and themes
Content analysis may also produce qualitative data.
Researchers may describe:
Recurring themes.
Different interpretations.
Relationships between ideas.
The context surrounding a statement.
How participants explain their experience.
For example, a researcher may identify that students describe revision distraction as involving:
Habitual phone checking.
Fear of missing messages.
Boredom.
Avoidance of difficult topics.
This provides greater depth than reporting only the number of references to distraction.
Strength of qualitative analysis
Qualitative analysis may:
Preserve participant meaning.
Identify unexpected ideas.
Show how themes are connected.
Provide context.
Avoid reducing every response to a simple count.
It may be especially useful when little is known about a topic.
Limitation of qualitative analysis
Qualitative interpretation relies heavily on researcher judgement.
Researchers may disagree about:
Which themes are present.
Which quotations are important.
How categories should be named.
What a statement means.
Whether an example supports the conclusion.
This may reduce objectivity and reliability.
Combining quantitative and qualitative analysis
A content analysis may use both forms of data.
For example, the researcher could:
Count how many responses refer to distraction.
Describe the different ways participants explain that distraction.
Include brief examples.
Compare themes across groups.
This combines:
Quantitative summary.
Qualitative depth.
The distinction between the two types is explored in Quantitative and qualitative data.
Reliability in content analysis
What is reliability?
Reliability concerns consistency.
A content-analysis procedure is reliable when coders apply the coding system consistently.
Questions about reliability include:
Would the same coder reach the same decisions later?
Would another coder classify the material similarly?
Are category definitions clear?
Can the coding procedure be repeated?
Inter-coder reliability
Inter-coder reliability, also called inter-rater reliability, concerns the level of agreement between people coding the same material.
A common procedure is:
Two coders receive the same coding frame.
They code the same sample independently.
Their coding decisions are compared.
The level of agreement is assessed.
Categories are revised if agreement is too low.
A high level of agreement suggests that coding is not dependent on one person’s interpretation.
This is closely related to inter-observer reliability and consistency across research methods.
Example of agreement
Suppose two researchers code 20 items.
They agree on 17 items and disagree on 3.
The percentage agreement is:
1720×100=85%\frac{17}{20}\times100=85\%2017×100=85%
This produces 85% agreement.
The percentage provides a simple indication of consistency.
However, researchers should still examine why the disagreements occurred.
Causes of low coding reliability
Low agreement may result from:
Vague categories.
Overlapping definitions.
Too many categories.
Inadequate training.
Ambiguous source material.
Different interpretations of context.
Inconsistent coding units.
Unclear instructions about repeated content.
Improving coding reliability
Reliability may be improved by:
Operationalising categories clearly.
Giving examples and exclusions.
Training coders.
Piloting the coding frame.
Reducing category overlap.
Using consistent coding units.
Coding independently.
Reviewing disagreements.
Repeating the reliability check after revision.
Reliability does not guarantee validity
Two coders may apply the same category consistently, but the category may not measure the intended concept.
For example, researchers might define “anxiety” only as use of the word “worried”.
They could record this reliably.
However, people may describe anxiety using:
Fear.
Panic.
Tension.
Physical symptoms.
Avoidance.
The category may therefore have limited validity despite high reliability.
Validity in content analysis
What is validity?
Validity concerns whether the coding system and interpretation genuinely represent the feature the researcher claims to investigate.
A valid content analysis requires:
Appropriate source material.
Categories that match the aim.
Accurate coding.
Conclusions supported by the material.
Careful interpretation of context.
Validity of the source material
The researcher must ask whether the selected material represents the issue being studied.
For example, newspaper articles may show:
How newspapers present mental health.
They do not necessarily show:
How people with mental health difficulties experience their condition.
How the whole population understands mental health.
How accurate the articles are.
The researcher’s conclusion must match the source.
Validity of categories
Categories should capture the intended concept.
Suppose a researcher investigates supportive behaviour in interview transcripts.
A category that records only the word “help” may miss statements such as:
“She listened to me.”
“He stayed with me.”
“They reassured me.”
“My teacher gave me extra time.”
The category may need to include a wider range of relevant content.
Loss of context
Coding may remove words or statements from their context.
Consider:
“I thought the exam would be impossible, but it was not.”
A coder searching for the word “impossible” might classify this as a negative description.
The full sentence communicates a more positive final judgement.
Researchers should examine enough surrounding content to interpret the unit correctly.
Researcher interpretation
Researchers decide:
Which material to select.
Which categories to create.
How categories are defined.
Which examples are included.
How frequencies are interpreted.
Which themes are emphasised.
These decisions may reflect the researcher’s expectations.
Content analysis is systematic, but it is not automatically free from bias.
Improving validity
Validity may be improved by:
Selecting material relevant to the aim.
Using clearly justified inclusion criteria.
Piloting categories.
Including sufficient context.
Using more than one researcher.
Considering alternative interpretations.
Combining frequency counts with qualitative explanation.
Avoiding conclusions beyond the selected material.
Comparing findings from different sources or methods.
The wider principles are developed in Validity.
Strengths of content analysis
Strength: existing material can be studied
Content analysis can investigate material that has already been produced.
Researchers may analyse:
Historical records.
Published media.
Archived interviews.
Public communications.
Existing programmes.
Earlier written accounts.
This allows psychologists to study material from:
Different time periods.
Different cultures.
Groups that are difficult to access.
Events that cannot be recreated.
Strength: behaviour may be unaffected by the researcher
When researchers analyse material that already exists, its creators may not have known that a psychological study would later examine it.
Their original communication cannot be changed by:
Demand characteristics.
Investigator effects during data collection.
Awareness of the later coding process.
For example, a historical diary cannot alter its content because it is being analysed years later.
However, the material may originally have been written for an audience, which could still have influenced its content.
Strength: large amounts of material can be summarised
Coding allows researchers to organise extensive material into manageable categories.
Hundreds of:
Articles.
Posts.
Advertisements.
Transcript sections.
Television scenes.
can be compared using the same coding frame.
This makes it possible to identify broad patterns that may be difficult to recognise through informal reading.
Strength: systematic procedure
A content analysis can use:
Defined coding units.
Operationalised categories.
Standard instructions.
Independent coders.
Reliability checks.
This makes the process more systematic and transparent than relying on personal impressions.
Another researcher can examine:
How material was selected.
How categories were defined.
How coding decisions were made.
Strength: quantitative data can be produced
Frequency counts and percentages allow researchers to:
Compare categories.
Compare sources.
Examine changes over time.
Present findings clearly.
Apply numerical analysis.
This may increase objectivity because conclusions are linked to recorded data.
However, numerical presentation does not remove subjective choices made when constructing categories.
Strength: qualitative meaning can be retained
Content analysis does not have to reduce every response to a number.
Researchers can examine:
Context.
Themes.
Explanations.
Contradictions.
Participant language.
This allows content analysis to combine systematic organisation with detailed interpretation.
Strength: non-intrusive research
Analysing existing material may require no direct contact with participants.
This can reduce:
Participant burden.
Disruption.
The need to arrange appointments.
Behavioural reactivity.
Some forms of physical risk.
Ethical responsibilities remain, particularly where material is private or identifiable.
Strength: comparisons across time
Researchers may compare content from different periods.
For example, they could investigate whether representations of psychological treatment have changed between two decades.
The same coding categories can be applied to material from each period.
However, changes in language, publishing practices and social context may complicate the comparison.
Strength: replicability
A study may be replicated if researchers report:
The selected sources.
The sampling dates.
Inclusion and exclusion criteria.
Coding units.
Category definitions.
Recording procedures.
Reliability checks.
Existing material can sometimes be re-examined by another researcher.
Exact replication may be difficult if online content is edited, deleted or inaccessible.
Limitations of content analysis
Limitation: researcher subjectivity
The researcher makes judgements when:
Selecting the material.
Constructing categories.
Interpreting ambiguous content.
Choosing quotations.
Explaining patterns.
Expectations may influence these decisions.
Two researchers might develop different coding systems for the same material and reach different conclusions.
Limitation: reduction of complex material
Converting rich communication into categories may oversimplify it.
For example, a detailed account of examination stress may be reduced to:
Workload present.
Time pressure present.
Physical symptom absent.
This loses:
The participant’s personal meaning.
The relationship between ideas.
The strength of feeling.
The sequence of events.
Contradictions within the account.
Quantitative coding gains simplicity but may lose depth.
Limitation: category overlap
A single statement may fit several categories.
For example:
“My parents expect high grades, so I spend every evening revising and never have enough time.”
This might involve:
Family expectations.
Workload.
Time pressure.
Reduced leisure.
Researchers need clear rules about whether:
Multiple categories may be recorded.
One category should take priority.
Each statement is coded only once.
Unclear rules reduce reliability.
Limitation: meaning depends on context
Sarcasm, humour and implied meaning can be difficult to code.
For example:
“Brilliant, another three hours of revision.”
The word “brilliant” appears positive, but the statement may be negative or sarcastic.
Automated or word-frequency coding may miss this meaning.
Human coding may preserve context but introduces greater interpretation.
Limitation: source material may be unrepresentative
The selected material may not represent:
All communication on the topic.
The wider population.
Private views.
People who do not publish content.
Other time periods.
Other cultures.
For example, public social-media posts may represent people willing to post publicly rather than everyone experiencing the issue.
Limitation: authenticity cannot always be checked
Researchers may not know whether the material is:
Accurate.
Honest.
Complete.
Produced by the stated author.
Edited by another person.
Created for entertainment.
Written to influence an audience.
Content analysis examines what the material contains.
It does not automatically establish whether the content is truthful.
Limitation: cannot establish cause and effect
A content analysis may identify patterns but does not normally manipulate an independent variable.
For example, a researcher may find that negative descriptions occur more frequently in one type of newspaper than another.
This does not establish why the difference occurred.
Possible explanations include:
Editorial policy.
Intended audience.
Topic selection.
Time period.
Individual writers.
Differences in the events covered.
Content analysis is generally descriptive or correlational rather than experimental.
Limitation: time-consuming coding
Large samples may require:
Selecting material.
Preparing transcripts.
Developing categories.
Training coders.
Coding each item.
Checking reliability.
Resolving disagreements.
Analysing frequencies and themes.
Qualitative coding is especially time-consuming.
Limitation: changes in online material
Online content may be:
Deleted.
Edited.
Reposted.
Removed from public access.
Presented differently by an algorithm.
Difficult to archive consistently.
This can make replication difficult.
Researchers should record clearly when and how the material was accessed.
Ethical issues in content analysis
Public and private material
Material being accessible to a researcher does not necessarily mean that its creator expected it to be used in psychological research.
Researchers should consider:
Whether the content is genuinely public.
Whether access requires group membership.
Whether the person expected privacy.
Whether the topic is sensitive.
Whether quotations could identify the author.
Informed consent
Consent may be straightforward when:
Participants created material specifically for the study.
Interviewees agreed that their transcripts could be analysed.
An organisation granted appropriate access.
Consent is more difficult when researchers use existing material produced without the study in mind.
The ethical decision depends on:
Public availability.
Identifiability.
Sensitivity.
Potential harm.
Reasonable expectations of privacy.
Confidentiality
Researchers should protect individuals by:
Removing names and usernames.
Avoiding identifiable descriptions.
Storing material securely.
Reporting group patterns.
Paraphrasing where direct wording could reveal identity.
Restricting access to private data.
An anonymous username may still be traceable through an online search.
Protection from harm
Publishing a content analysis could harm individuals or groups if it:
Reveals private information.
Reinforces stigma.
Misrepresents a community.
Draws attention to sensitive posts.
Allows authors to be identified.
Presents personal accounts without context.
Researchers should consider the possible impact of both the data collection and the final report.
Sensitive material
Coders may also be exposed to upsetting content.
For example, material may involve:
Violence.
Abuse.
Discrimination.
Serious illness.
Trauma.
Researchers should consider the wellbeing of the people analysing the content as well as those who produced it.
The broader ethical requirements are covered in Ethics in psychological research.
Practical issues in content analysis
Selecting material
The researcher needs clear inclusion criteria.
For example:
Publication dates.
Type of source.
Language.
Intended audience.
Topic.
Minimum length.
Availability of complete material.
Without clear criteria, researchers may select content that supports their expectations.
Sampling material
Analysing only a few convenient items may produce a biased sample.
Researchers should decide whether to use:
Random sampling.
Systematic sampling.
Stratified sampling.
All eligible material within a period.
A carefully defined purposive selection.
The selected method should suit the aim and be reported clearly.
Preparing transcripts
Audio or video material may need to be transcribed.
The researcher must decide whether to record:
Exact words.
Pauses.
Repetition.
Laughter.
Tone.
Overlapping speech.
Non-verbal behaviour.
The amount of detail required depends on the coding system.
Storing material
Researchers may need to store:
Text files.
Images.
Audio.
Video.
Coding sheets.
Identifying information.
Storage must be:
Secure.
Organised.
Accessible only to authorised researchers.
Managed in line with the consent given.
Automated coding
Computer software may help researchers:
Search for words.
Count frequencies.
Organise coded sections.
Store coding decisions.
Compare categories.
However, automated systems may struggle with:
Context.
Sarcasm.
Ambiguity.
Cultural meaning.
Spelling variation.
Implied ideas.
Technology can support coding, but the researcher remains responsible for the validity of the analysis.
Content analysis compared with other methods
Content analysis and questionnaires
Content analysis | Questionnaire |
Examines communication or material | Asks participants written questions |
May use existing content | Usually produces new self-report data |
May involve no direct participant contact | Requires participant responses |
Can analyse open questionnaire answers | Questionnaire is the original data-collection technique |
A researcher may first collect open questionnaire responses and then use content analysis to code them.
Content analysis and interviews
Content analysis | Interview |
Organises and interprets communication | Produces verbal self-report data |
May analyse an interview transcript | Involves direct researcher-participant interaction |
Coding occurs during analysis | Questions are asked during data collection |
Researcher effects may occur during interpretation | Interviewer effects may occur during data collection |
Content analysis and observations
Content analysis | Observation |
Studies the content of communications or media | Studies behaviour |
Uses coding categories | Uses behavioural categories |
May use material created earlier | Usually records behaviour during the study |
Meaning may be central | Observable action is central |
Content analysis and correlations
A content analysis may provide numerical data that are later used in a correlation.
For example, researchers could record:
Frequency of negative words in each diary.
The diary writer’s reported anxiety score.
They could then investigate whether the two measures are related.
The content analysis produces one co-variable, while the correlation analyses the relationship.
Review this distinction in relationships between co-variables.
Selecting content analysis for a research scenario
When content analysis is appropriate
Content analysis may be suitable when the researcher wants to investigate:
How an issue is represented in communication.
Patterns across many documents.
Material from an earlier period.
Existing media without contacting its creators.
Themes within open responses.
Differences between sources.
Changes in language over time.
When content analysis may be less suitable
It may be less suitable when the researcher needs to know:
Whether the content is truthful.
Why a person produced the material.
How the audience understood it.
Whether one variable causes another.
What people currently think but have not expressed publicly.
Another method may be needed.
For example:
An interview could investigate the creator’s intentions.
A questionnaire could examine audience responses.
An experiment could test causal effects.
An observation could record behaviour directly.
Scenario: newspaper coverage
A psychologist wants to investigate whether newspapers present psychological therapy positively or negatively.
Content analysis may be appropriate because:
Newspaper articles provide analysable communication.
Researchers can create categories for positive, negative and neutral presentation.
Different newspapers can be compared.
Existing articles are unaffected by the later analysis.
However:
Tone may be difficult to classify.
Articles may not represent public attitudes.
The coding categories may reflect researcher expectations.
Scenario: interview transcripts
A psychologist has completed unstructured interviews about loneliness and wants to identify common experiences.
Content analysis may be appropriate because:
The transcripts contain detailed qualitative data.
Recurring themes can be coded.
Frequencies can be calculated.
Examples can provide context.
The researcher should:
Define categories clearly.
Preserve participant meaning.
Use a second coder.
Protect confidentiality.
Scenario: television programmes
A researcher wants to investigate the representation of helping behaviour in children’s television.
A suitable procedure might be:
Select a defined sample of programmes.
Treat each scene as a coding unit.
Define helping behaviour precisely.
Record different forms of help.
Use two independent coders.
Compare their coding.
Present frequencies and percentages.
Avoid claiming that the programmes directly cause children’s behaviour.
A method for answering content-analysis design questions
Step 1: Identify the material
State exactly what will be analysed.
Step 2: Link the material to the aim
Explain why the source is appropriate.
Step 3: State the coding unit
Identify whether coding applies to:
Words.
Sentences.
Posts.
Scenes.
Articles.
Complete responses.
Step 4: Construct categories
Categories should be:
Relevant.
Clear.
Operationalised.
Distinct.
Practical to record.
Step 5: Explain the recording method
State whether the researcher will record:
Frequency.
Presence or absence.
Duration.
A rating.
A qualitative theme.
Step 6: Check reliability
Use at least two coders and compare their independent coding.
Step 7: Address validity
Explain how categories reflect the intended concept and preserve sufficient context.
Step 8: Address ethics
Consider:
Consent.
Privacy.
Confidentiality.
Sensitive material.
Identifiability.
Step 9: State an appropriate conclusion
Describe patterns within the selected content without claiming more than the evidence supports.
Writing an effective definition
A strong definition might state:
Content analysis is a research method in which communication, such as written, spoken or visual material, is systematically examined. Researchers create coding categories and record the presence, frequency or nature of relevant content.
This definition identifies:
The material.
The systematic procedure.
Coding.
The type of information recorded.
Writing an effective coding explanation
A strong explanation might state:
Researchers first identify a coding unit, such as a sentence or social-media post. They then construct clear coding categories based on the research aim. Each unit is examined and placed into the appropriate category, allowing frequencies or themes to be recorded. Categories should be operationalised and piloted so that different coders interpret them consistently.
Writing an effective evaluation paragraph
A developed strength might state:
One strength of content analysis is that it can examine communication that already exists. The original authors may not know that their material will later be analysed, so they cannot change it in response to the researcher’s aim. This reduces demand characteristics during the production of the material. However, authors may still have altered their communication for its original audience, so the content cannot automatically be treated as a completely honest account.
A developed limitation might state:
One limitation is that coding may involve subjective interpretation. Researchers decide which categories to create and how ambiguous material should be classified. Their expectations may therefore influence the results. Using operationalised categories, training coders and checking inter-coder reliability can reduce this problem, although agreement does not guarantee that the categories are valid.
Overall evaluation
Content analysis is valuable because it can:
Examine written, spoken and visual communication.
Use material that already exists.
Summarise large amounts of data.
Produce quantitative frequencies.
Retain qualitative themes.
Compare sources or time periods.
Use systematic and replicable coding procedures.
Its main limitations are:
Researcher interpretation.
Loss of context.
Oversimplification through coding.
Difficulty establishing valid categories.
Possible low inter-coder reliability.
Unrepresentative or inauthentic material.
Inability to establish cause and effect.
Ethical concerns involving private or identifiable communication.
The method is strongest when:
The source material is selected systematically.
Coding units are clear.
Categories are operationalised.
Coders are trained.
Reliability is checked.
Context is preserved.
Conclusions remain limited to what the material supports.
Key Words 🔑
Key word | Student-friendly definition | How it may be used in an exam |
Content analysis | A systematic method of examining written, spoken or visual communication. | Define the method or identify it in a research scenario. |
Content | The communication or material being examined. | Identify the source data used in an investigation. |
Coding | The process of classifying parts of the content using defined categories. | Explain how material is converted into analysable data. |
Coding category | A clearly defined type of content that researchers record. | Construct or evaluate a coding system. |
Coding unit | The piece of material to which a code is applied. | Explain whether words, sentences, scenes or posts are being coded. |
Coding frame | The organised categories, definitions and instructions used during coding. | Explain how a standardised coding procedure is created. |
Operationalisation | Defining a category clearly enough for it to be identified and recorded. | Improve vague or subjective categories. |
Frequency | The number of times a category occurs. | Summarise quantitative content-analysis data. |
Theme | A recurring idea or pattern identified in qualitative material. | Explain qualitative content analysis. |
Quantitative data | Numerical information such as frequencies or percentages. | Identify data produced by counting coded content. |
Qualitative data | Descriptive information about meaning, context or experience. | Identify data produced through analysis of themes. |
Primary data | Material collected specifically for the current investigation. | Classify researcher-produced interviews or diaries. |
Secondary data | Material originally produced for another purpose. | Classify archived articles or existing media. |
Inter-coder reliability | The level of agreement between researchers coding the same material. | Explain how coding consistency can be assessed. |
Reliability | The consistency of a procedure or measurement. | Evaluate category definitions and coder agreement. |
Validity | The extent to which the analysis represents the concept it claims to examine. | Evaluate source selection, categories and interpretation. |
Researcher bias | The researcher’s expectations influencing coding or interpretation. | Evaluate subjectivity in content analysis. |
Context | The surrounding material needed to understand the meaning of a coding unit. | Explain why isolated words may be misleading. |
Pilot study | A small trial used to identify problems before the main analysis. | Explain how a coding frame can be improved. |
Replicability | The extent to which another researcher can repeat the procedure. | Evaluate clear source selection and coding instructions. |
Informed consent | Agreement to participate after receiving sufficient information. | Discuss analysis of participant-created material. |
Confidentiality | Protecting identities and private information. | Explain how transcripts or online content should be reported. |
Generalisability | The extent to which findings apply beyond the selected material. | Evaluate the representativeness of the content sample. |
Common Mistakes ⚠️
Mistake: Describing content analysis as simply reading some material.
Why this is incorrect:Content analysis is systematic and uses defined coding categories.
How to improve:Explain how material is selected, coded, recorded and analysed.
Mistake: Saying content analysis can examine only written text.
Why this is incorrect:It may examine written, spoken or visual communication.
How to improve:Refer to appropriate sources such as transcripts, images, programmes or online posts.
Mistake: Confusing content analysis with observation.
Why this is incorrect:Observation records behaviour, whereas content analysis examines communication or media content.
How to improve:Identify what is being analysed and the aim of the study.
Mistake: Using vague coding categories such as “bad behaviour”.
Why this is incorrect:Different coders may interpret the category differently.
How to improve:Define the exact words, actions or ideas that should be recorded.
Mistake: Creating categories that overlap.
Why this is incorrect:The same content may fit several categories, producing inconsistent coding.
How to improve:Make the categories distinct and provide clear coding rules.
Mistake: Listing categories without defining them.
Why this is incomplete:A category title alone may not tell coders when it applies.
How to improve:Provide an operational definition and an example.
Mistake: Saying coding removes all researcher bias.
Why this is incorrect:Researchers choose the material, categories and interpretations.
How to improve:Explain that coding may reduce subjectivity when categories are clear and reliability is checked.
Mistake: Saying two coders automatically produce reliable data.
Why this is incorrect:The coders may disagree substantially.
How to improve:State that their independent decisions should be compared.
Mistake: Claiming high inter-coder reliability proves validity.
Why this is incorrect:Coders may agree consistently while using a category that does not represent the intended concept.
How to improve:Evaluate reliability and validity separately.
Mistake: Ignoring the coding unit.
Why this is incorrect:Coders need to know whether they are classifying words, sentences, scenes or complete items.
How to improve:State and justify a consistent coding unit.
Mistake: Assuming the most frequent category is the most important.
Why this is incorrect:Frequency does not necessarily represent psychological significance.
How to improve:Interpret frequency alongside context and the research aim.
Mistake: Treating media content as a direct measure of reality.
Why this is incorrect:Media show how an issue is presented, not necessarily how common or accurate it is.
How to improve:Limit conclusions to the selected content.
Mistake: Claiming content analysis establishes cause and effect.
Why this is incorrect:The researcher normally records existing patterns without manipulating an IV.
How to improve:Describe associations, differences or representations rather than causal effects.
Mistake: Assuming public online content creates no ethical issues.
Why this is incorrect:Authors may be identifiable, the content may be sensitive and expectations of privacy may vary.
How to improve:Consider consent, privacy, confidentiality and possible harm.
Mistake: Removing statements from their context.
Why this is incorrect:The meaning of words may depend on the surrounding sentence, tone or situation.
How to improve:Use a coding unit large enough to interpret the content accurately.
Mistake: Reporting percentages without showing the relevant total.
Why this is incomplete:The reader cannot judge what the percentage represents.
How to improve:Report both the frequency and the sample size, and show working when asked.
Exam-Style Questions ✍️
Question 1
Which one of the following best describes content analysis?
A. Manipulating an independent variable and measuring a dependent variable
B. Systematically examining communication using coding categories
C. Measuring a relationship between two co-variables
D. Asking every participant the same verbal questions
[1 mark]
Question 2
Define coding as it is used in content analysis.
[2 marks]
Question 3
Explain the purpose of a coding unit in content analysis.
[3 marks]
Question 4
A psychologist wants to investigate how examination stress is represented in online articles.
Write two suitable coding categories and operationalise each one.
[4 marks]
Question 5
A researcher uses the following category when analysing interview transcripts:
“Negative experience”
Explain one problem with this category and suggest how it could be improved.
[4 marks]
Question 6
Two researchers independently code the same 25 social-media posts. They agree on the coding of 20 posts.
a) Calculate the percentage agreement.
[2 marks]
b) Explain what the result suggests about inter-coder reliability.
[2 marks]
Question 7
A psychologist analyses 60 newspaper articles about psychological therapy.
Explain one strength and one limitation of using content analysis for this investigation.
[6 marks]
Question 8
A researcher finds that 70% of selected articles contain references to academic workload.
Explain one conclusion that can be drawn and one conclusion that cannot be drawn from this finding.
[4 marks]
Question 9
A psychologist plans to analyse messages posted within a private online support group.
Explain two ethical issues that should be considered.
[6 marks]
Question 10
Discuss content analysis as a psychological research method.
Refer to coding, reliability and validity in your answer.
[8 marks]



Comments