Showing posts with label Critical Appraisal. Show all posts
Showing posts with label Critical Appraisal. Show all posts

Thursday, March 12, 2026

The cognitive effects of dark chocolate: critically appraising a randomised trial

 

Disclaimer: this blog post has been created as an example of critical appraisal and doesn’t constitute medical advice. For expert advice, please speak to your health professional!

With Easter approaching, you may be thinking of chocolate eggs as an Easter treat. But did you know, some studies have explored the health benefits of chocolate eaten in small amounts regularly? In the library team, we encourage and provide training for critical appraisal of research studies. So, we decided to present a critical appraisal of a study about the cognitive effects of eating chocolate – how reliable is the research? Read on to find out…

Photo by Elena Leya on Unsplash

What is critical appraisal?

Critical appraisal is evaluating information to determine its trustworthiness and its relevance to a particular context. There are several published checklists that provide guidance for this evaluation, such as those produced by CASP, JBI and CEBM. Each checklist is usually tailored to a particular type of research study. So first we want to identify – what type of study is the article that we’re looking at?

The article

We’ll be looking at Sub-Chronic Consumption of Dark Chocolate Enhances Cognitive Function and Releases Nerve Growth Factors: A Parallel-Group Randomized Trial by Sumiyoshi et al., published in 2019. It’s a randomised trial so, if taking the CASP checklists as an example, the closest fit will be the checklist for randomised controlled trials.

The article explores the cognitive effects of regularly eating dark chocolate in the medium term, in comparison to white chocolate. If you’re a white chocolate fan, the bad news is that white chocolate contains fewer of the nutrients that have previously been found to have beneficial health effects. The nutrients specifically mentioned in this article are flavonoids – namely epicatechin and catechin - and methylxanthines - namely theobromine and caffeine - with regard to dark chocolate.

The researchers recruited participants to be divided into two groups: one group eating a set daily amount of dark chocolate and the other group eating a set daily amount of white chocolate, over 30 days. They were instructed not to eat any other chocolate during this time and not to exceed more than 3 cups of caffeinated drink per day. At the start of the trial, at the end of the trial and 3 weeks after the trial finished, the key measurements were taken. These included two cognitive tests: the Stroop Colour Word Test (modified), which uses names of colours printed in different coloured ink to test reaction to cognitive interference (disruption to usual thinking/processing), and the Digital Cancellation Test, which involves finding numbers amongst a sheet of digits, in order to measure attention, processing speed and executive functioning. The researchers also measured some physical attributes such as weight, BMI, heart rate and various chemicals in the blood.

Image by Willfried Wende from Pixabay


CASP Section A: Is the basic study design valid for a randomised controlled trial (RCT)?

This study isn’t claiming to be an RCT, rather a randomised trial. This is probably because the factors aside from chocolate consumption (like a participant having extra cake, biscuits etc.!) weren’t absolutely controlled, they were only given limited instructions to follow at home. The lack of control over these other factors could allow them to influence the findings – something to bear in mind.

On the other hand, participants were assigned to white chocolate or dark chocolate randomly using a computer generator, which is good practice. Randomisation is important because it minimises bias. For example, if the researchers were to allocate participants non-randomly, they might assign healthier-looking participants to dark chocolate for better likelihood of outcomes.

Another aspect to mention in terms of study design is the sample size. Twenty students were recruited for the study and two of them dropped out. This is a very small sample which means the reliability of results hasn’t been tested on enough people to be really trustworthy. It suggests that the results might not reflect the wider population. Seeing as the whole sample was made up of healthy undergraduate students (aged 20-31 years), this is even more the case. It means we don’t know for sure if the findings could be applied to other types of people with different ages or health conditions.

Image by mcmurryjulie from Pixabay


CASP Section B: Was the study methodologically sound?

Can you taste the difference between dark chocolate and white chocolate? Probably very easily yes, because white chocolate is much sweeter. So, in this study it wouldn’t have been possible to ‘mask’ or ‘blind’ the participants, to prevent them from knowing which chocolate intervention they received. ‘Blinding’ or ‘masking’ is normally carried out in RCTs when possible, because it minimises bias caused by participants knowing which type of treatment they’ve been given. For example, if a person knows that they are only receiving a placebo (‘dummy drug’) rather than a real drug, their mind and body could actually respond differently than if they believed they were receiving the real drug (known as the placebo effect).

That said, the participants weren’t told the exact objective of the study until after the course of chocolate was complete. In a way, this acted as an alternative form of masking, limiting the influence of participants’ knowledge on outcomes – unless of course they guessed what the experiment might be about! The article also says that the researchers were unaware of which participants were consuming which type of chocolate. This limits the potential bias from researchers’ knowledge of groupings.

Next, we want to know – were the characteristics of the participants allocated to dark chocolate similar to the participants allocated to white chocolate, at the start of the study? If one group was particularly sporty and fit, but the other wasn’t, that wouldn’t be a fair test!  The article tells us that there were no significant differences, so we’re assuming they distributed genders etc. evenly between groups. From the supplementary tables we can see that the weight, BMI, heart rate, blood composition etc. of the 2 groups didn’t vary significantly. This all means there’s low likelihood of any of these factors affecting the results. In that case, if the results show differences between groups, it’s more likely down to the type of chocolate.

Similarly, we want to know if the groups were treated any differently during the trial, besides being given different types of chocolate. Unless the participants had some sneaky extra chocolate or coffee, we know that they followed instructions to limit caffeine intake, refrain from other chocolate and avoid intense exercise. Again, this would prevent other factors influencing results. In other aspects, such as other food or drink intake while at home, we don’t know whether this varied between groups during the trial.

 

Image by Memed_Nurrohmad from Pixabay


CASP Section C: What are the results?

According to the scores from the cognitive function tests, there was some improvement in outcomes for the dark chocolate group, but results were mixed. For example, in the Stroop Colour Word Test, the dark chocolate group showed improved performance after starting the course of chocolate; but the white chocolate group also improved on the colour part of the test. In the third part of the Digital Cancellation Test, the dark chocolate group showed significant improved total performance, but not in the first and second parts of the test (neither did the white chocolate group). Basically, in this study, there are signs of a link between dark chocolate and improvement in cognitive tests, but the pattern isn’t really consistent or strong enough to be completely certain.

The measures of theobromine (a nutrient which previous studies have linked with cognitive performance) from blood plasma show it rose significantly in the dark chocolate group but not in the white chocolate group. Caffeine and other chemicals didn’t show such a significant increase. So, if dark chocolate does improve cognitive function, theobromine is probably at least one of the reasons behind it. Another contributing element that the authors describe is Nerve Growth Factor, which was also elevated in the dark chocolate group compared to the white chocolate group.

The researchers calculated the statistical significance of their results, to show whether the findings were likely to be due to chance (p-values). Overall, some of the measurements couldn’t be said to be statistically significant, but others were, including the increased number of correct answers for the Stroop test in the dark chocolate group.

Image by Mohamed Hassan from Pixabay
 

CASP Section D: Will the results help locally?

So, overall, what should we make of this study? If you are a healthy student (or young adult), aged 20-31, eating around 24g of dark chocolate with at least 70% cocoa every day (cocoa is the key ingredient, so higher percentages may make a difference!), it’s quite likely that your cognitive functioning would be higher than a similar person eating white chocolate (0% cocoa) daily.

However, it’s not so certain whether this applies to the same extent if you are of a different age etc. or if you consume a significantly lower or higher amount of the chocolate per day. There could be some anomalies in cognitive performance even if you mirrored the participants in the study. Meanwhile, you might not be particularly bothered about cognitive performance. You might be more interested in other health factors related to chocolate consumption, such as cardiovascular function, diabetes etc. which aren’t covered in this study. Some other studies have looked at these sort of aspects, both positive and negative (e.g. Morze et al. (2020), Amoah et al. (2022) and Yuan et al. (2017), to name just a few). It really is a case of weighing up the strengths and weaknesses of this study and other studies then coming to a critical judgement depending on your circumstances (and possibly professional advice - health professionals rather than confectionery professionals, that is!). 

Whatever your own conclusions, we hope you can enjoy the spring/Easter season!

(No chocolate was harmed in the making of this blog post…probably.)

We hope you found this article interesting. If you’d like to find further evidence about the health effects of chocolate, you could search on the NHS Knowledge and Library Hub using keywords such as ((chocolate OR cacao OR cocoa) AND (health OR wellbeing OR benefit* OR risk* OR effect* OR outcome*)).

Your library team are always happy to support with locating articles.


Thursday, November 20, 2025

CASP Checklists in 10 Minutes

You have just read an article claiming a "breakthrough" treatment for a condition you manage regularly, but you are thinking: "How do I know if this research is actually any good?"

This is where CASP comes in. 
CASP (Critical Appraisal Skills Programme) checklists are a series of checklists involving prompt questions to help you evaluate research studies. They are designed to help systematically assess the trustworthiness, value, and relevance of published research studies.

All CASP checklists are structured around three main questions to guide the appraisal process: 
  1. Are the results of the study valid? (Assesses methodological rigour and bias).
  2. What are the results? (Examines the reported outcomes and their clinical importance).
  3. Will the results help locally (in my setting)? (Evaluates the relevance and applicability of the findings to a specific context).

CASP provides free checklists for the most common types of research you'll encounter:

  1. Systematic Reviews with Meta-Analysis of Observational Studies
  2. Systematic Reviews with Meta-Analysis of RCTs
  3. Randomised Controlled Trial / RCT 
  4. Systematic Review
  5. Qualitative Studies 
  6. Cohort Study
  7. Diagnostic Study
  8. Case Control Study
  9. Economic Evaluation
  10. Clinical Prediction Rule
  11. Cross-Sectional Studies

Download PDF or Word versions at: https://casp-uk.net/casp-tools-checklists/

How to Use CASP Checklists in 10 Minutes

Let's break down the systematic review checklist as an example (most commonly used).


The Three Sections

As already mentioned, every CASP checklist has three parts:

Section A: Are the results valid? (Screening questions)
Section B: What are the results? (Detailed questions)
Section C: Will the results help locally? (Applicability)

Section A: Screening Questions (2 minutes)

These are your "deal breakers." If the study fails here, you can stop—it's not worth continuing.

Question 1: Did the review address a clearly focused question?

Look for PICO:

  • Population: Who was studied?
  • Intervention: What was done?
  • Comparison: Compared to what?
  • Outcome: What did they measure?

Example of a good focused question: "In adults with Type 2 diabetes (P), does metformin (I) compared to placebo (C) reduce cardiovascular events (O)?"

Example of a poor question: "Does medication help diabetes?" (Too vague!)

If YES → Continue. If NO → Stop here, study is too broad/unclear


Question 2: Did the authors look for the right type of papers?

For a systematic review about treatment effectiveness, you'd want to see:

  • ✅ Randomised controlled trials (RCTs)
  • ✅ High-quality studies
  • ✅ Relevant to the question

If they are including case reports or opinion pieces for a treatment question, that's a red flag.

If YES → Continue to Section B. If NO → Major concerns about reliability

Time check: 2 minutes spent. Should you continue? If both screening questions = YES, proceed.


Section B: Detailed Assessment (5 minutes)

Now you're diving deeper into the quality of the research.

Question 3: Do you think all the important, relevant studies were included?

Look for:

  • ✅ Comprehensive search strategy (multiple databases)
  • ✅ Clear inclusion/exclusion criteria
  • ✅ Hand searching of reference lists
  • ✅ Attempts to find unpublished studies
  • ❌ Only searched one database = incomplete
  • ❌ Only English language papers = potential bias

Question 4: Did the review's authors do enough to assess the quality of included studies?

Look for:

  • ✅ Used validated quality assessment tools (like CASP!!)
  • ✅ At least two reviewers assessed each study independently
  • ✅ Quality scores reported
  • ❌ No quality assessment = you don't know if they included junk studies

Question 5: If the results have been combined, was it reasonable to do so?

Check:

  • ✅ Studies were similar enough to combine (similar populations, interventions, outcomes)
  • ✅ Statistical heterogeneity assessed
  • ❌ Combined apples and oranges (e.g., different age groups, different interventions)

Question 6: What are the overall results of the review?

Now you're getting to the findings:

  • What is the main result?
  • Is there a clear effect size?
  • Are confidence intervals reported?
  • How certain are the results?

Example: "Intervention reduced mortality by 20% (95% CI: 10-30%)" = clear, useful result

Question 7: How precise are the results?

Look at confidence intervals:

  • Narrow = precise, confident
  • Wide = uncertain, less reliable

Example:

  • Precise: Risk reduction 20% (CI: 18-22%) = we're pretty sure it's around 20%
  • Imprecise: Risk reduction 20% (CI: 2-38%) = could be anywhere from barely effective to very effective

Time check: 7 minutes total. Almost done!


Section C: Will the results help locally? (3 minutes)

This is where you decide: "Should I change my practice?"

Question 8: Can the results be applied to the local population?

Consider:

  • Is your patient population similar to the study population?
  • Are there important differences (age, comorbidities, setting)?
  • Is the intervention feasible in your setting?

Example: Study in a community settings with limited resources might not apply to an acute hospital setting with intensive monitoring. 

Question 9: Were all important outcomes considered?

Check:

  • Did they measure what matters to patients?
  • Did they only report positive outcomes (cherry-picking)?
  • What about adverse effects, quality of life, cost?

Question 10: Are the benefits worth the harms and costs?

The final question:

  • What is the balance of benefits vs risks?
  • Is it cost-effective?
  • What do patients value?
  • Are there alternative interventions?

Time check: 10 minutes total. Done!


CASP Checklists for Different Study Types


Randomised Controlled Trial (RCT) Checklist

Use when: Evaluating treatment effectiveness studies

Key screening questions:
  1. Did the trial address a clearly focused issue?
  2. Was the assignment of patients to treatments randomised?
  3. Were all patients who entered the trial properly accounted for at its conclusion?
Red flags:
  • ❌ No randomisation
  • ❌ High dropout rates (>20%)
  • ❌ No intention-to-treat analysis
  • ❌ Unblinded when blinding was possible
Time: 10 minutes


Cohort Study Checklist

Use when: Looking at prognosis, outcomes, or risk factors

Key screening questions:
  1. Did the study address a clearly focused issue?
  2. Was the cohort recruited in an acceptable way?
  3. Was the exposure accurately measured to minimise bias?
Red flags:
  • ❌ Selected cohort (not representative)
  • ❌ Short follow-up period
  • ❌ High loss to follow-up
  • ❌ No adjustment for confounding factors
Time: 10 minutes


Qualitative Research Checklist

Use when: Understanding patient experiences, perspectives, or complex phenomena

Key screening questions:
  1. Was there a clear statement of the aims?
  2. Is a qualitative methodology appropriate?
Red flags:
  • ❌ Quantitative question disguised as qualitative
  • ❌ No description of methods
  • ❌ Researcher bias not considered
  • ❌ No participant quotes/data
Time: 10 minutes


Case Control Study Checklist

Use when: Investigating causes of disease or rare outcomes

Key screening questions:
  1. Did the study address a clearly focused issue?
  2. Did the authors use an appropriate method to answer their question?
  3. Were the cases recruited in an acceptable way?
Red flags:
  • ❌ Cases and controls from different populations
  • ❌ Recall bias not addressed
  • ❌ No matching or adjustment for confounders
  • ❌ Small sample size for rare exposure
Time: 10 minutes

Further help:


For further help on using checklists contact the Clinical Librarians at mtw-tr.clinical.librarians@nhs.net  or sign up for our Critical Appraisal training course on MTW Learning or iLearn (KMMH)