1. Tyamka
  2. Blog
  3. What does the Stroop test measure, and what can an online score tell you

What does the Stroop test measure, and what can an online score tell you

· 7 min read · Tyamka team

The Stroop test measures how much a conflicting word slows you down. How researchers score it, why one result says little, and what our points mean.

The Stroop test measures one thing well, the cost of conflict. It compares how fast and how accurately you name the ink of a word when the word agrees with the ink and when it does not. The gap between the two is usually called Stroop interference. Researchers use it as one window into how people hold back a strong habit, reading, in favor of a weaker task, naming a color. If you have not tried the task yet, the Stroop test online takes a minute a level. The effect itself is described in the Stroop effect explained. This article is about the number that comes out of it.

The short answer

A Stroop test gives you times and errors for at least two kinds of items.

The main number is a difference between them. Studies with a neutral baseline split it in two. Incongruent minus neutral shows how much the conflicting word slows you down, and this part is called interference. Neutral minus congruent shows how much a matching word speeds you up, called facilitation, and it is usually much smaller. Studies without a neutral baseline simply take incongruent minus congruent. Standardized card tests have their own scoring. Errors on incongruent items are counted too, because a person can keep the times low simply by answering carelessly.

The difference is not meant to measure how fast you are in general. A person who is slow at everything will be slow on congruent and incongruent items alike, and the gap between them can still be small. Subtracting one from the other removes part of the influence of general speed. It is not a pure measure of conflict, because strategy and random noise stay in it.

What researchers use it for

In psychology the Stroop task is one of the standard ways to study cognitive control, the ability to follow a goal when a habit points elsewhere. In the well-known study by Akira Miyake and colleagues in 2000, which looked at how different executive tasks relate to each other, the Stroop task was one of three tasks chosen to represent inhibition of strong responses, next to the antisaccade task, where you look away from a sudden light instead of at it, and the stop-signal task, where you have to stop a keypress you have already started.

It is used in many other ways too. Researchers change the words to study reading, meaning and emotion, change the response from voice to keys to study how answers are prepared, and add cues that switch between reading and naming to study task switching. Colin MacLeod's 1991 review in Psychological Bulletin covers the first fifty years of this work, and the task is still in use in new studies every year.

In clinics there are standardized versions, such as the Stroop Color and Word Test by Charles Golden, published in 1978, with printed pages, fixed timing and tables of norms. A specialist uses them as one part of a wider assessment, together with an interview and other tests. No clinician would draw a conclusion from one Stroop score alone.

How the test is scored

There is no single Stroop test, and the way it is given changes what the number means.

FormatHow it worksWhat you get
CardA whole page of items, read or named aloud, timed with a stopwatchSeconds per page, or items done in a fixed time
Computer, voiceOne item at a time, answer spoken into a microphoneTime per item, errors
Computer, keys or touchOne item at a time, answer with a key or a buttonTime per item, errors

Card versions measure a whole page, so they mix the conflict with reading speed, eye movements and how well you keep your place. Computer versions measure each item, which is cleaner but depends on the device. With key or touch answers the effect is often reported to be smaller than with spoken answers. One likely reason is that you first have to map the color onto a button, though researchers propose other explanations too. None of this is a problem for research, as long as everyone in a study does the same version. It does mean that numbers from different versions cannot be compared.

Why one result says little about a person

The Stroop effect is one of the most reliable findings in psychology at the level of groups. Put a hundred people through it, and nearly all of them will be slower on conflicting words. But the size of the effect for one person is much less stable. In 2018 Craig Hedge, Georgina Powell and Petroc Sumner called this the reliability paradox. Tasks such as Stroop and the flanker task produce strong, stable effects for groups precisely because people differ little in them, and that makes the individual differences they do show hard to measure reliably. When the same people did the Stroop task weeks apart, their interference scores were only moderately consistent.

What moves your score from one day to the next is mostly ordinary.

So a single online score is a snapshot of one attempt. It is not a diagnosis, it says nothing about intelligence, and it is not a way to check a health condition. If you have a real concern about your attention or memory, a doctor is the right person to ask, not a web page.

What our game counts

The Stroop test in Tyamka is a game built on the task, not a lab version. One color word appears on a dark panel, in red, blue, yellow or white ink, and you tap the name of the ink. It keeps score like this.

Your points therefore mix three things, how accurate you were, how fast you were, and how many conflicting words the level had. The game does not split your times into congruent and incongruent, so it does not report an interference score. We think that is honest for a game. A split from a few dozen taps on a phone would look precise without being so.

What the score is good for is comparing you with yourself on the same level. If your best on level 8 goes up over a week, you got better at level 8. That is all we claim, and does brain training work explains why we stop there.

Common questions

Is there a normal Stroop score? Standardized versions have norm tables, which apply only to that exact version and the group it was normed on. Our game has no norms, and scores from different sites are not comparable.

Does a big Stroop effect mean poor attention? No. The effect is a normal result of fluent reading, and nearly every reader shows it. Its size in one attempt changes too much to say anything about a person.

Why is it harder in my second language? Usually it is the other way around. Words in a language you read less fluently tend to interfere less. We cover this with examples in Stroop test examples.

Is the arrow version the same test? It works on a similar idea, a habit against a rule, but with direction instead of color. See the flanker task and the arrow Stroop.

When you are ready, play a few levels of the Stroop test and keep an eye on your own best. Other games are on the all games page.