Back to BlogPsychological Tests

What Does a Psychological Test Score Actually Mean?

Reading a scale score correctly: severity bands, cut-offs, and the "it is not a diagnosis" reality.

June 10, 20266 min read1

The number at the end of a psychological test is one of the most over-read pieces of information in mental health. On its own it carries almost no meaning. What it means depends on three things: which instrument produced it, how that instrument builds a total, and what clinical question it is being asked to answer. This article covers the rules for reading a scale result properly and the mistakes that show up most often.

What a scale score actually is

A psychological scale converts standardized questions into a quantity. How that quantity is presented differs by instrument, and mixing up the three formats is the commonest source of misreading.

  • A total with severity bands. Instruments such as the Beck Anxiety Inventory produce one number falling into a minimal, mild, moderate, or severe range.
  • A multi-dimensional profile. Broad screens like the SCL-90 produce a shape across several symptom dimensions; the informative feature is the contour, not the height.
  • A descriptive typology. On tools such as the Big Five or the Enneagram, high and low are not good and bad; they describe a tendency.

Reading a typology output as if it were a severity band does the most damage, because it turns a description of style into a statement about pathology.

Raw score, band, percentile: three different languages

The same answers can be reported in several currencies, and reports rarely say which one is on screen.

Score typeWhat it tells youWhere it misleads
Raw totalThe bare sum of item responsesNot comparable across instruments; 30 is low on one scale and alarming on another
Severity bandThe clinical level the raw total corresponds toA score at the edge of a band is clinically almost identical to one just below it
Percentile rankWhere the person sits within a reference sampleMeaningless if the reference sample does not resemble the person tested
Subscale profileWhich domain the symptom load is concentrated inOne elevated subscale gets read as if it described the whole picture

Cut-offs are decisions, not properties

A cut-off is the threshold above which a result is treated as worth a closer look; on the GAD-7, a total of 10 or above is a widely used example. The same instrument appears with different thresholds in different settings, and that is not sloppiness.

Every threshold trades two properties against each other. Lowering it catches more of the people who need follow-up, at the cost of flagging people who do not; raising it does the reverse. In a large community screen, missing people is the expensive error, so the threshold goes down. In a service with limited assessment capacity, it goes up. The cut-off belongs to the purpose, not to the scale — crossing it is an invitation to look, not a verdict, and staying under it certifies nothing.

A case: the same total, two clinical routes

Two adults complete the same anxiety measure in the same week and arrive at an identical total.

The first accumulates most of it on bodily items — racing heart, breathlessness, trembling — and describes discrete episodes that arrive abruptly and subside within twenty minutes. The sensible next steps are a closer look in the direction of panic and the exclusion of medical contributors.

The second accumulates the same total on cognitive items: persistent worry, an inability to switch it off, anticipating the worst, running for two years and tied to no single topic. A generalized-anxiety framing fits better, and a brief screen for low mood is worth adding.

Identical numbers, entirely different clinical paths.

Reading change, not just level

The most valuable use of a scale is the second measurement, not the single snapshot — though not every difference is real change. Measurement carries error, and small movements can come from sleep, time of day, or reading the items in a different mood. Hold three things constant before interpreting a difference: the same instrument, the same reporting window, and ideally a comparable time of day. One or two points is noise; a shift that crosses a band and matches clinical observation is signal.

The rule that matters most

Screening instruments quantify symptoms. They do not diagnose. A diagnosis weighs duration, functional impact, developmental and medical history, and the exclusion of alternative explanations, and only a licensed professional can carry it out. A scale result is where that process starts.

Five checks before you read a result out loud

  1. Read the number together with its band. A bare figure carries no scale of reference.
  2. Check the direction of the instrument. On self-esteem and distress-tolerance measures, a high score is the favorable one.
  3. Check the reporting window. "In the past week" and "in general" are different questions and produce different totals.
  4. Look at the response pattern. An identical option marked down the whole form, or a half-finished questionnaire, can invalidate the total.
  5. Look at the clinical picture rather than a single instrument. Scales complement each other; none replaces another.

Frequently asked questions

I scored high. Does that mean I am ill?

No. A high score means symptoms were reported as intense during the window the instrument asked about. Whether that reflects a disorder, a period of strain, or a medical cause is what clinical assessment is for.

Should a result change over time?

Yes, and it usually does. Screening scales ask about a bounded period, so as circumstances shift the total shifts with them. That sensitivity is what makes them useful for tracking treatment response.

Two scales on the same topic gave me different answers. Which is right?

Most often both are. If one weights bodily symptoms and the other weights worry cognitions, divergence is expected. The gap is a finding, not a contradiction.

Is a test I took online valid?

The instrument may be sound, but conditions of administration and the person interpreting the output determine what the result is worth. A number obtained without professional interpretation is, at best, a reason to book an appointment.

For a worked comparison of two instruments that measure the same construct differently, see BAI vs. GAD-7. For how scores are collected and reported in practice, see the digital test administration guide. Every instrument on the platform is listed in the scales catalog.

This article is for information only. The scales and tests mentioned here are screening and assessment instruments; on their own they do not establish a clinical diagnosis. Administration and interpretation belong to licensed professionals trained in the relevant instrument. Item texts, stimulus cards, and norm tables of copyrighted tests are never published on this page.