Skip to main content

How to Choose a Personality Test You Can Trust

A personality test can win your trust the moment you recognize yourself in the result, and that feeling says very little about the test. Look instead at its questions, its retest evidence, the meaning of its numbers and the provider’s documentation.

By Traitium5 min read
A clear glass lens suspended above layered sage and brass topographic terrain.

When a result sounds like you

In a classroom demonstration published in 1949, the psychologist Bertram Forer gave his students a personality test, then handed each of them a sketch presented as an individual reading of their answers. Every student got the same sketch, and the class rated it as a good description of themselves.

Forer called the mistake the fallacy of personal validation, and it’s now usually known as the Barnum effect. Agreeing with a description can’t validate the test behind it, because statements general enough to fit a wide range of people will fit you as well. A results page might tell you that you hold yourself to high standards and can be hard on yourself when you fall short. Picture the same line on your manager’s results page, or your brother’s, and it may well fit them too.

Start with the questions

Somewhere in the middle of a quiz you get to an item like “I love big parties and meeting new people.” Maybe you like meeting people one at a time and dread parties. There’s one answer for both halves of the sentence, so whatever you pick, your score records something you didn’t quite say.

Other items have an answer that sounds better than the rest. How you rate yourself on “I’m a good listener” depends partly on how you’d like to come across. The Standards for Educational and Psychological Testing, published by AERA, APA and NCME, count a response bias, such as playing down your anxiety on an anxiety test, among their examples of construct-irrelevant variance: scores shifted by something the test isn’t meant to measure.

Forty items that keep rewording the same few situations add minutes to a test and may steady the score, but they don’t cover any more of the trait. A test can also leave out parts of what it claims to measure, which the same document calls construct underrepresentation. Its example is a test meant to measure anxiety that asks only about physical reactions and skips the emotional, cognitive and situational components. So check whether the items for one trait cover different facets of it: for sociability, how much you enjoy a crowded room as well as how drained you feel after a long day with people.

Reliable and valid are separate claims

Reliability, as the Standards use the term, is how far scores agree across repetitions of a testing procedure: another sitting, a comparable set of items, another person doing the scoring. A retest study estimates the first kind by having a group of people take the test twice and comparing the two sets of scores. That describes the test across many people, and your own scores can still move between sittings on a test that does well by this measure.

A provider that has run a study like this can show it. The Myers & Briggs Foundation, for instance, reports that 1,721 adults took the MBTI Step I twice, 6 to 15 weeks apart, with test-retest coefficients of 0.81 to 0.86. You don’t need to know what a good coefficient looks like to see why that claim is checkable: it names the instrument, the sample, the gap between sittings and the statistic.

Those figures are about consistency and nothing else. Validity asks whether the evidence supports a particular use of the scores. The Standards are strict about the wording: validity belongs to interpretations of scores for specified uses, so the unqualified phrase “the validity of the test” is, in their terms, incorrect. When a landing page says “scientifically validated,” ask what for. A score you use to start a conversation about how you plan your week and a score someone uses to decide who gets a job aren’t held to the same bar, and the document adds that the need for precision increases as the consequences of decisions grow.

Percentile or position on a scale

A score on a continuous scale can keep what a category throws away: the difference between leaning slightly toward one end and leaning hard. Whether a few points mean anything depends on how precisely the test measures, and before any of that you need to know what kind of number you’re looking at. “Introversion: 72” could be a percentile, meaning your answers came out more introverted than those of 72 percent of some comparison group. It could also be a position on the test’s own 0–100 scale, worked out from your answers alone, with no comparison group at all.

For a score on the test’s own scale, you need to know which end is which and what a score near the middle looks like in daily life. For a percentile, find out who the comparison group was and when they were tested. The Standards expect publishers that report norms to specify the population they sampled and the dates of testing.

What a provider should put in writing

The Standards ask test publishers to document a test’s rationale, its recommended uses, the evidence for those uses and the information people need to interpret the scores, and to warn against misuses they can anticipate. That’s addressed to publishers, but you can read a provider’s About or methodology page against that list before you answer a single question.

Illustration

Six things to look for before you start

  • One thing per question

    Each item asks about a single behavior or situation you can picture.

  • No flattering answer

    None of the options sounds obviously better than the others.

  • Range within each trait

    Items for the same trait cover different situations and behaviors.

  • Retest evidence

    Who took the test twice, how far apart, and what was calculated.

  • A stated use

    What the scores are meant for, what they shouldn’t be used for, and the evidence behind those uses.

  • Readable scores

    Whether a number is a percentile or a scale position, and which end is which.

A reading guide for any provider’s site, Traitium’s included.

Reading Traitium the same way

Traitium asks 50 questions, each a statement you rate on a seven-point scale, and gives you a score from 0 to 100 on each of 14 continuous dimensions. We chose 50 questions and continuous scores; neither is evidence of validity. The scores are positions on each dimension’s scale: a 70 on Planning Intensity means your answers lean toward the methodical end, and it isn’t compared with anyone else’s score. Every dimension has its own page describing both ends and the middle range, so you can see how a scale is defined before you take the test, and each insight in your profile carries a “Why we identified this” note that names the scores behind it.

The same encoded answers produce the same scores in English, Traditional Chinese and Japanese, which makes the scoring reproducible. It says nothing about whether you’d answer the same way next month, or whether an item means the same thing in all three languages, which is part of the validity question.

Our stated use is personal reflection. The methodology page describes Traitium as a self-discovery tool and says plainly what it isn’t: a clinical instrument or a diagnosis.

Sources and notes

  1. [1]AERA, APA & NCME — Standards for Educational and Psychological Testing (2014)

    Chapter 1 (validity, construct underrepresentation, construct-irrelevant variance) and chapter 2 (reliability/precision) for the definitions used here; standard 5.9 for what norming reports should specify; standards 7.1 and 7.2 for what test documentation should cover.

  2. [2]Myers & Briggs Foundation — Reliability and Validity

    Defines reliability as results staying consistent over time and reports the MBTI Step I test-retest figures quoted above.

  3. [3]B. R. Forer — The Fallacy of Personal Validation (1949)

    The original classroom demonstration. Journal of Abnormal and Social Psychology, 44(1), 118–123.

See where your answers land

Traitium’s free test has 50 questions and gives you a score from 0 to 100 on each of 14 dimensions.

Start the free test

All guides