16 Types, 64 Types and Beyond: How Finely Can Personality Be Measured?
Going from 16 types to 64 gives each result more to say, but a category still can’t tell someone near its edge from someone at its center. A continuous score shows how far you lean, and that changes how you have to read it.

When the description gets half of you right
Say your four-letter result described someone who likes everything decided in advance. That’s you at work, with the shared calendar color-coded and every project plan carrying dates. Then you go on holiday, book a hotel for the first night and leave the rest of the week open.
A 64-type version of the same idea looks built for that gap. With four times as many categories there’s room for a planner who leaves holidays unplanned, and the longer description you get might name exactly that. Each extra split promises to catch something the broader label missed, and sometimes it does.
Where the extra categories come from
Sixty-four is two to the sixth power, the count you’d get from six either-or splits where the familiar system has four. A test can get there in two quite different ways, and the label alone won’t tell you which one produced yours.
One way is to cut the same answers more finely. Take two of the four letter pairs and divide each into four bands instead of two halves, so that leaning slightly toward one side counts separately from leaning hard. That makes 64 categories without a single extra question. Whatever the new name adds, it comes from answers that were already counted.
The other way is to ask about something new. A fifth and sixth split, each with its own questions, bring in information the 16 types never had. Whether those questions capture what the new label says they do then depends on that particular test’s evidence.
Illustration
Three ways a result can be reported
16 categories
Four either-or splits, one letter for each. Easy to remember and to compare with a friend’s.
64 categories
Six splits, or four with some cut finer. More names for more differences, though everyone inside a category reads the same one.
Continuous scores
A position from 0 to 100 on each dimension. The distance from the middle shows, and you judge how much a few points matter.
What a finer label still drops
However many boxes a system sorts answers into, every box has edges, and the label doesn’t say how close to one your answers came. Two people who share one of 64 labels read the same description, even if one only just crossed into that category and the other is near its center. In a smaller box, two people sharing a label can’t be as far apart. Finer categories also add edges, though, and every new edge is another place for a set of answers to land right beside one.
Whether a repeat of the test would put you in the same category again has a name in the Standards for Educational and Psychological Testing: decision consistency. The Standards tie it to two things: how precise the scores are and where the cut-off falls.
A number instead of a name
A label is easy to carry around. Four letters fit in a dating profile or an Instagram bio, and a list of scores doesn’t. A continuous score gives up that convenience to report a position: 78, say, on a 0 to 100 scale. Someone at 51 and someone at 89 stop sharing a description.
The same Standards describe the trade-off. Use too few score points and information is discarded; use too many and people may try to interpret differences that are small relative to the measurement error in the scores. Splitting a scale into four bands instead of two adds points, and a 0 to 100 score adds far more.
We think the number is worth having, as long as you read 78 as a region well toward one end of the scale and give a gap of two or three points little weight. Whether a scale measures the trait it’s named for is a validity question, and the format of the result can’t answer it.
Fourteen scales instead of one code
Traitium gives you a score from 0 to 100 on each of 14 dimensions, from 50 questions. Two of them bear on the holiday planner at the top of this page: Structure Preference and Planning Intensity. Both describe a middle range where people fix some things and leave the rest open. The Structure Preference page’s example is “a booked hotel but a free afternoon,” and it adds that which way people in that range lean can depend on how much is at stake. Someone with a color-coded work calendar and an open week away may sit nearer the middle than a single letter suggests.
Every dimension has a page like that, covering both ends and the middle, where you can check what 78 on Social Confidence describes before reading anything into it. Each score comes from your answers alone, and the same answers give the same scores in every language, whether you take the test in English, Traditional Chinese or Japanese.
What the extra scales add is coverage: Risk Tolerance and Independence describe tendencies the four letters don’t cover. Related tendencies such as Social Confidence and Leadership can rise and fall together, so 14 scores don’t add up to 14 unrelated findings.
Sources and notes
- [1]AERA, APA & NCME — Standards for Educational and Psychological Testing (2014)
Chapter 2 defines decision consistency and links it to score precision and cut-score location (p. 40). Chapter 5 notes that too few score points discard information, while too many invite interpreting differences that are small relative to the measurement error in the scores (p. 95).

