The trait half of this assessment is ordinary: fifty public-domain statements that measure the Big Five, scored the standard way. The unusual half is five more statements that ask how you are answering, and a correction built on what they reveal. Most quizzes take you at your word. This one does not.
Here is the whole machine, in order, without the jargon.
Ask someone to rate their own honesty and you get a number. The trouble is that the number is produced by the same mind that benefits from a flattering answer.
This is not lying. Almost nobody sits down to a questionnaire planning to deceive anyone. It is quieter than that: the answer that comes to mind first is the one that fits the story we already tell about ourselves, and that story was built to be told. The gap opens up without anyone deciding to open it.
In the Negev Desert there is a bird called the Arabian babbler. Babblers feed each other, groom each other, and take turns standing sentinel on an exposed branch, watching for hawks while the rest of the group forages in safety.
What the biologist Amotz Zahavi noticed, watching them for decades, is that they do not merely take turns. They compete for the turn. A dominant babbler will shove a rival off the sentinel post and force food on a subordinate that did not ask for it. The generosity is a contest, and the prize is rank.
That is the logic the whole assessment is built on. A signal is believable in proportion to what it costs to send. Anyone can say they are generous. Being generous when it is expensive, unwitnessed, and uncredited is a different claim entirely, because the cheap version of it is not available.
Where humans differ from the bird. The babbler's display is honest because it is physically costly. Ours has an extra layer: we hide the competitive motive from ourselves. A motive you have never consciously noticed is one you cannot accidentally leak through a pause or a glance, which makes the performance considerably more convincing. This idea, and the babbler with it, comes from The Elephant in the Brain by Kevin Simler and Robin Hanson.
You rate each statement from 1 to 5. They look like one questionnaire. They are really two, and the second exists to check the first.
These are deliberately ordinary. They are the IPIP Big-Five Factor Markers, fifty plain statements like "I am always prepared" and "I have a vivid imagination", ten for each trait. They are in the public domain, they have been used for decades, and they are the exact items behind the million-response dataset this project's model is trained on. Using anything bespoke here would mean the model had nothing to learn from. Half of them are reverse-worded, so agreeing with everything does not produce a high score on anything.
This is where the project's own argument lives. They ask about the machinery rather than the person: whether you find it easy to admit a mistake, whether you are genuinely humble or only performing it, whether you behave the same with friends as at work, whether your own motives are ever opaque to you, and whether you can recall a specific time you acted selfishly.
They are deliberately awkward. Someone answering carefully produces a mixed, slightly uncomfortable pattern here. Someone reaching for the good answer every time produces a suspiciously tidy one, and a tidy pattern sitting next to five confident trait scores is exactly the contradiction the correction is looking for.
Each trait averages its ten items and lands on a scale from 0 to 100. There is nothing inventive about this half, and that is deliberate. Using the standard instrument is what lets your scores mean the same thing they mean elsewhere, and what lets a model trained on a million other people's answers say anything useful about yours.
Curiosity, imagination, and appetite for the unfamiliar.
"I have a vivid imagination." "I am full of ideas."
Order, follow-through, and attention to detail.
"I am always prepared." "I shirk my duties." (reversed)
Sociability, assertiveness, and energy around people.
"I start conversations." "I keep in the background." (reversed)
Warmth, empathy, and concern for other people.
"I sympathise with others' feelings." "I insult people." (reversed)
Proneness to worry, tension, and shifting moods.
"I worry about things." "I am relaxed most of the time." (reversed)
Half the items run backwards. Twenty-two of the fifty are worded so that agreeing means a lower score, and they are scattered through the questionnaire rather than grouped. Someone who agrees with everything therefore lands near the middle on every trait rather than at the top, which is one of the reasons this instrument has lasted.
Once the five raw scores exist, the system compares them against the cross-checks and asks a single question: does this set of answers hang together, or is it too good to be internally consistent?
Suppose someone scores very high on agreeableness, then says they find it hard to admit a mistake and cannot recall ever acting selfishly. Each answer is defensible on its own. Together they describe a person who has never once been at odds with themselves, which is not a kind of person that exists. That combination raises the adjustment.
The correction is not applied evenly. It lands hardest on traits that stand out from the rest of your profile, because an unusually high score is where a flattering answer would have shown up. It is also weighted per trait, and neuroticism moves the other way: agreeableness and conscientiousness are the ones people talk up, while neuroticism is the one they play down, so the correction raises it rather than lowering it.
This is an opinionated step, and worth saying plainly. Subtracting an estimated amount of self-flattery from someone's self-report is this project's central bet, not a settled method in psychology. Both numbers are kept and both are shown to you, so you can disagree with the correction and read the originals.
From the moment you press submit, your answers pass through six stages. Each one hands its output to the next, and every intermediate result is kept.
Your fifty answers are fed through a small network trained on just over a million other people's answers to the same fifty statements. It compresses them to sixteen numbers, and the nearest of six patterns found in that data becomes your result.
The map is not a picture of anything you would recognise. It is a compact space where profiles that answered similarly end up near each other, which lets the system express "this profile resembles that one" as a distance rather than a rule. The six patterns were not chosen in advance: they are clusters found in the training data, and they are named after whichever trait sets each one apart. The confidence figure on your results is how much closer you sit to your nearest pattern than to the next one along.
A note on the six patterns. They are reference points for describing a profile, not boxes people belong in. Two people matched to the same pattern can differ a great deal, and the five trait scores underneath carry far more information than the name on top of them.
For the curious: the site is a plain static front end talking to a small Python API. Nothing runs in your browser except the questionnaire itself and the code that draws your results, and nothing your answers touch leaves that server. There is no external model API in this system at all.
A method this opinionated should be equally clear about where it stops.
The trait half uses a long-established public instrument, and the model behind the patterns is trained on a public dataset of over a million responses. The correction on top of it has not been checked against any outside measure of behaviour.
Nothing here diagnoses anything. It should not inform decisions about hiring, admissions, treatment, or anyone's standing.
Correcting a self-report for estimated self-flattery is this project's argument, not established practice. Your uncorrected scores are kept and shown alongside.
Fifty-five statements on one afternoon capture a mood as much as a disposition. Read the result as a prompt for reflection, not a verdict.
On the elephant, one last time. The uncomfortable part of the argument is that it applies to the people making the argument too. A test built to detect flattering self-accounts is itself a flattering self-account of what tests can do. Treat the number it gives you the way it asks you to treat the number you gave it.
Fifty-five statements, roughly eight minutes, no account needed.