What happens to your answers

The trait half of this assessment is ordinary: fifty public-domain statements that measure the Big Five, scored the standard way. The unusual half is five more statements that ask how you are answering, and a correction built on what they reveal. Most quizzes take you at your word. This one does not.

Here is the whole machine, in order, without the jargon.

The problem

You are the least neutral person to ask about you

Ask someone to rate their own honesty and you get a number. The trouble is that the number is produced by the same mind that benefits from a flattering answer.

This is not lying. Almost nobody sits down to a questionnaire planning to deceive anyone. It is quieter than that: the answer that comes to mind first is the one that fits the story we already tell about ourselves, and that story was built to be told. The gap opens up without anyone deciding to open it.

The gap between a self-rating and the behaviour behind it Two panels side by side. On the left, a self-rating shown as a bar filled close to the top. On the right, the same trait measured by behaviour, shown as a shorter bar. A bracket between them is labelled "the gap". What you would say "I help because I want to help." Self-rating 8.5 What you would do …when nobody finds out you helped. Behaviour 6.4 the gap what we estimate One trait, measured two ways
The assessment never observes your behaviour directly. It infers the size of this gap from the shape of your answers, then shifts your score by that estimate.
Where the idea comes from

A small bird that fights for the right to be generous

In the Negev Desert there is a bird called the Arabian babbler. Babblers feed each other, groom each other, and take turns standing sentinel on an exposed branch, watching for hawks while the rest of the group forages in safety.

What the biologist Amotz Zahavi noticed, watching them for decades, is that they do not merely take turns. They compete for the turn. A dominant babbler will shove a rival off the sentinel post and force food on a subordinate that did not ask for it. The generosity is a contest, and the prize is rank.

A sentinel bird on an exposed branch above a sheltered flock An illustration of a single bird standing on a high, bare branch in the open, watching a distant hawk, while three other birds forage safely on a lower branch below. Annotations mark the exposed position as costly and the lower position as safe. a hawk Exposed first to be eaten Safe below fed, groomed, and outranked
The sentinel is not being kind at its own expense by accident. The expense is the message. Only a bird in genuinely good condition can afford to stand where the hawk can see it and still survive the week.

That is the logic the whole assessment is built on. A signal is believable in proportion to what it costs to send. Anyone can say they are generous. Being generous when it is expensive, unwitnessed, and uncredited is a different claim entirely, because the cheap version of it is not available.

Where humans differ from the bird. The babbler's display is honest because it is physically costly. Ours has an extra layer: we hide the competitive motive from ourselves. A motive you have never consciously noticed is one you cannot accidentally leak through a pause or a glance, which makes the performance considerably more convincing. This idea, and the babbler with it, comes from The Elephant in the Brain by Kevin Simler and Robin Hanson.

Step one

Fifty-five statements, two different jobs

You rate each statement from 1 to 5. They look like one questionnaire. They are really two, and the second exists to check the first.

The 55 statements, fifty scored across five traits plus five cross-checks Five rows of ten squares, one row per Big Five trait, labelled O for openness, C for conscientiousness, E for extraversion, A for agreeableness and N for neuroticism. Below a dividing line, a sixth row of five warm orange squares marks the cross-checks. O Openness 10 C Conscientiousness 10 E Extraversion 10 A Agreeableness 10 N Neuroticism 10 Cross-checks 5, none of them scored as a trait
50 scored statements, ten per trait 5 cross-checks, which drive the correction
The fifty are the IPIP Big-Five Factor Markers, a public-domain instrument. The five are this project's own, and they never appear in your trait scores.

The fifty scored statements

These are deliberately ordinary. They are the IPIP Big-Five Factor Markers, fifty plain statements like "I am always prepared" and "I have a vivid imagination", ten for each trait. They are in the public domain, they have been used for decades, and they are the exact items behind the million-response dataset this project's model is trained on. Using anything bespoke here would mean the model had nothing to learn from. Half of them are reverse-worded, so agreeing with everything does not produce a high score on anything.

The five cross-checks

This is where the project's own argument lives. They ask about the machinery rather than the person: whether you find it easy to admit a mistake, whether you are genuinely humble or only performing it, whether you behave the same with friends as at work, whether your own motives are ever opaque to you, and whether you can recall a specific time you acted selfishly.

They are deliberately awkward. Someone answering carefully produces a mixed, slightly uncomfortable pattern here. Someone reaching for the good answer every time produces a suspiciously tidy one, and a tidy pattern sitting next to five confident trait scores is exactly the contradiction the correction is looking for.

Step two

Five traits, scored the ordinary way

Each trait averages its ten items and lands on a scale from 0 to 100. There is nothing inventive about this half, and that is deliberate. Using the standard instrument is what lets your scores mean the same thing they mean elsewhere, and what lets a model trained on a million other people's answers say anything useful about yours.

O

Openness

Curiosity, imagination, and appetite for the unfamiliar.

"I have a vivid imagination." "I am full of ideas."

C

Conscientiousness

Order, follow-through, and attention to detail.

"I am always prepared." "I shirk my duties." (reversed)

E

Extraversion

Sociability, assertiveness, and energy around people.

"I start conversations." "I keep in the background." (reversed)

A

Agreeableness

Warmth, empathy, and concern for other people.

"I sympathise with others' feelings." "I insult people." (reversed)

N

Neuroticism

Proneness to worry, tension, and shifting moods.

"I worry about things." "I am relaxed most of the time." (reversed)

Half the items run backwards. Twenty-two of the fifty are worded so that agreeing means a lower score, and they are scattered through the questionnaire rather than grouped. Someone who agrees with everything therefore lands near the middle on every trait rather than at the top, which is one of the reasons this instrument has lasted.

Step three

The adjustment, and how big it gets

Once the five raw scores exist, the system compares them against the cross-checks and asks a single question: does this set of answers hang together, or is it too good to be internally consistent?

Suppose someone scores very high on agreeableness, then says they find it hard to admit a mistake and cannot recall ever acting selfishly. Each answer is defensible on its own. Together they describe a person who has never once been at odds with themselves, which is not a kind of person that exists. That combination raises the adjustment.

The adjustment scale, from light to substantial A horizontal bar running from zero to one, shaded from pale to deep teal, divided into three bands labelled light, moderate and substantial, with a marker showing an example reading. How much correction is applied Light answers hang together Moderate some tension between them Substantial the picture is too clean none maximum an example reading
A light reading is not a compliment and a substantial one is not an accusation. It describes the answers, not the person: it says how much the six numbers should be trusted at face value.

What the adjustment does to your scores

The correction is not applied evenly. It lands hardest on traits that stand out from the rest of your profile, because an unusually high score is where a flattering answer would have shown up. It is also weighted per trait, and neuroticism moves the other way: agreeableness and conscientiousness are the ones people talk up, while neuroticism is the one they play down, so the correction raises it rather than lowering it.

Three areas before and after adjustment Grouped horizontal bars for three traits. In each pair the upper bar is the score as answered and the lower bar is the score after adjustment. Agreeableness falls from 86 to 71, conscientiousness from 79 to 70, and neuroticism rises from 24 to 31. Agreeableness As answered: 86 out of 100 86 After adjustment: 71 out of 100 71 Conscientiousness As answered: 79 out of 100 79 After adjustment: 70 out of 100 70 Neuroticism As answered: 24 out of 100 24 After adjustment: 31 out of 100 31 0 100
As answered After adjustment
Illustrative figures, not real results. The two traits people talk up come down; neuroticism, the one people play down, goes up. That asymmetry is the mechanism.

This is an opinionated step, and worth saying plainly. Subtracting an estimated amount of self-flattery from someone's self-report is this project's central bet, not a settled method in psychology. Both numbers are kept and both are shown to you, so you can disagree with the correction and read the originals.

End to end

The whole journey, in six moves

From the moment you press submit, your answers pass through six stages. Each one hands its output to the next, and every intermediate result is kept.

The six-stage processing pipeline Six boxes connected left to right by arrows: your answers, five trait scores, consistency check, adjusted scores, pattern match, and written summary. Each box carries a short description of what happens there. 1 Answers 55 answers, 1 to 5, plus age and sex 2 Trait scores five averages, ten items each 3 Cross-check how far the story holds together 4 Adjust per-trait correction, largest at the peaks 5 Match nearest of six learned patterns 6 Write-up a plain-language summary IN YOUR BROWSER ON THE SERVER Everything from stage 2 onward is saved, so your results page can be reopened later.
Stages 2 through 4 are ordinary arithmetic and give the same answer every time. Stage 5 multiplies your answers through a small trained network. Stage 6 writes the reading, on the same server, with no external service involved.
Step five

Finding the nearest neighbourhood

Your fifty answers are fed through a small network trained on just over a million other people's answers to the same fifty statements. It compresses them to sixteen numbers, and the nearest of six patterns found in that data becomes your result.

The map is not a picture of anything you would recognise. It is a compact space where profiles that answered similarly end up near each other, which lets the system express "this profile resembles that one" as a distance rather than a rule. The six patterns were not chosen in advance: they are clusters found in the training data, and they are named after whichever trait sets each one apart. The confidence figure on your results is how much closer you sit to your nearest pattern than to the next one along.

Six learned patterns with one profile placed among them Six soft circles, each named for a pattern found in the training data, spread across a rectangular field. A single marked point labelled "your profile" sits between two of them, with a dashed line to the nearer one. 1 Explorer 2 Organiser 3 Connector 4 Observer 5 Diplomat 6 Anchor your profile nearest next nearest The gap between "nearest" and "next nearest" is what the confidence percentage on your results reports. A profile sitting midway between two neighbourhoods gets a lower confidence figure, not a wrong answer.
Six patterns, learned rather than invented. Landing between two of them is common and is reported honestly rather than rounded away.

A note on the six patterns. They are reference points for describing a profile, not boxes people belong in. Two people matched to the same pattern can differ a great deal, and the five trait scores underneath carry far more information than the name on top of them.

Under the hood

How the pieces are actually wired

For the curious: the site is a plain static front end talking to a small Python API. Nothing runs in your browser except the questionnaire itself and the code that draws your results, and nothing your answers touch leaves that server. There is no external model API in this system at all.

System architecture across four layers Four stacked layers. The browser layer holds three pages. The API layer holds the endpoints for starting, submitting and reading an assessment. The engine layer holds the scoring engine, the pattern model and the summary writer. The storage layer holds the saved records. Arrows run downward from browser to API to engines to storage. YOUR BROWSER static pages, no framework, no build step Landing and this page explains the method Questionnaire collects the 55 answers Results draws the charts and panels requests over HTTPS API a small Python service Start an assessment creates the record, returns an id Submit answers runs the whole pipeline Read results returns everything computed ENGINES each one independent, so a failure in one does not take the page down Scoring averages, cross-check, correction Pattern model trained weights, run in numpy Summary writer runs locally, no external service STORAGE your answers, both sets of scores, the adjustment, the match, and the write-up
The results page loads your headline result first and fills in the deeper panels afterwards, so a slow or missing panel never blocks the part you came for.

Plain-language glossary

Raw score
The average of your ten answers for one trait, reverse-keyed where needed, before anything is subtracted.
Cross-check
One of the five statements that probes how you answer, rather than what you are like.
Adjustment
How much the system thinks your answers were shaded, expressed as a number between none and maximum.
Adjusted score
The raw score with the estimated shading removed. Both numbers are shown, and both are kept.
Percentile
Where your trait score falls in the reference sample. A 70th percentile means you scored higher than 70% of the people in it.
Confidence
How clearly your profile sits nearer one reference pattern than the others. Low confidence means you are between two, not that something went wrong.
Novelty
How far your answers sit from the patterns in the million-response reference sample. Unusual is not better or worse.
Honest limits

What this is, and what it isn't

A method this opinionated should be equally clear about where it stops.

It is a research prototype

The trait half uses a long-established public instrument, and the model behind the patterns is trained on a public dataset of over a million responses. The correction on top of it has not been checked against any outside measure of behaviour.

It is not a clinical instrument

Nothing here diagnoses anything. It should not inform decisions about hiring, admissions, treatment, or anyone's standing.

The adjustment is a hypothesis

Correcting a self-report for estimated self-flattery is this project's argument, not established practice. Your uncorrected scores are kept and shown alongside.

A single sitting is a thin sample

Fifty-five statements on one afternoon capture a mood as much as a disposition. Read the result as a prompt for reflection, not a verdict.

On the elephant, one last time. The uncomfortable part of the argument is that it applies to the people making the argument too. A test built to detect flattering self-accounts is itself a flattering self-account of what tests can do. Treat the number it gives you the way it asks you to treat the number you gave it.

See where your answers land

Fifty-five statements, roughly eight minutes, no account needed.

Start Assessment Back to Home