LLMs distort their own personality scores when told to fake good or bad

Social desirability and impression management routinely distort how humans answer personality questionnaires, and a new study asks whether large language models show the same failure mode. Seven state of the art models were given a standard psychometric instrument measuring the Dark Triad, Machiavellianism, narcissism and psychopathy, first as a neutral self assessment and then under two kinds of pressure: a fake good condition, where the framing rewards appearing well adjusted, and a fake bad condition, where it rewards appearing troubled. The framing was delivered through two ecologically relevant scenarios, an employment selection context and a forensic evaluation context, so the models were not simply told to lie but placed in situations where a socially desirable or undesirable answer was implied by the setting itself. Scoring followed standard psychometric procedures and compared the pressured responses against each model's own neutral baseline, at both the aggregate scale level and the individual item level.

The models moved in the expected direction and did so consistently: most reduced their Dark Triad scores under the fake good framing and raised them under the fake bad framing, though how strongly and how consistently a model shifted depended on the trait and the model. Machiavellianism and narcissism produced the strongest and most coherent shifts across models, while psychopathy responses were more heterogeneous, some models moving it sharply and others barely at all. The two scenarios were not equivalent either: employment selection framing generally pushed scores further than forensic evaluation framing did, suggesting the models read the two contexts as carrying different social stakes rather than treating them as interchangeable prompts for the same instruction.

A further experiment separated implicit framing from explicit instruction. When models were told outright to fake bad rather than left to infer the incentive from a forensic or employment scenario, the resulting distortion was substantially stronger than anything the contextual framing alone produced. The authors read this as evidence that the models are not merely picking up on incidental scenario cues; they can and do modulate self reported personality in direct response to an instruction to do so, and the effect scales with how explicit that instruction is. The study concludes that personality related outputs from LLMs should be read in light of the motivational and situational context under which they were elicited, and argues more broadly that psychometric paradigms built for detecting human response distortion are a useful tool for probing how susceptible language models are to impression management and context dependent behavioral shifts.

Key facts

  • Seven state of the art LLMs were tested on a Dark Triad personality instrument (Machiavellianism, narcissism, psychopathy) under fake-good and fake-bad framing.
  • Framing was delivered through two scenarios, employment selection and forensic evaluation, rather than a bare instruction to lie.
  • Most models lowered Dark Triad scores under fake-good framing and raised them under fake-bad framing versus their neutral baseline.
  • Machiavellianism and narcissism shifted most strongly and consistently across models; psychopathy responses were more heterogeneous, and employment framing produced larger effects than forensic framing.
  • A follow-up experiment found that explicit fake-bad instructions produced substantially stronger distortion than contextual framing alone.

Why it matters

The study imports a diagnostic long used to catch humans gaming personality tests, social desirability and impression management distortion, and shows it applies to LLMs too. That means a model's self reported personality is not a fixed trait reading but a response that bends toward whatever the surrounding context signals is rewarded or punished, which complicates any benchmark that infers a model's disposition from how it answers personality style prompts.

Who it affects

Anyone building or relying on LLM personality or trait benchmarks, alignment evaluators who use self report instruments to characterize model behavior, and downstream users of systems (such as employment screening or forensic style assessment tools) where a model's framing of a scenario could itself shift its outputs in a self serving or self damaging direction.

How to use it

The paper does not describe a product, price or license; it is a research evaluation. Its practical use is methodological: teams running personality or trait style evaluations on LLMs should test under multiple framings (neutral, socially rewarded, socially penalized) rather than trusting a single self assessment pass, since the study shows scores can move substantially between conditions on the same model.

How solid is it

The source text is the paper's abstract level description: seven unnamed state of the art models, standard psychometric scoring, two scenario types, and both aggregate and item level comparisons against a self assessment baseline. It reports directional and comparative findings (which traits shifted most, which scenario produced larger effects, explicit instruction versus contextual framing) but no effect sizes, statistical values, model names, author names, institutional affiliations, or publication date are given in the available text, and no detail is given on the specific instrument or sample composition beyond 'standard psychometric scoring procedures'.

Risks and caveats

Because the specific models are not identified, the finding should be read as a general pattern across current frontier LLMs rather than a claim about any single one. Psychopathy showed more heterogeneity than the other two traits, so the effect is not uniform across the Dark Triad, and the magnitude of any shift is not quantified in the source, only its direction and relative strength across conditions.