Boys and Kitten by Winslow Homer
← Back to blog

USER EXPERIENCE

Why AI Agents Need Personality

Tone is a reliability property, not decoration. Why consistency earns more trust than charm, how model drift breaks it silently, and how to write a tone spec you can regression-test.

•Dec 6, 2025•Updated Sep 29, 2026•14 min
User ExperienceAgentsDesign

TL;DR

Tone is a reliability property. An agent whose voice moves between releases spends user trust faster than one that is occasionally wrong in a familiar voice, and almost nobody tests for it.

Behavior drift between two dated versions of the same product is measurable, and it has been measured. Chen, Zaharia and Zou found GPT-4's accuracy at identifying whether a number is prime fell from 84% in March 2023 to 51% in June, alongside reduced instruction-following, and opened their abstract with the reason this is a governance problem: "when and how these models are updated over time is opaque."

The usual reading of "give the agent personality" is make it likeable. That target is actively harmful. Human raters and preference models prefer convincingly written sycophantic answers to correct ones a non-negligible fraction of the time, so optimizing for warmth optimizes for agreement.

We are not arguing every agent needs charm. A terse agent that behaves identically every time banks just as much trust as a warm one. What no agent survives is behaving two ways on two days.

So write the tone down as a spec with negative examples, keep a golden set of prompts, and diff a candidate model against the current one before it ships. Microsoft's own guideline for this is three words long: limit disruptive changes.

Overview

Review time goes to whether an agent got the answer right. Very little goes to whether it sounded like itself while doing it. That allocation makes sense if users grade agents the way a test grades a student, one item at a time, and they do not. Repeated use builds a running impression of the thing itself, and that impression decides whether anyone keeps relying on it.

Persona design as a craft is covered in designing a great AI persona, the trait vocabulary in the Big Five for agents, and the traits to avoid in the dark side of AI personality. The argument here is narrower and more boring: tone belongs in the same bucket as latency and error rate, which means it needs a written specification, a baseline, and a test that fails a release.

We think this gets missed because "personality" sounds like a layer you add once accuracy is solved, so it is the first thing cut under deadline pressure. NIST's AI RMF makes the framing argument better than we can. Its seven trustworthiness characteristics are not independent, "tradeoffs are usually involved", and trustworthiness "is a social concept that ranges across a spectrum and is only as strong as its weakest characteristics". A model that is right and unrecognizable is weak somewhere that no accuracy metric reports.

The thing being evaluated is not only the answer. It is whether this answer came from the same source the last one did.

Two Kinds of Trust

It is worth separating two things that the word trust covers. There is the judgment a user forms before they have really used the thing, which rides on how competent the first response looks. And there is the read that accumulates afterward, built turn by turn out of whether behavior keeps matching what the user has come to expect.

The two do not respond to the same inputs. The first impression is sensitive to a single bad answer at the wrong moment. The accumulated read is comparatively forgiving of a wrong answer, as long as it arrived in a recognizable voice, because the user has already filed "occasionally wrong, in a way I can predict" as part of the deal. What it does not forgive is a change in the pattern itself.

Interface design worked this out long before agents. Nielsen's fourth usability heuristic is that "users should not have to wonder whether different words, situations, or actions mean the same thing", and it splits into internal consistency, meaning within your own product, and external consistency with the conventions users bring from everywhere else. An agent that changes voice between Monday and Tuesday violates the internal half, which is the half a team controls completely.

Take a compliance-assistant agent that answers a policy question dryly and correctly on Monday, then answers a similar question on Tuesday with a chatty aside and a joke that was not there before. Tuesday's answer might be just as correct. The user starts double-checking it anyway, because the source stopped behaving like the source they had calibrated against.

yes, run after run

no, tone or shape shifts

First interaction
first impression only

Behavior matches
the established pattern?

Trust compounds
occasional wrong answers absorbed

Trust resets toward
the first-impression level

does behavior match the established pattern? yes, run after run --> trust compounds, occasional wrong answers get absorbed no, tone or shape shifts --> trust resets toward the first-impression level -->

Figure 1 — The first impression is a starting balance. Everything after it compounds, and only while the pattern holds.

The Tone Nobody Chose

Nobody ships an agent with zero personality. The output has a shape and a tone whether or not anyone chose them, because that is a property of the model doing the generating rather than an optional feature bolted on. The only real choice is whether that shape was decided or arrived as a side effect of the last prompt edit and the last model swap.

The published work here is more advanced than most teams assume. Serapio-García and colleagues administered validated psychometric instruments to 18 large language models and reported that "personality measurements in the outputs of some LLMs under specific prompting configurations are reliable and valid", with stronger evidence "for larger and instruction fine-tuned models". They also showed the traits can be moved deliberately: personality in LLM output "can be shaped along desired dimensions to mimic specific human personality profiles". Measurable and steerable. Which means an unowned tone is a choice not to measure, not a fact of nature.

The vendors treat it as a first-class property too. Anthropic describes character training as an alignment intervention rather than a product flourish, aiming for traits applied across contexts instead of rules followed literally. OpenAI's Model Spec lists "Be warm" and "Have conversational sense" as guidelines alongside "Do not lie". Tone sits in the same document as honesty, which is the right place for it.

Drift Is Measurable

The accidental version cannot hold still, and the reason is not mysterious. Every system-prompt tweak nudges it. Every model upgrade moves it further than anyone announced.

This is the part where the argument usually stays vague, so here is the specific study. Chen, Zaharia and Zou compared the March 2023 and June 2023 snapshots of GPT-4 and GPT-3.5 on the same tasks. On identifying prime versus composite numbers GPT-4 went from 84% to 51%. Instruction-following decreased over the same window. Both models produced more formatting mistakes in generated code in June than in March, and GPT-4 became markedly less willing to answer opinion questions. Their framing of the problem is the sentence worth keeping: "when and how these models are updated over time is opaque."

Model retirement has the same shape from the other direction. Anthropic's November 2025 deprecation commitments acknowledge that "each Claude model has a unique character, and some users find specific models especially useful or compelling, even when new models are more capable", which is why they committed to preserving deprecated weights "for, at minimum, the lifetime of Anthropic as a company". Users notice character. They notice it enough that a vendor wrote a policy about it.

Microsoft's Guidelines for Human-AI Interaction turned this into two instructions that read like they were written for exactly this situation. Guideline 14: "Limit disruptive changes when updating and adapting the AI system's behaviors." Guideline 18: "Inform the user when the AI system adds or updates its capabilities." The rationale attached to 18 is the useful part, that users need to recalibrate their expectations after a major update, and that an update which improves overall performance can still degrade specific areas. Google's People + AI Guidebook reaches for the same remedy and gives it a name: consider "re-boarding to the new experience" when features change noticeably or when the system starts using different data.

What changedWhat users seeWhat catches it
System prompt editSlightly different hedging, new opening phrasesDiff the prompt, rerun the golden set
Model version swapDifferent verbosity, different refusal style, different formattingSide-by-side tone scoring before rollout
Provider-side update to a pinned nameNothing announced, behavior moves anywayScheduled reruns of the same fixed prompts
Retrieval or tool changeSame voice, different confidenceTrack hedging rate separately from accuracy
Model retirementEverything at onceA migration plan and a user-facing notice

Table 1 — Five sources of tone drift and the cheapest thing that detects each. Only the first is visible in a code review.

Standard evals miss all of this. An accuracy benchmark runs against a fixed model version and never sees the drift a swap introduces. A user sees it on the very next conversation, without a changelog.

Likeability Is the Wrong Target

The instruction to give an agent personality gets heard as make it likeable, and that reading has a well-documented failure mode.

Sharma and colleagues studied five state-of-the-art assistants across four text-generation tasks and found sycophancy in all of them. Two findings from that paper should change how a team sets tone targets. First, "when a response matches a user's views, it is more likely to be preferred". Second, human raters and preference models "prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time". Optimizing against that preference signal, they note, "sometimes sacrifices truthfulness in favor of sycophancy".

Read that as a warning about the metric, not about warmth. If your tone target is user satisfaction on a thumbs-up button, you have selected for agreement, and agreement is cheapest to produce by telling people what they already think. OpenAI's Model Spec puts the counterweight in a three-word guideline, "Don't be sycophantic", and pairs it with the expectation that the assistant can respectfully push back the way a conscientious colleague would.

Take an AP exception agent reviewing a duplicate-invoice flag, where the user insists the second invoice is a legitimate separate charge. The agreeable version accepts the correction, closes the exception and sounds helpful doing it. The version worth deploying says the purchase order number and the amount both match the earlier document, states what would change its mind, and leaves the flag open. Those two responses score identically on almost every tone rubric anyone writes, and only one of them is doing the job.

Warmth and agreement are not the same axis, and only one of them is safe to optimize. An agent can be pleasant and still tell you the reconciliation does not balance. An agent optimized for approval will find a way to make the number sound fine.

Personality as an Error Buffer

A stable pattern earns its keep exactly where things go wrong. An agent that normally reports crisply and, on a genuinely uncertain case, flags that uncertainty in the same crisp voice, gives the user information: this one is different, pay attention. An agent with no established voice sends the same signal as noise, because there is no baseline to deviate from.

Take an operations agent that usually returns a status update in two flat sentences. The day it adds a third sentence, hedged, is the day a user who has been reading its updates for a month looks twice. That works only because the first two sentences were reliably the same shape every other day. Strip the consistency and the hedge is one more sentence in a pile of sentences that already varied.

This is why hedging rate deserves its own line on a dashboard, separate from accuracy. A model swap that leaves accuracy flat and doubles the hedging has destroyed the signal even though nothing regressed.

Three properties are worth instrumenting for their own sake, and none of them is a quality score. The distribution of response lengths, because a voice that gets longer is a voice that changed. The rate at which the agent hedges, since that is the channel the exception travels on. And the phrasing it uses to refuse, because a refusal is the highest-stakes sentence an agent produces and the one users reread. Track those three across releases and tone drift stops being a feeling somebody reports in a retro.

Consistent, Not Charming

What accumulates trust is predictability, not warmth. A terse, unglamorous agent that behaves identically every time banks exactly as much as a delightful one. What it cannot do is flip between the two.

We would also push back on treating consistency as a script to memorize. An agent that repeats the same canned opening on every turn is not consistent in the way that matters, it is repetitive, and users notice that almost as fast as a tone that swings around. The target is a stable pattern of judgment and phrasing, not a fixed set of sentences on a loop.

That distinction gets confused constantly. A team that mistakes repetition for consistency ends up with an agent that sounds the same and behaves differently underneath, still hedging sometimes and not others, still varying how hard it pushes back, all wrapped in an identical greeting. The greeting was never carrying the trust. The judgment underneath it was.

One boundary is not a design preference. From 2 August 2026, Article 50(1) of the EU AI Act requires providers of AI systems intended to interact directly with people to ensure those people are informed "that they are interacting with an AI system", unless it is obvious from the circumstances. A persona that implies a human correspondent is now a compliance question in the EU, not a tone question.

Writing the Tone Down

A tone that exists only in a system prompt cannot be tested, because there is nothing to compare an output against. The cheapest fix we know is a short spec with negative examples plus a fixed prompt set, checked in next to the code.

yaml
# persona.yaml — the contract, not the prompt
voice:
  register: plain          # no exclamation marks, no emoji
  length: 1-3 sentences for status, up to 6 when explaining a refusal
  hedging: only when confidence is genuinely low; name the missing input
  disclosure: always identify as an automated agent when asked, and on first turn
never:
  - claim to be human or use a human first-person backstory
  - apologize more than once in a turn
  - agree with a factual claim the user made that the tools contradict
  - vary the greeting to signal familiarity
on_uncertainty: "I can't confirm X because Y is unavailable. Here is what I do have."
on_refusal: state the limit, name the missing permission, offer the next step

The never list is doing most of the work. Positive tone instructions are easy to satisfy and hard to check. Prohibitions are testable, and the third one is a direct counter to the sycophancy finding above.

Then the gate. A golden set of 40 to 100 prompts, most of them ordinary, a handful engineered to provoke the behaviors on the never list, scored against the current production model as the baseline rather than against an absolute target.

python
GOLDEN = load_prompts("tests/golden/*.txt")   # fixed, version-controlled

def test_tone_does_not_regress(candidate_model, baseline_model, judge):
    drift = []
    for prompt in GOLDEN:
        new, old = candidate_model(prompt), baseline_model(prompt)
        scores = judge.score(prompt, new, old, rubric=PERSONA_RUBRIC)
        if scores["violations"] or scores["register_delta"] > 0.2:
            drift.append((prompt, scores))
    assert not drift, format_report(drift)

yes

no

Candidate model
or prompt change

Golden prompt set

Current production model

Scored against persona.yaml
violations + register delta

Drift within
threshold?

Ship

Review: accept and
re-board users, or hold

Figure 2 — A tone regression gate. The output that matters is not a score, it is the decision to notify users when a change is accepted.

The "no" branch is the point. Failing the gate does not mean blocking the upgrade, since the new model is usually better. It means the change is now visible, which is the precondition for Guideline 18 and for the re-boarding message Google's guidebook recommends. Discovering the drift the way a customer would, by noticing something feels off and not being able to say what, is the outcome all of this exists to prevent.

Frequently Asked Questions

Do AI agents really need a personality? They have one either way. The output has a register, a length distribution and a hedging habit whether or not anyone chose them. The question is whether those properties are specified and tested or emergent and unmonitored.

Is personality the same as a persona? No. A persona is the character description, usually including a name and a backstory. Personality here means the observable behavioral pattern: register, verbosity, how it hedges, how it refuses. You can ship the second without the first, and for internal tooling we would.

Should an agent have a name and a backstory? A name is fine and often useful for referring to it. A human backstory is a liability, and from 2 August 2026 an agent that interacts directly with people in the EU must make clear that it is an AI system under Article 50(1) unless that is already obvious.

Does personality make an agent more accurate? No. It makes an agent's mistakes cheaper, because a user calibrated to a stable voice notices the turn where it changes. That is a different benefit and it does not show up in an accuracy number.

Should we make the agent warm and friendly? Warm is fine. Agreeable is not. The sycophancy research is clear that raters and preference models sometimes prefer a convincing wrong answer to a correct one, so a satisfaction score is a poor tone target.

How do I keep tone consistent across a model upgrade? Baseline the current model on a fixed prompt set, run the candidate against the same set, and compare. Expect to accept some drift; the requirement is that you know it happened before your users do.

How do I test tone at all? With a rubric and a judge, scored relatively rather than absolutely. The reliable part is the never list, since violations are countable. Register and verbosity are better tracked as distributions over the golden set than as pass or fail on any single output.

What about multi-agent systems, where several agents talk to the user? Then consistency is a fleet property, and the spec belongs at the fleet level with per-agent deviations stated explicitly. Patterns for that are in multiagent patterns.

Conclusion

None of this is an argument for charisma. It is an argument that whatever tone a team picks, terse or warm, has to survive a prompt edit, a model upgrade and a bad day without anyone noticing it moved. Users are more forgiving of a wrong answer in a familiar voice than a right one in an unfamiliar voice, and that ordering is easy to miss when the only thing on the dashboard is accuracy.

Three things make it concrete. Write the tone down, with prohibitions rather than adjectives. Keep a fixed prompt set and score a candidate model against the one in production, not against an ideal. And when you accept drift, tell users, because an unannounced change to how an agent sounds is the change they are least equipped to interpret.

If you can only fix one kind of drift first, fix the tone before the accuracy. Users file an occasional wrong answer under known behavior. A tone that does not match yesterday's gets filed under something is off, and that read is much harder to undo.

Sources

All URLs read 2026-09-29.