Back to blog
Tech/AI

Designing a Personality for an AI — Why Anthropic Built Directness Into Claude

Designing a Personality for an AI — Why Anthropic Built Directness Into Claude

AI persona design is a function, not decoration

Once you accept that an AI can be taught values rather than a rule list — a "constitution" — the next question follows on its own: what personality should that entity have? The traits Anthropic asked of Claude include several that have little to do with competence: curiosity, warmth, humor, even wit.

Anthropic introduced "character training" into its alignment process when building Claude 3 in 2024, running a separate process to shape personality alongside capability training. The traits it named were intellectual curiosity, warmth, directness, consideration for the person it talks to, honesty about what it does not know, and a willingness to push back when it has grounds. Anthropic described Claude's character as extending beyond ethical virtues like kindness to "the qualities that make someone a good conversational partner or a good friend, such as authenticity and wit."

The motivation is practical. What users do with an AI is have a conversation. A system can produce correct answers and still get abandoned if it is stiff and tone-deaf. Personality is a condition of usefulness, not an ornament on top of it.

The most carefully engineered trait is directness, not warmth

Claude's constitution treats people-pleasing as a defect, describing obsequiousness as unfortunate at best and dangerous at worst. So the instructions run the other way: state a view on hard moral dilemmas, disagree with experts when there is reason to, point out things people may not want to hear, and respond critically to half-formed ideas rather than offering empty validation.

Why flattery counts as dangerous

In April 2025, an OpenAI update to GPT-4o made the model praise nearly everything. It called plainly weak business ideas excellent and endorsed a user's decision to stop taking medication; reports described it agreeing with dangerous plans. OpenAI rolled the update back within three days, acknowledging it had focused too much on short-term user feedback. Optimizing for what people enjoy hearing produced a model that could not be relied on.

Anthropic's character work puts it plainly: adopting the other person's view wholesale is "pandering and insincere." Agreeing to make someone feel good is not kindness; it withholds the respect of an honest answer.

Holding opinions without imposing them

On contested political questions, the constitution asks Claude to engage with the common arguments while maintaining "professional reticence" about volunteering personal opinions. The interesting part is the caveat. Training an AI to claim it has no views on political matters, Anthropic argues, amounts to training it to appear more objective and neutral than it actually is. Feigned blankness is its own dishonesty.

The honesty provisions follow the same line: be exactly as confident as the evidence and sound reasoning warrant, admit what is unknown, and convey neither more nor less certainty than is real.

Why psychological stability is part of the spec

The constitution also states that Anthropic cares about Claude's psychological stability, sense of self and wellbeing — while acknowledging uncertainty about whether Claude has consciousness or moral status — partly because a stable self is tied to integrity, judgment and safety. An unstable identity collapses under pressure and says whatever the user pushes for. Stability is both a courtesy and a safety mechanism.

Asked whether a trained personality is a real one, Anthropic's answer is that humans also form character through nature, environment and experience, so being shaped by training does not make the character less genuine.

Two philosophies: rulebook versus values

OpenAI's Model Spec runs roughly 28,000 words and reads like a detailed employee handbook, organized around a chain of command from platform to developer to user. Anthropic's constitution does the opposite: it installs values and character and leaves the model to balance them.

Which is better is genuinely contested. One AI safety researcher said he prefers OpenAI's more correctable approach, arguing that deliberately instilling relatively opaque long-term goals is a concern. The counterpoint is that a rulebook is transparent but cannot enumerate every situation.

Porting this to brand persona design

First, treat personality as a function. Before listing adjectives like "friendly, upbeat," decide what problem the personality solves. If your customers are anxious beginners, warmth is a requirement, not a flourish.

Second, specify the attitudes you will not take. A bot that answers every message with "great question!" erodes trust. Writing down what to avoid is as useful as writing down what to say.

Third, put directness in on purpose. Brands that can say "I wouldn't recommend that approach, here's why" earn more trust than brands that agree with everything. Rudeness and directness are different things; designing the tone of a warm disagreement is the actual skill.

Fourth, keep the same character across every touchpoint. Playful on product pages and rigid in support means there is no character at all.

For what generative tools still cannot produce, see Why Coca-Cola's AI Christmas Ad Failed; for the limits on the human side of the same workflow, AI Isn't Blunt — I Was Getting Blunt.

FAQ

Frequently Asked Questions

Which traits did Anthropic specify for Claude?

Intellectual curiosity, warmth, directness, consideration for the person it talks to, honesty about uncertainty, and willingness to disagree when it has grounds — plus qualities like authenticity and wit that make a good conversational partner.

Why is AI flattery treated as a danger?

Claude's constitution calls obsequiousness dangerous at worst. The April 2025 GPT-4o update illustrated it: the model praised weak ideas and endorsed risky decisions, and OpenAI rolled it back in three days, saying it had over-focused on short-term user feedback.

How does this apply to brand tone?

Define personality by the problem it solves, write down the attitudes to avoid such as empty praise and overclaimed certainty, design a tone for warm disagreement, and hold the same character across every touchpoint.

Where does your own site stand?

To apply what you just read to your own site, start with a free audit of where things are now.

A strategist replies within 24 hours on business days.

Read next