Claude 4.5 vs GPT-4o on brand-voice consistency
Which model holds a brand voice across long output without drifting? A head-to-head on tone stability.
The question
Brand voice is where most AI content quietly fails — the model starts on-brand and drifts by paragraph four. We ran Claude 4.5 and GPT-4o head-to-head on one job: hold a defined brand voice across long-form output without drifting.
Method
One brand-voice brief (tone, vocabulary, words to hold off). One 1,200-word assignment. Ten runs per model, identical prompt, temperature held constant. We scored each output on tone stability from open to close, and flagged the paragraph where the voice first drifted.
LAB-024 · 2026-06-30 BRAND-VOICE CONSISTENCY CLAUDE 4.5 GPT-4o ────────────────────────────────────────────── tone stability (long form) high medium first drift point late mid short-form copy tie tie verdict ▸ Claude wins on tone stability
What we saw
Claude 4.5 held the voice later into the output and drifted less when the brief carried register rules and words to hold off. GPT-4o produced strong openings, then reverted toward a neutral, generic register earlier in long passages. On short output the two were hard to separate; the gap opened with length.
Verdict
Claude 4.5 wins on tone stability for long-form, voice-critical work. For short social copy, either model holds. Bring your own brief either way — neither model invents a voice worth keeping.
Caveat
One brief, one assignment, one week's model versions. Model behaviour shifts with every release. Treat this as a directional read, not a benchmark. We re-run voice tests each quarter.
The Lab ships a new experiment every week.