Extracting your brand voice from your own existing content (and why most tools get this wrong)
By Kane Harrison
Most AI marketing tools that say they support "brand voice" want you to upload a brand voice guide.
This is backwards. The guide is the last artefact you should optimise for. The first artefact you should optimise for is your existing content — the website, your top-performing blog posts, your last few approved campaigns. That content has your real voice, encoded in usage patterns the writer probably can’t articulate.
The 5-attribute extraction
We use five attributes for voice:
- Specificity — concrete vs vague language
- Calmness — measured vs hype-driven energy
- Commercial focus — outcome language vs feature language
- Control posture — emphasis on guardrails vs creative freedom
- Professional register — operator/exec vs consumer/influencer
Each scored 0-100 against your existing corpus, with example phrases pinned to each score.
How extraction actually works
Three pipeline steps:
- Ingestion — chunk the corpus into ~400-token blocks, embed each with text-embedding-3-small, store in pgvector.
- Attribute scoring — for each attribute, classify a sample of chunks against canonical "high" and "low" examples we’ve curated. The score is the proportion of chunks closer to "high" than to "low".
- Phrase pinning — for the highest-scoring chunks per attribute, extract phrase n-grams (1-6 words) and surface them as "preferred phrases". These become retrieval-augmented additions to every generation prompt.
Why most tools get this wrong
Most tools either:
- Skip extraction and ask the writer to write a voice guide (slow, often wrong)
- Use a single generic "tone" parameter (too coarse to be useful)
- Train a fine-tuned model on your corpus (expensive, brittle, hard to update)
Retrieval against an embedded corpus, with attribute scoring + phrase pinning, gives you a voice profile that’s grounded in your real content, easy to update, and transparent to the writer.
The metric that matters
A working brand voice tool should let you generate copy, score it against your profile in real time, and surface specific drift ("the calmness score dropped because you used 'unlock' twice"). If your tool doesn’t do this, the brand voice claim is doing more work than the tool is.
This is one of the things Marketeer’s brand brain does at generation time. Score is on every variant in the messaging step. Phrases pinned from your highest-voice content are auto-preferred. Forbidden phrases auto-blocked.
The point isn’t magic. The point is that the voice you ship is the voice you’ve been shipping all along — just more consistently.
Want a marketing team without the headcount?
Start free and run a real campaign with the whole team, or talk to us for a quick walkthrough.
Start free