A versioned prompt library built from your own accepted and rejected passages, with an eval harness that catches voice drift before publication.
About this service
Brand voice defined as ten pairs of accepted and rejected passages from your own site, not three adjectives on a slide. That is the deliverable: a versioned prompt library in your git repository with an evaluation harness that fails a prompt change the way a test suite fails a code change. Two weeks of the engagement goes into collecting the exemplars, because everything downstream is only as good as that set.
Why adjectives fail:
Confident, human, never salesy produces different output from every model and every writer who reads it. Paired examples do not. A rejected passage next to the accepted rewrite of the same paragraph carries the actual rule, including the rules your team has never written down: that you do not name competitors, that you never state a return window without its exception, that Arabic pages open with the answer rather than the greeting.
What is in the library:
System prompts as files under semantic versioning, one per content type, each declaring the exemplar set it depends on. Task prompts that reference them rather than copying them. A golden set of roughly 60 outputs labelled by your own editors as ship, fix or reject. An LLM judge scored against that golden set, used only once its agreement with the human labels holds. If the judge disagrees with your editors, the judge is wrong and gets rebuilt.
Every prompt change runs the golden set and produces a diff of scores. Voice drift becomes a number people can argue about in a review, instead of a feeling somebody raises three weeks after publication.
Model changes are treated as prompt changes. When a version is deprecated or a provider updates one quietly, the same run shows what moved. On fashion and DTC catalogue copy this has turned out to be the main use of the harness, well past the initial build.
Arabic:
Arabic is a separate library, not a translation of the English prompts. Register is decided per surface: Modern Standard for published pages across the Gulf, Gulf dialect only where your social team already writes that way. Exemplars are Arabic originals from writers you already trust. I do not accept translated English exemplars into an Arabic library; they teach the model to write English in Arabic script, which is exactly the output your readers skip.
Not included:
No content production. No fine-tuning for voice; it costs more, reverts slower, has to be redone on every base model change, and hides the rule inside weights where no editor can read it. No prompt library living in Notion or a shared doc. If it has no version history and no review, it is folklore, not a library. No tone wheels, messaging houses or brand workshops.
Who this is not for:
Companies without editors. The harness measures agreement with human judgement, so it needs humans whose judgement is settled. If your team still disagrees about whether a paragraph is on brand, that is an editorial decision to make internally first, and I would rather say so in a call than sell you a harness that averages the disagreement.
Handover:
The repository, the exemplar set, the golden set, the judge rubric, and a written record of every threshold with the reasoning behind it. You can run the whole thing without me, and the reason the reasoning is written down is so you can also change it without me.