How We Handed "Copywriting + A/B Testing" Over to AI: A Manager's Perspective
Anyone who has led a mobile product team has probably experienced this scenario: a new feature is about to launch, and you need push notification copy, splash screen onboarding text, and subscription page conversion copy — three versions of each. The product manager writes two versions, operations writes two, the designer says "the headline is too long," and the boss drops a comment in the group chat like "it doesn't feel impactful enough." A week goes by and the copy still isn't finalized. In the end, it's usually the person with the highest job title who makes the call — and then nobody knows which version actually performed better.
That's what my team looked like six months ago. Today I want to share what happened after we handed "copywriting + A/B testing" over to AI — the process changes, the collaboration changes, and how we kept the risks under control, from a manager's perspective.
1. Laying Out the Pain Points First
From a manager's perspective, the waste in copywriting comes down to three things:
- Wasted manpower: Copy output depends on specific individuals. If one person takes a vacation, the copy gets stuck. Writing copy isn't hard — what's hard is the endless back-and-forth communication and revisions.
- Decisions based on gut feeling: Which version is better ultimately comes down to who argues loudest, not data. Even when A/B tests are run, the conclusions are often unreliable due to too few variants and arbitrary test durations.
- Nobody owns the risk: Exaggerated claims, prohibited words, or sensitive phrasing in copy usually get discovered only after launch. When problems arise, nobody can explain who reviewed that version of copy or by what standards.
To sum it up in one sentence: copy production is artisanal, copy decisions are based on intuition, and copy risk control is nonexistent.
2. What the Process Looks Like After AI Steps In
We didn't just let AI "generate copy with one click" and call it a day — that would only have moved the chaos earlier in the process. What actually works is embedding AI into the workflow, with humans only gating the key checkpoints. The whole process has four steps:
Step 1: Build a Copywriting Standards Asset
First, we had AI consolidate the copywriting styles scattered across historical versions into a "brand tone and compliance guidelines" document, including character limits, a list of prohibited words, naming conventions, and tone of voice. This took one afternoon, but it's the prerequisite for all automated generation that follows — AI generation without standards is an assembly line with no quality control.
Step 2: Batch-Generate Multiple Sets of Candidate Copy
For each functional module (push notifications, splash screens, subscription pages), generate 3–5 candidate sets, each with an attached "design intent explanation" — for example, this version leverages loss aversion, that one emphasizes concrete numbers. The intent explanation matters: it means the review discussion becomes "is the hypothesis correct" rather than "does this sentence read smoothly."
Step 3: AI Pre-Review + Human Final Review
AI first runs a compliance pre-review based on the standards from Step 1, flagging potentially exaggerated claims and sensitive words. Humans only review what AI flags plus the final shortlisted versions. The manager's time goes where it matters most, not into reading twenty push notifications word by word.
Step 4: Hook Into A/B Testing, with AI Assisting with Data Analysis
Test variants, traffic allocation, and observation periods are defined as rules and written into documentation up front. AI is responsible for aggregating conversion data across groups on schedule and providing significance judgments. AI reads the data; humans judge whether the conclusions are trustworthy — this is now the only agenda item where we discuss copy in our regular meetings, and it usually wraps up in ten minutes.
3. A Reusable Prompt Template
Here's the multi-variant copy generation template we use internally — take it, adapt it, and put it to work:
# 角色
你是移动端应用的文案专家,服务于[产品名称],目标用户是[用户画像]。
# 任务
为[功能模块:推送通知/开屏页/订阅页]撰写 N 套候选文案,每套包含:
1. 主文案(不超过 X 字)
2. 辅助文案(如有)
3. 设计意图(说明该版本基于什么心理动机,如损失厌恶/社会证明/具体收益)
# 约束
- 符合以下语调规范:[粘贴品牌语调规范]
- 禁用词/禁用表述:[粘贴禁用词清单]
- 不得出现绝对化用语、夸大承诺、虚构数据
- 不得使用"最""第一""100%"等违规修饰
# 输出格式
以表格输出:版本号 | 主文案 | 辅助文案 | 设计意图 | 自查(是否触碰约束项)The key to this template is the "design intent" and "self-check" columns: the former gives the A/B test a hypothesis to validate, and the latter has AI run its own compliance check first, directly cutting the human review workload in half.
4. Before and After AI
| Dimension | Before AI | After AI |
|---|---|---|
| Copy production cycle | 2–3 versions, ground out over a week by hand | 5 candidate versions generated at once, ready for review within half a day |
| People involved | PM, operations, and design all dragged into discussions | One person works with AI output; everyone else attends a ten-minute review |
| A/B testing | Done occasionally, conclusions based on impressions | Standard for every launch, data aggregated automatically on schedule |
| Compliance risk | Discovered via user complaints after launch | AI pre-review + standards checklist intercepts issues up front |
| Knowledge retention | Standards lived in veteran employees' heads | Standards are a document that AI references on every generation |
As for hard numbers — we've internally observed a noticeable drop in the number of copy-related meetings and rework rounds, but the actual conversion lift varies widely across products. I won't make up numbers here; I'd suggest testing with your own product.
5. The Three Things Managers Actually Need to Manage
Once the process runs smoothly, my takeaway is that in the AI era, managers need to hold the line on three fronts:
- Standards are an asset: Tone-of-voice guidelines and prohibited-word lists should be maintained as formal documents — they determine the floor of AI output quality.
- AI doesn't sign off: AI can generate and pre-review, but final sign-off on the launch version must land on a specific person. The chain of accountability cannot be broken.
- Testing has rules: Sample size, observation period, and decision criteria must be locked in up front, to prevent the old habit of "just pick whichever looks nicer" from coming back from the dead.
On tooling costs: mainstream LLM APIs today are pay-as-you-go, and the cost of batch-generating a set of copy is essentially negligible for a team (check the official pricing pages for specifics). The real cost has shifted to standards building and process design — which is precisely what managers should be doing.
If you want to get this workflow running too, you'll need a stable API gateway that lets you switch between multiple models. Check out this platform: https://api.thistoken.ai/register — after signing up, you can use the template above to kick off your first round of AI-powered copy A/B testing.
---
Every example in this post runs with a single API key — get yours at https://api.thistoken.ai/register and start in minutes.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key