I. A Regex That Almost Caused a Production Incident
Our team maintains a legacy project that contains hundreds of regular expressions: phone number validation, ID card numbers, emails, URL rewrite rules, log field extraction... Many of them were written years ago by a long-departed colleague — a single line of over two hundred characters, zero comments, zero test cases.
What managers fear most isn't hard technical problems, but "code nobody dares to touch." Once, operations reported that a certain category of phone numbers was failing validation. I asked around in the group chat, and nobody could confirm whether changing the boundary logic of that regex would affect other channels. In the end, a senior colleague had to spend an entire day manually breaking it down and verifying it piece by piece before we dared to commit a two-character change.
This is not an isolated case. Regular expressions are a classic case of "a joy to write, a nightmare to read":
- Knowledge gap: The person who writes it and the person who reads it are rarely the same, so knowledge never gets institutionalized;
- Opaque risk: Changing a single character might affect hidden groups or assertions — you can't see it with the naked eye;
- Missing documentation: Test cases are incomplete, and regression testing is done by gut feeling.
So I established a rule: any regex change must first go through an AI analysis step, producing a structured explanation and boundary test cases, and pass review before the code can be modified. This article shares this workflow and the changes it brought to the team.
II. What AI Can Do Here
Let's be clear about the positioning: AI is not an automated machine that fixes regexes for you — it's about turning implicit knowledge that "lives in one person's head" into explicit, team-shared assets. Specifically, it can do four things:
- Structural breakdown: Explains a long regex section by section — groups, assertions, quantifiers — annotating the intent of each part;
- Boundary case generation: Automatically generates example strings that should and should not match, covering typical boundaries (empty strings, special characters, overly long input);
- Risk flagging: Points out potential hazards like backtracking traps, greedy matching ambiguity, and character set omissions;
- Reverse generation: You describe the requirement, it produces candidate regexes with verification cases for human confirmation.
The fourth point deserves special emphasis: reverse generation must always be paired with human confirmation. Regex is precise down to the character level — AI-generated candidates are drafts, not final versions. This is also part of the risk control built into the workflow.
III. The Team Workflow: Four Steps
I formalized this process into four steps and wrote it into the team wiki:
Step 1: Submit an analysis request. Anyone who receives a regex-related task pastes the original expression plus context (which language, where it runs) to the AI, using a unified template (see below).
Step 2: AI produces an analysis card. It contains: a section-by-section explanation, a description of the matching goal, several positive and negative test cases, and risk warnings.
Step 3: Human review of test cases. This is the critical risk-control gate. The reviewer must weed out inaccurate cases and add business-specific boundaries (for example, whether our system should support the +86 prefix for phone numbers — AI doesn't know that, but the business side does).
Step 4: Archive. The analysis card and its test cases are stored in the code repository alongside the regex file. The next time someone needs to touch that regex, they read the card first, then the code.
We've been running this workflow for over a month. The most visible change: in review meetings, discussing regex went from "staring blankly at characters" to "checking off test cases one by one" — the discussion now has a shared language.
IV. A Reusable Prompt Template
Here's the template I use — take it and adapt it directly:
你是一名资深正则表达式工程师。请解析以下正则表达式,输出结构化报告。
【正则表达式】
{{regex_here}}
【运行环境】
语言/引擎:{{如 JavaScript / Python re / PCRE}}
使用场景:{{如接口入参校验 / 日志字段提取 / URL 重写}}
【请输出】
1. 逐段拆解:按分组、字符类、量词、断言分段,说明每段作用
2. 功能概述:用一句话说明这个正则的整体匹配目标
3. 匹配示例:至少 5 个应命中的字符串,覆盖典型与边界情况
4. 不匹配示例:至少 5 个不应命中的字符串,并说明原因
5. 风险提示:指出回溯风险、贪婪/懒惰歧义、字符集遗漏等潜在问题
6. 修改建议:如有更清晰的等价写法或隐患修复方案,请列出
【约束】
- 示例必须是可直接用于单元测试的完整字符串
- 不确定的地方明确标注"待人工确认",不要猜测The "pending human confirmation" constraint was added after I stepped on that landmine — in the early version, the AI would confidently fill in business rules, making test cases look correct but fail in actual execution. After adding this line, the AI proactively leaves blanks where it's unsure, handing judgment back to humans.
V. Before and After AI
| Dimension | Before | After |
|---|---|---|
| Understanding an unfamiliar regex | Senior colleague, ~half a day | AI analysis + human review, ~half an hour |
| Regex change review | Arguing from experience | Verifying case by case against test cases |
| Knowledge retention | In individual brains | Analysis cards stored with the code |
| Onboarding new members | Afraid to touch the regex module | Complete independently per workflow, with review as the gate |
| Regression risk | Testing your luck in the test environment | Run positive and negative cases first |
One detail worth mentioning: "whoever wrote it maintains it" used to be the default rule, turning regexes into some people's private territory. Now that analysis cards are stored in the repo, any team member can take over — staff turnover is no longer a risk point. For managers, this may be worth more than the hours saved — the process reduces the organization's dependence on specific individuals.
VI. A Few Reminders
- AI's analysis can also be wrong, especially with obscure syntax and engine differences — the human review step cannot be skipped;
- Don't let AI directly modify production regexes — its role is to "write the manual" and "write the exam questions"; whether and how to change things is decided by humans;
- If you use a commercial model service, costs are per the official pricing page. At our team's scale, it's negligible compared to labor costs, but I recommend logging usage for accounting purposes.
Final Thoughts
Complex regexes used to be the darkest corner of our codebase. The AI analysis workflow has turned them into a "controlled zone" with documentation, test cases, and review. For small teams, this kind of transformation requires no infrastructure investment — one prompt template plus one process rule is enough to get started.
If you want to systematically identify this kind of "analysis work AI can take over" in your team, you can try connecting to a unified model service via https://api.thistoken.ai/register, turn the workflow into tooling, and make the process truly stick.
---
Every example in this post runs with a single API key — get yours at https://api.thistoken.ai/register and start in minutes.
Bạn muốn thử Token.AI?
Tạo API Key cấp dự án, bật kênh trong bảng điều khiển và định cấu hình định tuyến, ngân sách và nhật ký kiểm tra.
注册 ThisToken.AI 并获取 API Key