## The Pain Point: Those "It Works, Don't Touch It" Legac...
The Pain Point: Those "It Works, Don't Touch It" Legacy Scripts
Any developer who has taken over a five-year-old project has probably seen this kind of file: a deploy.sh with two hundred-plus lines, with a sixty-character regex embedded in the middle, used to extract version numbers from logs. The person who wrote it has long since left, the script runs on a weekly schedule, and nobody dares to touch a single line—because no one is sure what will break if they do.
Our team's situation was even more typical:
- The repository contained 30+ shell scripts with regexes scattered all over, the vast majority with zero comments;
- Every time we troubleshot a failed cron job, it took 40 minutes to an hour on average to read through the script, dig through Git history, and guess what the variables meant;
- Onboarding new members relied entirely on verbal hand-offs, which was costly, and the process itself drained senior members' time.
A rough calculation: at 6 troubleshooting sessions per month, 50 minutes each, that's 5 hours of pure overhead per month—60 hours a year—and that doesn't even count onboarding and the time new members spend figuring things out on their own. Hiring someone to rewrite these scripts would be riskier and less cost-effective.
Writing comments was the only low-risk, high-reward option. But it was always "important but not urgent," and never made it onto the schedule.
The AI Workflow: Three Steps to Make Legacy Scripts Readable
My approach is simple—three steps, all using a single prompt template with a consistent format.
Step 1: Batch inventory. Use a simple script to scan the repository and list all .sh files and the regexes in the code (patterns in grep/sed/awk). No AI involved in this step—just file traversal.
Step 2: Feed each item to AI to generate comments. Send each script or regex, along with the necessary context (which file calls it, sample inputs and outputs), to a large language model, requiring output in a fixed format: section-by-section comments + a breakdown of the regex + flagged risk points. The key is to require the AI to only explain, not modify the code—comments are documentation, not refactoring.
Step 3: Manual spot-check and merge. Manually verify the accuracy of the comments for the 5 most critical scripts, and merge the rest through the normal code review process. AI comments occasionally "hallucinate" intent—for example, explaining a buggy-but-luckily-working regex in a very plausible way—so a human safety net is essential.
Here is the prompt template, ready to copy and adapt:
你是一名资深 Shell 与正则表达式专家。请为下面这段存量代码写注释,严格遵守以下规则:
1. 只添加注释,不修改任何逻辑,即使你认为有 bug,也只注释不改动;
2. Shell 脚本:按逻辑分段,每段上方写 2-3 行中文注释,说明"做什么"和"为什么";
3. 正则表达式:逐段拆解每个部分的作用,用如下格式:
正则:xxx
- ^\d+ :匹配开头的连续数字,对应日志中的进程号
- \.(tar|gz) :匹配压缩包后缀
整体作用:一句话总结;
4. 对有风险或难以理解的地方,单独加"⚠️ 注意:"前缀标注;
5. 如果代码行为依赖特定环境(如 bash 版本、文件路径),明确指出;
6. 输出格式:先给完整带注释的代码,再给一段 100 字以内的整体说明。
代码如下:
---
(粘贴脚本或正则,附上输入输出示例)
---Before and After: The Numbers
Using our legacy assets as a sample, the comparison looks roughly like this:
| Task | Manual Only | AI-Assisted |
|---|---|---|
| Commenting a single 100-line script | ~60-90 minutes (including recalling context) | ~10 minutes (AI generation + human verification) |
| Breaking down a complex regex | 15-25 minutes | 2-3 minutes |
| Full coverage of 30 scripts | ~35-40 hours; realistically never fit into the schedule | Done in 3 days of scattered time |
| Troubleshooting cron job failures | 50 minutes on average | ~15-20 minutes once comments are in place |
The two most noticeable changes: First, what used to be "never on the schedule" became something you could knock out in your spare time—AI generates the first draft, humans only verify, lowering the barrier from "writing" to "reviewing." Second, onboarding time for new members dropped significantly. Deployment scripts that used to require an entire afternoon of verbal explanation can now be understood by reading the comments alone, and senior members went from "lecturers" to "Q&A support."
As for cost, the API fees for commenting this entire batch of scripts came out to less than one hour of the team's labor cost (see the official pricing page for exact billing details). Compared to the 30+ hours saved, this pays for itself no matter how you calculate it.
A Few Lessons Learned
- Don't let AI modify the code. During the commenting phase, AI should only explain. It occasionally explains a bug as a feature; keep the "⚠️ Note" markers in the comments and leave the fix decisions for later.
- Providing context matters more than prompt engineering tricks. Attaching a real sample of input and output dramatically improves the accuracy of regex comments.
- Add tests while the comments are fresh. With comments in place, it becomes much easier to add minimal smoke tests for the critical scripts—only then can you safely talk about refactoring.
Legacy code debt doesn't disappear on its own, but the cost of paying it down can be compressed to a "just do it in passing" level. If you want to get this workflow running, you can start by registering for an API service and trying the template above: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Bạn muốn thử Token.AI?
Tạo API Key cấp dự án, bật kênh trong bảng điều khiển và định cấu hình định tuyến, ngân sách và nhật ký kiểm tra.
注册 ThisToken.AI 并获取 API Key