How I Used AI to Cut a Legacy Code Refactor from 19 Person-Days to 8
1. The Kind of Project Every Independent Developer Fears
Taking over a legacy project with no documentation, no tests, and an original author who vanished long ago is one of the biggest headaches for independent developers and small teams. The code mixes five different naming conventions, database queries are inlined directly into page logic, and some core function is over seven hundred lines long that nobody dares to touch.
My situation was typical: an internal management system that had been running for six years, eight thousand-plus lines of PHP plus cron jobs that barely worked. The client wanted a migration to a new architecture, with the budget calculated in person-days. On my first read-through, I estimated that manually mapping all the call relationships, identifying risk points, and writing regression tests would take at least three weeks. The contract gave me ten days.
The pain points are very concrete:
- Afraid to touch anything: no test coverage, so changing one line could break something anywhere;
- Can't see through it: business logic and technical implementation are tangled together, so before making changes you first have to understand "what this code is actually doing";
- Can't explain it clearly: when explaining risks and workload to the client, you can only wing it based on experience, with no deliverable evidence.
These three things are exactly what AI is good at sharing the load for.
2. What AI Can Do for You in the Refactoring Process
My approach was not "let AI rewrite everything with one click" — that's even riskier. Instead, I broke the refactoring down into steps that AI can reliably complete, with humans only making judgments and acting as gatekeepers.
Step 1: Generate a code map. Feed the code to an LLM in chunks, and have it output a module inventory, function responsibilities, call relationships, and a dependency table. This step used to take two days of manual diagramming; now I get a first draft in half a day, then verify it by hand.
Step 2: Risk annotation. Have AI scan according to preset rules: SQL string concatenation, hardcoded secrets, uncaught exceptions, implicit type conversions, timezone handling, deprecated APIs. The output is a risk table with line numbers.
Step 3: Draw safety boundaries. This is the most critical step. Have AI classify the code into three categories: a "green zone" that can be safely refactored, a "yellow zone" that requires tests before touching, and a "red zone" that should only be wrapped and isolated — never rewritten. AI provides the classification suggestions, and I confirm each one.
Step 4: Generate verification checklists and regression test skeletons. For the yellow and green zones, have AI produce a before/after behavior comparison checklist and runnable draft test cases. AI-written tests will have omissions and naive assertions, but editing tests is much faster than writing them from scratch.
Step 5: Migrate in batches + comparative verification. Each time a module is migrated, have AI compare the behavioral differences between the old and new implementations, outputting a three-column comparison table of "input — old output — new output," and manually spot-check the critical paths.
3. A Reusable Prompt Template
This is the template I repeatedly refined for the risk annotation step, suitable for legacy code in any language:
你是一名资深代码审计工程师。我会提供一段遗留代码,请完成以下任务:
1. 【功能摘要】用不超过 5 句话说明这段代码的业务职责。
2. 【风险清单】逐条列出以下类别的风险点,标注行号和严重程度(高/中/低):
- SQL 拼接或注入风险
- 硬编码的密钥、IP、路径
- 未处理的异常和边界情况
- 依赖已废弃或不再维护的 API
- 隐式类型转换、时区、字符编码问题
- 副作用(写文件、发请求、改全局状态)
3. 【外部依赖】列出这段代码依赖的外部服务、数据表和全局变量。
4. 【重构分级】对每个函数给出建议:
绿区=逻辑清晰可安全重构;黄区=需先补测试;红区=建议只封装不重写。
并用一句话说明理由。
5. 【验证要点】如果重写这段代码,重构后必须验证哪些行为?列出可执行的检查项。
要求:不确定的地方明确标注"不确定",不要猜测。输出为 Markdown 表格。
代码如下:
<在这里粘贴代码>The last instruction — "say 'uncertain' when you're uncertain" — matters a lot. It significantly reduces the AI confidently fabricating call relationships.
4. Before and After: The Numbers Add Up Clearly
Estimated comparison of two approaches on the same project (based on my own actual work records, for reference only):
| Step | Pure Manual Estimate | Actual with AI | Savings |
|---|---|---|---|
| Code mapping and call relationships | 5 days | 1.5 days | 70% |
| Risk identification and classification | 3 days | 0.5 days | 83% |
| Writing regression tests | 4 days | 1.5 days | 62% |
| Refactoring implementation and verification | 5 days | 4 days | 20% |
| Documentation delivery | 2 days | 0.5 days | 75% |
| Total | ~19 person-days | ~8 person-days | ~58% |
A few observations:
- The biggest gains weren't in "writing code" but in "reading code" and "writing tests." The refactoring implementation itself only saved 20%, because core logic still needs human confirmation. This matches expectations: AI's value lies in lowering comprehension costs and covering blind spots, not making architecture decisions for you.
- Deliverable quality actually improved. The risk classification table and verification checklist I delivered let the client, for the first time, understand "where the money goes and where the risks are." That's not time saved — that's an additional billable deliverable.
- Cost is almost negligible. The total LLM API spend for the whole project is a tiny fraction of the ten-odd person-days saved (check the official pricing pages for specific model prices).
5. Lessons Learned the Hard Way
- Don't have AI output "the complete refactored code" and replace everything wholesale. Small batches, comparable output, and easy rollback — that's the safe pace.
- When in doubt, leave the red zone alone. For those seven-hundred-line functions, wrapping them in an adapter is enough; the risk of rewriting far outweighs the benefit.
- AI's risk list is a first-pass filter, not a conclusion. On average I remove about 15% false positives and add about 10% missed points on top of its output — but both steps are far faster than starting from scratch.
- Write the verification checklist into the contract. Changing the definition of "refactoring complete" from "the code runs" to "every item on the checklist passes" protects both parties.
Final Thoughts
Legacy code refactoring has never been a technical problem — it's a problem of risk and trust. AI can't undo the historical baggage of the code itself, but it can turn "assessment by gut feeling" into a controlled process with checklists, classifications, and verification — and for developers taking on freelance work, that's both time saved and the confidence to take on such projects.
If you're also doing code audits, refactoring assistance, or bulk documentation generation, you can try running this workflow through a stable model API. Registration here: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Vous voulez essayer Token.AI ?
Créez une API Key au niveau du projet, activez les canaux dans la console et configurez le routage, les budgets et les journaux d'audit.
注册 ThisToken.AI 并获取 API Key