A Real Dilemma
At the end of last year, I took over an old project that had been maintained for three years—40,000 lines of code with less than 8% test coverage. Opening the test directory revealed only a handful of files, with the last commit made a year and a half ago.
The task in front of me was clear: the upcoming large-scale refactoring had to have tests as a safety net. But if I wrote tests the traditional way, I estimated the workload—there were over 300 functions in the project, each taking an average of 20 to 40 minutes to test manually (including reading code, understanding logic, writing test cases, and running verification). Completing everything would take over 200 hours. For a solo developer or a three-person team, that meant two to three months without touching any new features.
This is the classic dilemma of legacy projects:
- Time black hole: Writing tests is important but not urgent, so it never makes it onto the schedule
- Heavy mental burden: Reading code written by others (or even yourself three years ago) is exhausting
- Outsourcing isn't cost-effective: Outsourcers can't understand the business context, so the tests they write are either trivial or don't pass
- Inverted risk: Without tests, refactoring is like walking a tightrope without a net
AI Changed the Math
My approach wasn't to have AI "fully automatically" generate tests, but to let AI handle the most time-consuming part: generating test skeletons, leaving human effort for the parts that truly require judgment.
Step 1: Batch-Extract the Function Inventory
First, I wrote a script using AST parsing to extract the signatures, parameters, and return types of all functions in the project into a list. Ten minutes of script work gave me a clear view of 300+ to-do items.
Step 2: Feed to AI in Batches
I grouped the functions by module and gave the AI the code for one module at a time, asking it to generate test skeletons. Note: skeletons—including test file structure, describe/it organization, edge case checklists, and mock suggestions, not complete tests claiming to run out of the box.
Step 3: Human Review and Completion
The AI-generated skeletons averaged 60% to 70% completion: the test structure was reasonable and the edge case coverage approach was correct, but the specific assertion values and business expectations often needed manual calibration. All I had to do was fill in the flesh on the skeleton and verify the business logic.
Measured Time Comparison
| Item | Pure Manual | AI-Assisted |
|---|---|---|
| Average time per function | 20-40 min | 6-10 min |
| Total for 300 functions | ~200 hours | ~45 hours |
| Sustainable daily output | 12-15 functions | 40-50 functions |
In the end, I used two weeks of fragmented time to complete work originally estimated at three months, raising test coverage from 8% to 63%. The 150+ hours saved is equivalent to a month of development work for a small team. Calculated by API usage, the model invocation cost for the entire task was roughly equivalent to a few cups of coffee (check the official pricing page for specifics)—negligible compared to the time saved.
The Prompt Template I Used
A directly copyable version, adaptable to most languages:
你是一位资深测试工程师。我会给你一段项目代码,
请为其中指定的函数生成单元测试骨架。
要求:
1. 使用 {测试框架}(如 pytest / JUnit / Jest)
2. 每个函数生成测试文件结构、describe/分组组织
3. 列出正常路径、边界值、异常输入三类用例的清单,
每个用例写清「输入 → 期望行为」,断言处用 TODO 标注
4. 识别函数的外部依赖(数据库、网络、时间),
给出 mock/patch 建议,但不要过度 mock
5. 不确定的业务逻辑,明确标注「需人工确认」,
不要编造期望值
6. 输出格式为可直接保存的代码文件
以下是待测代码:
{粘贴代码}A few key takeaways:
- Ask for skeletons, not finished products. Letting AI fabricate assertion values is the biggest trap—expected values must be confirmed by someone who knows the business
- Marking uncertainty is more valuable than pretending to be certain
- Batching by module works better than dumping everything on the AI at once; with too much context, quality degrades noticeably
- A cheap model is perfectly adequate for skeleton generation; save the expensive models for complex logic analysis
Before vs. After: More Than Just Speed
The problem with the pure manual approach isn't just slowness—it's also activation resistance. The mere thought of reading tens of thousands of lines of old code makes you instinctively push it to "next quarter." With AI assistance, the activation cost drops to "copy-paste a chunk of code," transforming the task from a "huge project" into "a quick task" in your mental accounting.
More importantly, the quality floor has been raised. The edge case checklists generated by AI often reminded me of cases I'd missed: empty arrays, negative numbers, concurrent calls, timezone boundaries. When writing tests manually, these corner cases are exactly what people skip when tired.
Of course, don't deify it either: AI doesn't understand your business history or know that a seemingly bizarre if branch is actually a patch for a production incident from 2019. The human review step cannot be skipped. I position AI as a tireless junior test engineer—it handles the volume, and you do the quality control.
Conclusion
For writing tests in legacy projects, the answer used to always be "no time"; now the answer is "it can be done in two weeks of fragmented time." For independent developers and small teams, this kind of highly repetitive, clearly patterned work is exactly where AI delivers the most value—it doesn't think for you, but it burns the hours nobody wants to burn.
If you'd like to try handling this kind of task in bulk via API, you can register an account here to get started: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key