A Solo Developer's Pain Point
I maintain a content website on my own, with modest daily traffic—around a few hundred thousand page views. There's no dedicated security team, no ops team; I do everything myself.
One of the biggest headaches is web scrapers. Some crawlers behave properly and respect robots.txt; but others disguise themselves as normal users, slowly hammering APIs, bulk-scraping detail pages, and siphoning off data. By the time I notice, it's usually already too late—slow query alerts from the database, abnormal bandwidth usage, or the entire site's content scraped away.
My old approach was pretty primitive:
- Manually grep the logs, roughly filtering by User-Agent;
- Use Excel or throwaway scripts to count per-IP request frequencies;
- Judge anomalies by gut feeling—e.g., "block any IP exceeding 100 requests per minute."
The problems with this workflow are obvious: slow, crude, and high false-positive rates. Once, I blocked an entire university computer lab's egress IP range, because behind that IP were hundreds of real student users. Another time, a low-frequency but extremely suspicious crawler—only 30 visits a day, exclusively targeting paid content—lurked for two weeks before I noticed it.
Each routine analysis took me 2.5–3 hours, and the quality of the conclusions depended entirely on how I was feeling that day.
A Different Approach: Throw the Logs at AI
Later, I came to a realization: crawler detection is fundamentally a pattern recognition + explanation problem—and that's exactly what large language models excel at. I don't need AI to process millions of lines of logs in real time—that's neither realistic nor economical. I just need it to replace me in the "analyze, summarize, and produce explainable conclusions" step.
My current workflow looks like this:
Step 1: Local preprocessing. A 50-line Python script aggregates the Nginx logs into a structured summary: total requests per IP, time distribution (e.g., whether activity concentrates between 3–5 AM), path entropy (whether the visited URLs are highly regular), User-Agent distribution, HTTP status code ratios, and request interval variance. The raw logs have millions of lines, but after aggregation the summary is usually only a few thousand tokens.
Step 2: Feed it to AI for analysis. Send the summary along with a carefully crafted prompt to the model, and have it output a structured report.
Step 3: Manual review + enforcement. For the list of suspicious IPs the AI provides, I spot-check the raw logs with another script to confirm, then decide whether to rate-limit, block, or add a CAPTCHA.
The key point: preprocessing happens locally, while AI handles judgment and explanation. This keeps costs under control and avoids sending out user privacy data contained in raw logs.
The Prompt Template I Actually Use
你是一名Web流量安全分析师。我会给你一份访问日志的聚合摘要数据,
请从中识别可能的异常爬虫行为。
数据字段说明:
- ip: 来源IP
- req_total: 请求总数
- active_hours: 请求活跃时段(0-23小时分布)
- path_entropy: 访问路径熵值(越低越规律)
- interval_cv: 请求间隔变异系数(越低越像脚本)
- ua_count: 不同User-Agent数量
- status_4xx_ratio: 4xx状态码占比
请按以下要求输出:
1. 列出可疑IP,按可疑程度排序,附风险等级(高/中/低);
2. 对每个可疑IP,用一句话说明判断依据(如"请求间隔变异系数极低
且集中在凌晨时段,符合脚本特征");
3. 区分"确定是爬虫"和"可能是正常用户但行为异常"两类,避免误伤
NAT出口、机房用户;
4. 给出对应的处置建议:封禁 / 限流 / 加验证码 / 仅观察;
5. 最后总结本次日志中爬虫流量的整体占比和主要行为模式。
数据如下:
{粘贴聚合后的摘要数据}Before and After: Not Just Faster
I ran this workflow for a full month and tracked the actual numbers (my own project, not client data):
| Dimension | Before AI | After AI |
|---|---|---|
| Time per routine analysis | ~3 hours | ~20 minutes (including review) |
| False positives on normal users | 3 per month | 0 per month |
| Speed of detecting stealthy crawlers | Over a week on average | Same day |
| Explainability of conclusions | Gut feeling, hard to articulate | Every judgment backed by evidence |
The third row surprised me the most. Previously I only watched high-frequency IPs, but the AI noticed things like "a certain IP traversed all category pages in alphabetical order over three days—low frequency but abnormally low path entropy." Detecting this kind of lurking crawler manually is nearly impossible. AI expanded the detection dimensions from "frequency" to "behavioral structure"—that's a qualitative leap.
As for cost, the aggregated data is roughly tens of thousands of tokens per run, and the monthly cost is negligible at my scale (check the official pricing page for exact prices—they vary greatly between models; I recommend starting with a cheap model to get the workflow running).
Pitfalls I've Hit
- Don't throw raw logs directly at the model. First, it wastes tokens; second, logs may contain user information—de-identification and aggregation must be done locally first.
- AI can be overconfident. It occasionally flags CDN origin-pull IPs as suspicious. So the manual review step is indispensable—AI narrows a hundred IPs down to five, and you make the final call.
- Explicitly stating "avoid false positives" in the prompt matters a lot. At first, my prompt only said "find the crawlers," and the model wanted to flag nearly all traffic as suspicious. After adding the constraint "distinguish abnormal behavior by normal users," accuracy improved significantly.
- Make it output its reasoning, not just conclusions. Only an explainable report is safe to act on for blocking operations.
Final Thoughts
For solo developers and small teams, the biggest cost of security has never been tools—it's people's time. "Analyzing logs to fight crawlers" used to be one of those important-but-forever-postponed tasks; now it's a fixed 20-minute routine I do every week.
Delegating this kind of repetitive analytical work to AI while humans make the final decisions—this division of labor can be replicated in almost any scenario with "lots of data, nobody looking at it," well beyond log analysis.
If you want to try this workflow, you'll need a stable, reliable model API to support it. Check out this platform: https://api.thistoken.ai/register—sign up and you can start feeding your first log summary to a model.
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key