## The Business Pain Point: The Most Labor-Intensive "Inv...
The Business Pain Point: The Most Labor-Intensive "Invisible Process" in Knowledge Commerce Teams
Anyone in the video course business knows that between "recording finished" and "course published" lies an extremely tedious process:
- Chapter splitting: The instructor records three hours in a single take. Operations staff have to scrub back and forth through the timeline to find "this is where environment setup starts" and "this is where the first case study starts," manually marking split points.
- Chapter naming and summaries: After splitting, each chapter still needs a title and description for student navigation and SEO indexing.
- Subtitle summarization: Full transcripts often run 40,000–50,000 characters and need to be condensed into course highlights, catalog page copy, and preview chapter summaries.
Take a 3-hour course with roughly 40,000 characters of subtitles: a skilled operator needs 5–6 hours in practice to split chapters and write summaries. For a small knowledge commerce team publishing 10 courses per month, that's 50–60 hours of repetitive labor every month—roughly a week and a half of one person's full working hours. Worse, manual splitting produces inconsistent granularity: different operators produce anywhere from 8 to 20 chapters, creating a fragmented user experience.
Our goal was clear: compress per-course processing time from 6 hours to under 10 minutes, with consistent splitting granularity.
Architecture Design: A Three-Stage Pipeline
The overall architecture has three layers and can be built by an indie developer in two or three days:
[视频文件]
│
▼
① ASR层:视频 → 音频提取 → Whisper类模型转写 → 带时间戳的SRT字幕
│
▼
② 切分层:字幕按时序分块喂给LLM → 输出章节边界时间点 + 章节标题
│
▼
③ 摘要层:按章节切分文本 → LLM生成每章摘要 → 汇总课程亮点
│
▼
[结构化JSON:章节列表 + 时间戳 + 标题 + 摘要]Several key design decisions:
Use an open-source model for ASR, an API model for summarization. Transcription is compute-heavy but follows a fixed pattern—run Whisper locally or use a cheap transcription service. Chapter detection and summarization, on the other hand, require language understanding, making LLM APIs the better value. Don't use large models for the entire pipeline—that multiplies your costs several times over.
Use a sliding window for the splitting stage rather than feeding everything at once. A 40,000-character transcript exceeds most models' context windows. The approach is to feed roughly 15 minutes of subtitle segments at a time, have the model output chapter boundaries within each segment, and keep a 2-minute overlap between windows to avoid missing splits.
Route everything through a unified AI API gateway instead of connecting directly to each provider. This is the key to keeping maintenance costs minimal—more on this below.
Key Implementation Steps and Code
The core chapter-splitting prompt and invocation logic are as follows (Python sketch):
CHAPTER_PROMPT = """你是一位课程编辑。以下是带时间戳的字幕片段。
请判断片段内的话题切换点,输出JSON数组,每项包含:
- start: 章节开始时间(秒)
- title: 8-14字的章节标题
要求:只有话题真正切换时才切章,目标章节时长5-15分钟。
字幕片段:
{transcript}
"""
def split_chapters(srt_blocks, llm_client):
results = []
for window in sliding_window(srt_blocks, size=15*60, overlap=120):
resp = llm_client.chat(
model="gpt-4o-mini",
messages=[{"role": "user",
"content": CHAPTER_PROMPT.format(
transcript=format_srt(window))}],
response_format={"type": "json_object"}
)
results += json.loads(resp)["chapters"]
return dedupe_and_merge(results)Implementation checklist:
- Use
ffmpegto extract audio, then run Whisper to get timestamped SRT subtitles (takes roughly 0.3× the video duration). - Call the LLM with a sliding window for chapter splitting; deduplicate and merge results across windows.
- Split the text by chapter and generate a 50–80 character summary for each.
- Make one additional call to generate "course highlights" (3–5 items) and catalog page copy.
- Humans only do final proofreading: spot-check whether chapter boundaries are reasonable—about 15 minutes on average.
- Write the resulting JSON into your CMS to auto-generate the course catalog page.
Efficiency comparison: For a single 3-hour course, the old process took 5–6 hours → now about 8 minutes of machine processing + 15 minutes of human proofreading, a reduction of over 90%. At 10 courses per month, that saves roughly 600 hours of labor per year.
Why a Unified AI API Gateway Reduces Maintenance Costs
This solution requires calling at least two types of models (ASR can be gatewayed too; for LLMs you want at least one primary plus one backup). Indie developers' biggest fears are:
- Maintaining multiple vendor SDKs: OpenAI, Anthropic, and DeepSeek each have their own SDK, authentication, and error-handling logic, and version upgrades frequently conflict with each other.
- Models change constantly: Model A may be the best value today, but you might switch to Model B three months later. If the vendor is hardcoded, every model change means code changes and retesting.
- Scattered key management: Keys for different services live in different environment variables, making rotation and auditing a pain.
Once everything goes through a single OpenAI-compatible API gateway, your code has only one endpoint and one SDK. Switching models means changing only the model name parameter; if a provider goes down, you can switch to a backup route without touching code; and keys are managed centrally on the gateway side. For an indie developer maintaining multiple AI pipelines alone, this doesn't just save one-time development effort—it eliminates that half-day of refactoring work that recurs with every model iteration. Over a year, that easily adds up to several working days saved.
Lessons from Real-World Testing
- Always specify the target chapter duration in the splitting prompt, or the model will split too finely (our first version averaged 20+ chapters per course).
- For narration-heavy content full of filler words, run a subtitle-cleaning pass before feeding the model—it cuts token consumption by about 30%.
- A lightweight model is sufficient for summarization; only outputs like chapter titles that need to be "both accurate and catchy" are worth using a stronger model for.
If you're also planning to build an AI processing pipeline for video content, start with a unified access layer: after registering at https://api.thistoken.ai/register, you can call models from multiple providers through a single OpenAI-compatible interface. Get the pipeline running first, then gradually optimize your prompts and cost structure.
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key