Continues from the previous two: Empirical Breakdown, Pipeline Design. This post combines three things: a data-backed deep dive into对标 accounts, an analysis of the Chinese-language ecosystem niche, and an important course correction—the originally planned route of “reverse-engineering map data from L1 video frames” was scrapped after research, with reasons detailed in the second half.
I. English-Language Benchmarks: Five Channels, Five Models
| Channel | Update Cadence | Typical Views/Video | Monetization Core | Replicable Aspects |
|---|---|---|---|---|
| Green Dot Aviation | 0–2/month, 3–8 week gaps | 800K–2.5M; viral hits 3–8M+ | Ad revenue + VPN/education sponsors | Fault-chain narrative template |
| Mentour Pilot | ~1/week | 300K–1.5M | Ads + diversified sponsors + courses | “Accident recap → systems education → industry impact” closed loop |
| Kings and Generals | 2–4/week | 100K–600K; viral hits 3M+ | Industrialized ads + game sponsors + Patreon | Series-based playlists |
| Montemayor | 0–3/year | 1–4M;代表作 5–15M+ | Ads primary, sponsors restrained | Multi-perspective concurrent deduction |
| Battle Order | 1–3/month | 80K–400K | Ads + guidebooks/merch | Strong template with horizontal expansion |
Laying these five rows side by side reveals one pattern: update frequency is inversely proportional to per-video views. Montemayor does three videos a year, each hitting millions; Kings and Generals does four a week and each video only gets hundreds of thousands. Total traffic might be similar, but the asset profiles are completely different—low-frequency, high-quality channels produce “long-tail assets” (military enthusiasts rewatch repeatedly; a single Midway series can drive recommendation traffic for five years), while high-frequency series channels are “traffic machines” (playlist binge-watching turns one viewing session into three to five).
What this means for you: The natural form of an automated pipeline is high-frequency series production (machines don’t tire), so benchmark K&G’s structure + Green Dot’s per-video quality standard. Don’t benchmark Montemayor—its precision is built entirely through manual labor, precisely the part hardest to automate. Learning from it would be shooting yourself in the foot.
What to Actually Steal from Each Channel
Green Dot Aviation’s fault-chain template. Fixed structure per episode: flight background → crew/aircraft status → critical events → cockpit decisions → accident chain → investigation findings → safety improvements. The value for automation is that this is本质上 a causal graph—a model can extract it directly from investigation reports (detailed in M3 below). You don’t need a “storytelling genius”; you need a “causal-chain extraction model.”
Mentour Pilot’s content closed loop. Accident recap → systems education (CRM, TCAS, stalls) → industry impact—the three segments cross-promote each other. For your pipeline: the same event JSON can render three video types (accident film / systems explainer / impact retrospective), one extraction, three monetizations.
Kings and Generals’ playlist engineering. “Pacific War Week by Week” progresses week by week, breaking a war into bingeable episodes. Chinese military channels are almost entirely single-video scattered hits; playlist-level series are a structural gap. For the pipeline: the topic module supports this natively—build series queues by war/accident type, auto-append episode numbers.
Battle Order’s horizontal scalability. Each episode answers “how many people, what equipment” for a platoon/company/division, with fixed icons + hierarchy diagrams. Switch country and era, and you’ve got infinite new content. This is the极致 of “template-driven topics”—closest to full automation, but also the lowest ceiling (capped at a few hundred thousand views).
II. Chinese-Language Niche: Not Empty—Just a Supply-Quality Gap
This track in the Chinese-language space isn’t blank; it’s fragmented. The real distribution on Bilibili:
- Broad knowledge giants: Xiao John Khan (~8–10M), Banfo Xianren (~6–9M), X Investigations (~3–5M)—they occasionally do disaster/accident episodes as viral hits, but they’re not vertical specialists
- Mid-tier map-history channels: Sandbox Wars (~1–3M), various “Map-Based Wars” accounts (~100K–1M)—they have map animations but lag English-language leaders by one to two quality tiers
- Aviation-accident verticals: Most are stuck at 10K–300K followers—搬运-type, AI-narration-type, or research-compilation-type; virtually none produce original animations starting from L0 reports
Where’s the gap? Million-follower broad-knowledge channels occasionally cover aviation accidents (not vertical, not sustainable); vertical aviation-accident channels can’t grow (quality insufficient). The “vertical + high quality” niche in the middle is empty. This is the specific shape of the L1→L2 gap described in Part Two.
Common Patterns in Growth Trajectories (Xiao John Khan, Sai Lei Hua Jin, X Investigations, etc.)
Breaking down how top knowledge-narration channels grew, the common thread isn’t “good content”—it’s four system capabilities:
- Series-first approach: Weird Small Countries, Hardcore Ruffians, Comic Economics—audiences remember the series promise, not individual episodes. No direction changes in the first 30 episodes.
- Three-phase rhythm: Cold start with 2–5 videos/week to test the model (5–10 min) → growth phase with 1–3/week building series (8–15 min) → maturity phase with weekly premium pieces (15–30 min)
- Titles that create knowledge gaps: Not “Flight XXX Crash Analysis,” but “A Normally Flying Plane—Why Did It Suddenly Crash?”
- Multi-platform reuse: Bilibili long-form builds trust, Douyin clips expand reach, WeChat articles capture conversions—one topic, four monetizations.
This is structurally favorable for an automated pipeline: series = topic queue; three-phase rhythm = output parameters; knowledge-gap titles = LLM-generable template sentences; multi-platform reuse = one render, multiple format exports. System-level work maintained by human operators at top channels is precisely the structured output that pipelines handle best.
III. The Real Money
YouTube Side
| Audience/Categy | RPM Range | Notes |
|---|---|---|
| English aviation/engineering disasters | $2–8 | Framed as “engineering failure” is far more ad-friendly than “horrific disaster” |
| Chinese broad-audience traffic | $0.3–2 | Heavier disaster framing pushes it lower |
| Chinese-speaking North American audience | $2–7 | Best audience quality within Chinese-language content |
The hidden cost of disaster content is yellow-label risk: death details and sensational thumbnails trigger ad limitation. Telling the same crash story, “The Engineering Failure Behind Flight XXX” versus “300 People Died in the Most Horrific Crash” can yield a 3× RPM difference. This is why the pipeline’s copy module must include a built-in “engineering narrative vs. sensational narrative” style toggle.
Douyin Side
Xingtu pricing (knowledge/disaster narration tier): 100K–300K followers, ¥2K–15K/video; 500K–1M, ¥10K–40K; 1M–3M, ¥30K–150K. But disaster/military categories face 30–50% price suppression across the board—brand-safety discount. Best-fit categories: SLG games (easiest), books/courses (secondary), FMCG/beauty (essentially incompatible).
Unit Economics
Per-video pipeline cost: <$5 (VLM + ASR + LLM + TTS). Compared to revenue: mid-tier Chinese YouTube videos earn $100–300 each; Douyin Xingtu at the 300K-follower tier pays ~¥10K/video. Near-zero marginal cost capacity facing a monetization curve with a clear price floor—that’s the math behind this business. The break-even point isn’t about views; it’s about account-weight climb speed.
IV. Course Correction: M3’s Map Redo—My Original Plan Was Wrong
The route described in Part Two was “reverse-engineer event data (time/location/unit/trajectory) from L1 video frames, then programmatically reconstruct animations.” After deep research, this route is downgraded to an auxiliary tool; the primary route shifts to L0 direct ingestion. Reasons:
The Precision Wall of Frame Reverse-Engineering
Extracting event data from video frames has feasibility that depends heavily on visual type:
- Frames with latitude/longitude grids, scale bars, clear place names: positional accuracy to within tens to hundreds of meters
- Frames with schematic/illustrative maps only: errors at the kilometer scale
- Time with on-screen timecodes: second-level; inferred from narration only: minute-level
- Unit identification is affected by icon overlap, compression artifacts, and occlusion, requiring manual confirmation per unit
More critically, the manual verification load: for a 10-minute investigative video, after OCR + tracking runs, 30–70% of key events still require manual review; if the source is a news derivative animation rather than a GIS export, “review” is roughly equivalent to rebuilding from scratch. Frame reverse-engineering digs imprecise data out of pixels and then calibrates it—inverting the structure by feeding low-quality sources into a high-quality production pipeline.
L0 Direct Ingestion: Investigation Reports Are Already Structured Data
NTSB/AAIB reports have far higher extractable field density than video frames: accident numbers, UTC timestamps, flight phases, lat/lon coordinates, MSL/AGL altitudes, airspeed/vertical speed/heading, FDR parameters, ATC call timelines, CVR excerpts, alert-system logs. These are already tables and timelines within the reports. PDF parsing → entity normalization → event schema ({t, lat, lon, alt, speed, heading, actor, event_type, source_page}) → render, with every data point traceable to a page number.
New primary route:
| |
Frame reverse-engineering no longer serves as the data source; instead it does two things: visual alignment (comparing your animation’s visual language against top English channels to maintain genre familiarity) and gap filling (supplementing details the L0 report doesn’t cover, drawn from L1 video). Its role shifts from “mine” to “quality inspector.”
Render Tool Selection (Under This Route)
- Mapbox GL JS: Route/battlefield GeoJSON trajectories + time slider—the workhorse for data-driven rendering
- Motion Canvas: Code-based 2D animation—arrows, labels, legends, tactical diagrams; renders event JSON into explainer visuals
- Google Earth Studio: 3D terrain fly-through shots (openers/atmospheric segments)—weak programmatic batch control, used only as embellishment
- Composition logic: Mapbox handles the geographic layer, Motion Canvas handles the information layer, Earth Studio handles the cinematic layer—each plays to its strength
Revised Milestones
- M1 (Week 1): Windmill monitoring online, observation-only
- M2 (Weeks 2–3): Manual production of 3–5 sample videos—but with source intake changed to NTSB PDF direct ingestion + L1 video assistance; calibrate field completeness rates of the extraction schema
- M3 (Month 2): Mapbox + Motion Canvas render pipeline + direct CapCut draft writing
- M4: View data feeds back into the topic model
V. What This Correction Changes
Part Two’s instinct—that map redo should come from frame reverse-engineering—was right in direction but wrong in engineering: it placed the most unreliable data source (video frames) at the most critical production step. This post’s correction elevates L0 from “reference material” to “primary data source”: the tables and timelines in investigation reports are naturally structured, and LLM extraction from them is an order of magnitude more reliable than guessing from pixels. This also rewrites the entire pipeline’s copyright positioning—source material is publicly available government reports (no copyright issues), L1 video serves only as visual reference, and the legal gray area of translation/adaptation doesn’t come into play under the primary route.
The next post covers M1 implementation: Windmill script code, two weeks of monitoring data, and the unvarnished answer to “how many qualifying topics per week.”
