Featured image of post My AI Model Arsenal Inventory: Multi-Chart Weighted Ranking Across Video/Image/Voice Lines (2026-09)

My AI Model Arsenal Inventory: Multi-Chart Weighted Ranking Across Video/Image/Voice Lines (2026-09)

A full inventory of every video, image, and TTS model actually callable on my production line: 15 DashScope keys tested individually against 7 video endpoints each; on the video line, Wan3.0-video at Arena 1481 (Elo #3) outperforms the pipeline's original Wan2.7 (+53 points); on the image line, GPT-Image-2.5 Sunburst 1421 tops the Arena text-to-image chart—coincidentally the current default; on the TTS line, recommend switching from CosyVoice to Qwen3-TTS-Flash (97ms first packet, voice cloning/design, 9+ dialects). Data cross-referenced across LMArena, Arena.ai weekly, Artificial Analysis, and SuperCLUE; single-engine sources are all flagged with a discount.

TL;DR: Two of Three Category Leaders Are Already in Your Hands

This week, while wrapping up the漫剧 pipeline, I ran a full inventory: every video, image, and speech model I can actually call on my production lines, weighted and ranked against September 2026 blind-test leaderboards. The audit wasn’t “what models exist on the market” — it was “what’s actually wired into my Lumenx pipeline + CPA gateway and ready to receive requests today.”

Of the three category leaders, two are already running in production: video line Wan3.0-video (Arena I2V Elo 1481, #3 overall, +53 over the pipeline’s original Wan2.7), image line GPT-Image-2.5 Sunburst (Arena T2I 1421, #1). The only action item is the speech line — recommend switching the main voice from CosyVoice to Qwen3-TTS-Flash.

One side finding deserves a callout: the previous day’s verdict was “DashScope video keys are all dead.” This time I tested all 15 keys in CPA one by one against the video-synthesis endpoint — 7 passed cleanly.所谓 “通道死了” is often just “I didn’t test them all.”

Why I Did This: A “Keys Are Dead” Misjudgment

Timeline: On the evening of Sept 20, during a Lumenx production test, all three video channels (DashScope/Vidu/Kling) errored out. That night’s conclusion, written into memory: “waiting for user to fix keys.” Two days later, I fired a minimal request at each of the 15 keys in the CPA config (1-second duration, X-DashScope-Async: enable) — 7 returned 200 + task_id immediately. The one I’d tested before happened to be the broken one.

Lesson learned: channel-level verdicts require exhaustive testing — sampling lies. So the first order of business this time wasn’t checking leaderboards; it was getting a clear picture of “what do I actually have” — the model directory in my toolchain (Lumenx’s model_catalog.json, 38 models across 8 families), the channel list in the gateway (CPA config.yaml, 90+ model names across 16 channels), the audio registry (CosyVoice/Qwen3 dual-family voices in tts.py), plus the per-key live tests. Only after the inventory was complete did ranking come into play.

Method: Multi-Source Cross-Referencing, Single-Engine Sources Get Discounted

Ranking rules, stated upfront, no pretense of objectivity:

  • Primary weight to LMArena / Arena.ai blind-test Elo — millions of human blind votes, closer to “real aesthetic preference” than any technical metric;
  • Secondary weight to Artificial Analysis Video Arena, SuperCLUE Chinese boards, and vendor six-dimensional reviews, for cross-validation and gap-filling;
  • Data gaps honestly demoted: TTS has no mature Arena board; that section’s ranking is a features + reputation weighted judgment, explicitly flagged in the text;
  • On the retrieval day, the Tavily engine had SSL issues throughout; data was cross-checked using Grok + 秘塔 dual engines, with every key number requiring at least one independent source.

Models from the same vendor can clash across boards (e.g., Kling 3.0 scores 1092 on the T2V-with-audio sub-board but noticeably higher on the main board口径). Where they clash, I kept the range rather than hard-coding a single number.

Video Line: Wan3.0’s +53 Isn’t an Anomaly

First, the Arena.ai blind-test board for image-to-video, third week of September:

Weighted RankModelArena Elo (I2V)My Status
1Wan3.0-video1481 (#3), +53 over Wan2.7✅ Just接入 and set as default this week
2Seedance 2.01474 (#5); Seedance 2.5 has surged to 1477 (#4)✅ Channel live (mulerouter)
3Kling 3.0Main board口径 ~1460; sub-board口径 1092⚠️ SecretKey empty
4Wan2.7-r2v/i2v1426–1428 (#10–11)✅ 7 keys tested live, original主力
5Vidu Q3 Pro1363 (#18)⚠️ Key 401
6PixVerse V5.6/C1Formerly #2 on AA board (early口径)✅ Channel live
7happyhorse-1.0April “欢乐马” dominant record (Elo 1375, clear gap)✅ Pipeline default i2v
8grok-imagine-video480p tier 1381 (early口径)⚠️ Gateway submission no response, on hold

What does the top of the board look like: #1 MiniMax-H3 at 1494, #2 Gemini Omni 1.1 Flash at 1488, Wan3.0 at 1481 — only 7 points off, with a 57% win rate. And it only went live in late August, entering the board straight at #3. That 53-point gap falls exactly on my white-model anchor pipeline: Wan3.0 supports first_frame/last_frame with strong semantics (reference images and first/last frames are mutually exclusive — can’t be mixed), meaning programmatically rendered anchor frames can be fed in as “strict first frame / last frame,” rather than Wan2.7-r2v’s “multiple reference images for approximate anchoring.” Composition-locking capability upgrades from “approximate” to “strict.”

The scenario-specific selection口径 also交叉出来了: 智东西’s six-dimensional review of Seedance 2.0 (画质/运动/图像保持/语义/音频/文字) ranked first across all dimensions, scoring 3.31–3.70, with facial realism and precise action control beating Kling 3.0; Kling 3.0’s moat is native 4K60fps output and cinematic camera work. Mapped to my漫剧 scenarios: white-model anchoring主力 Wan3.0, close-up shots switch to Seedance, cinematic炫技 consider fixing Kling keys.

Image Line: The Current Default Is Already #1 — That’s Luck

The image line was the most effortless — after the audit, the current default config is already at the top:

Weighted RankModelArena Elo (T2I)My Status
1GPT-Image-2.5 (Sunburst/Flare)1421 / 1399; editing board 1520✅ .env default
2Gemini-3-Pro-Image-Preview口径 1050–1420,互有胜负 with GPT-ImageNot in production chain
3Qwen-Image-2.0/2.11029–1237; open-source line Qwen-Image-2512 is open-source #1✅ Pipeline default is same-family wan2.7-image-pro
4agnes-image-2.5Qwen Image Bench 57.53 (>Seedream 4.5/4.0)✅ CPA agnes channel
5Seedream 4.0/4.5Mid-low tier on AA boardNo direct connection
6grok-imagine-image-liteNo board data✅ CPA has it, lightweight tier

GPT-Image-2.5’s position is uncontested: T2I board Sunburst variant at 1421, leading the previous generation by ~40 points; the image editing board shows a 1520-vs-1461 gap, where multi-round editing fidelity and 4K support create a clear断层 in the editing scenario. Per新京报 reporting, Qwen-Image-2.0’s T2I score of 1029 (global #3) comes from February’s AI Arena; post-April versions (2512/2.1) topped the open-source line but still trail GPT by a full tier overall.

The practical conclusion is one line: default stays put. Add a scenario-specific alternative — for Chinese posters, text-dense images (漫剧 covers / title cards), switch to Qwen-Image; its Chinese text rendering is a genuine edge in this scenario.

Speech Line: The Only Action Item Worth Acting On

Ugly truth first: TTS has no authoritative blind-test board. Artificial Analysis’s Speech Arena covers speech-to-speech end-to-end, not pure TTS scoring; figures like 97ms first-packet and 17 voices are vendor口径. This section’s ranking is a features + reputation weighted judgment, with overall evidence strength a full tier lower than the previous two lines. Please read with that discount applied:

RankModelKey FeaturesMy Status
1CosyVoice-v3.5-plusAlibaba latest; voice cloning/design; same-family Qwen-Audio-3.0-TTS-Plus曾登 Speech Arena榜首✅ tts.py default
2Qwen3-TTS-Flash97ms first-packet, 17 voices, 9+ Chinese dialects, cloning/design, Dual-Track streaming✅ tts.py registered
3Gemini-3.1-Flash-TTSControllable narration tags, low latency, multilingual✅ CPA aistudio-gemini
4CosyVoice v2/v3-flash龙系 dozens of voices, older generation✅

Ranks 1 and 2 are actually the same vendor’s前后代 — the difference isn’t quality but scenario fit: Qwen3-TTS’s dialect coverage (9+ dialects) and instruction-based voice adjustment are more useful for character voice differentiation in古风漫剧, and 97ms first-packet is friendly for streaming synthesis. Switching cost is nearly zero — both families are already registered in tts.py, and the voice table (Cherry/Serena/Ethan/Chelsie — that batch of Qwen3 voices) is sitting right there.

Limitations & Reproduction

Three caveats, written here to prevent future self-deception:

  1. Elo is a living number. Arena weekly boards fluctuate ±10–20 points; Seedance 2.5 just pushed past 2.0 two weeks ago. The numbers in this post are fixed to the 2026-09-22 snapshot — re-verify after a month.
  2. TTS evidence strength is a tier lower, as flagged within the section.
  3. The “My Status” column is based on per-key live testing or source code verification. The 15-key endpoint tests used minimal requests (duration=1s error confirmed the auth and billing链路 is alive; actual output is a separate matter). A channel being “present” doesn’t mean “working” — Kling’s empty SecretKey and Vidu’s 401 are both live/dead checks one command away from verifying.

Reproduction path: model inventory comes from /config/model_catalog/generated/model_catalog.json (38 models) plus CPA config.yaml channel section; blind-test boards at arena.ai/leaderboard/image-to-video and lmarena.ai/leaderboard/image-to-video; Chinese-side cross-reference via SuperCLUE video board and Arena.ai Chinese weekly report. The key live-test request template is fully documented in Lumenx’s phase0-verdict doc.

One final number to close on: this audit took under two hours, and the most valuable action among them was those 15 one-by-one curls — it rewrote “video channel is dead” into “8 keys dead, 7 keys alive.”


Data Sources: LMArena I2V Board · Arena.ai official tweet: Wan3.0 I2V #3, 1481 · Arena.ai Week 36 weekly report (NetEase repost) · 智东西: Seedance 2.0 six-dimensional全第一 review · CometAPI: GPT-Image-2.5 benchmark · 新京报: Qwen-Image-2.0 launch · Qwen3-TTS technical deep dive · 澎湃: 欢乐马 dominant record