When picking an AI model for novel writing, there’s an ocean of opinions online—“Claude is the writing ceiling,” “DeepSeek has the strongest foreshadowing density,” “Kimi’s ultra-long-context continuation is in a league of its own.” Which of these are backed by real testing and which are marketing fluff? This article pulls all publicly available ranking data from September 2026, cross-evaluates across five dimensions that actually matter for novel writing, and gives you a practical selection guide you can follow.
Evaluation weights are set according to the real demands of serialized long-form fiction: Long-form Generation 30% · Long-context Recall 25% · Story Logic & Consistency 25% · Literary Style/Voice 15% · General Reasoning 5% (Coding is excluded—basically irrelevant for novel writing).
Bottom Line: Overall Rankings
| Rank | Model | Overall Score | One-sentence rationale |
|---|---|---|---|
| 🥇 | Claude Opus 5 | 92 | #1 in long-form writing (86.3, lowest slop at 5.6), Creative Writing Elo 2120, multi-needle long-context recall 93%@128K—the only model with no weak spots |
| 🥈 | GPT-6 Astra | 88 | Ceiling for general capability (GPQA 96), Creative Writing Elo 2163 best overall (tentative sample), but more long-form slop/repetition |
| 🥉 | Claude Fable 5 / 5.1 | 87 | Writing style on par with Opus 5 (Elo 2152), #2 on EQ-Bench Emotional Intelligence, but heavy quota consumption |
| 4 | Kimi K3 | 82 | Strongest in Chinese: EQ Creative Writing 2070 (non-US #1), EQ4 Emotional榜 #3, 1M-context continuation reputation unmatched |
| 5 | GLM-5.3 | 81 | Dark horse: EQ Creative Writing 2064, Long-form 81.8 ties GPT-5.6, AIME 2026 0.992 |
| 6 | GPT-5.6 Sol | 79 | Strong at writing but heavy “AI flavor” (slop 16.9), community consensus on readability gap |
| 7 | DeepSeek V4 Pro | 78 | 384K single-output longest in the field + top reputation for Chinese foreshadowing; weak in dense writing (slop 19.7) |
| 8 | Muse Spark 1.3 (Meta) | 77 | King of cost-effectiveness, long-form 82.8 ties GPT-6 Astra |
| 9 | Qwen3.8-Max | 76 | #1 in LongBench v2 long-form comprehension (66.3), MRCR 256K 8-needle 92.9 |
| 10 | Gemini 3.1 Pro | 74 | Delicate worldbuilding, but recall drops 50 points after 128K, weak Chinese language sense |
Overall score = five dimensions weighted and normalized, for ranking reference only—not precise measurement. Those marked with * are EQ-Bench tentative samples; ranks may shift.
One-line answer: If budget allows, Claude Opus 5; for pure Chinese serialized production, the combination of Kimi K3 + DeepSeek V4 Pro + GLM-5.3 delivers ~80% of the effect at a fraction of the cost.
Dimension 1: Long-form Generation Capability (Weight 30%)
1.1 EQ-Bench Creative Writing Longform (the most directly relevant “write chapters” benchmark)
This is currently the only public leaderboard specifically testing “long chapter writing” (judge is Claude Sonnet 4.6, 0–100 scale, with two penalty items: slop “AI tone” and repetition). Data source creative_writing_longform.js, 2026-09-07 snapshot, 134 models total.
| Rank | Model | Overall | Avg Chapter Length (tokens) | Slop↓ | Repetition↓ |
|---|---|---|---|---|---|
| 1 | claude-opus-5 | 86.3 | 6264 | 5.64 | 5.0 |
| 2 | claude-fable-5-1 * | 85.3 | 5777 | 7.59 | 5.4 |
| 3 | claude-fable-5 | 83.0 | 6295 | 8.31 | 4.4 |
| 4 | gpt-6-astra * | 82.8 | 5845 | 9.07 | 6.3 |
| 4 | muse-spark-1.3 * | 82.8 | 6253 | 10.73 | 4.8 |
| 6 | claude-opus-4-7 | 81.8 | 5552 | 9.06 | 4.6 |
| 6 | GLM-5.3 * | 81.8 | 5928 | 7.09 | 4.6 |
| 8 | gpt-5.6-sol | 81.7 | 6881 | 11.98 | 6.3 |
| 9 | muse-spark-1.2 * | 81.5 | 7258 | 11.43 | 4.5 |
| 10 | claude-opus-4-8 | 80.8 | 5460 | 9.39 | 3.8 |
| 11 | claude-sonnet-4-6 * | 79.9 | 6893 | 10.56 | 5.5 |
| 11 | ox-alpha * (hidden model) | 79.9 | 6302 | 7.43 | 4.0 |
| 13 | kimi-k3 | 79.6 | 7296 | 9.67 | 4.9 |
| 14 | Kimi-K2.6 | 78.5 | 6649 | 18.92 | 4.6 |
| 15 | gpt-5.4 | 78.3 | 8192 | 12.45 | 4.8 |
| 16 | claude-sonnet-5 | 78.3 | 5138 | 13.53 | 5.6 |
| 17 | gpt-5.5 | 78.2 | 8812 | 16.89 | 5.3 |
| 18 | gpt-5.6-terra | 78.0 | 7482 | 15.41 | 7.4 |
| 19 | GLM-5.2 | 77.9 | 5316 | 16.51 | 4.5 |
| 20 | claude-opus-4-6 * | 77.7 | 6189 | 16.86 | 4.7 |
| 21 | gemini-3.8-flash * | 76.8 | 7061 | 27.69 | 5.2 |
| 22 | DeepSeek-V4-Pro | 75.6 | 7516 | 21.83 | 3.9 |
| 24 | Kimi-K2.5 | 74.9 | 6315 | 17.28 | 4.8 |
| 27 | GLM-5.1 | 73.5 | 6271 | 24.94 | 4.4 |
| 31 | GLM-5 | 70.9 | 8065 | 24.04 | 4.4 |
| 35 | grok-4.20-beta | 68.5 | 8370 | 27.41 | 5.2 |
| 37 | gemini-3.1-pro-preview | 68.2 | 7433 | 38.12 | 4.9 |
| 40 | DeepSeek-V3.2 | 66.8 | 6496 | 41.24 | 4.9 |
| 41 | Qwen3-Max-2025-09-24 | 66.2 | 4403 | 49.19 | 7.7 |
| 42 | DeepSeek-V4-Flash | 66.0 | 5505 | 19.59 | 7.0 |
| 44 | DeepSeek-V4-Flash-0731 | 61.2 | 6221 | 31.44 | 7.6 |
| 46 | DeepSeek-R1 | 59.5 | 4035 | 56.44 | 6.4 |
| 51 | Qwen3.8-27B * | 53.3 | 6293 | 46.34 | 8.9 |
Note who’s missing from this leaderboard: Qwen3.8-Max, Qwen3.8-2.4T-A95B, grok-4.5/4.6 did not participate in longform evaluation—anyone quoting longform scores for these models is making it up.
1.2 Single-output token ceiling (determines whether a chapter can be written in one shot)
| Model | Max Output | Context Window |
|---|---|---|
| DeepSeek V4 Full Series (Pro/Flash/V4.1) | 384K (longest in field) | 1M |
| Claude Opus 5 / Fable 5.1 / Sonnet 5 | 128K | 1M |
| GPT-6 Astra | 128K | 1.05M |
| GLM-5.2 / 5.3 | 128K | 1M |
| Qwen3.8-2.4T-A95B (open-source weights) | Reasoning 262K / Prose 131K (recommended config) | 262K native, extensible to 1M |
| Kimi K3 | Not officially disclosed, community reputation: “seamless continuation after feeding 2M Chinese characters of prior context” | 1M |
DeepSeek’s 384K output ceiling means you could theoretically write 200K Chinese characters in one go—though you’d rarely use it that way. The real advantage is writing a complete 10–20K chapter in one shot without truncation and resumption.
Dimension 2: Literary Style / Voice (Weight 15%)
2.1 EQ-Bench Creative Writing v3 (main Creative Writing leaderboard, Elo-based)
2026-09-07 snapshot, 133 models. * = tentative sample.
| Rank | Model | Elo | Writing Score /20 | Slop↓ | Repetition↓ |
|---|---|---|---|---|---|
| 1 | gpt-6-astra * | 2163.9 | 16.80 | 8.41 | 3.54 |
| 2 | claude-fable-5-1 * | 2152.7 | 16.95 | 8.16 | 3.64 |
| 3 | claude-opus-5 | 2120.6 | 17.07 | 6.59 | 4.31 |
| 4 | kimi-k3 | 2070.6 | 16.85 | 9.70 | 3.73 |
| 5 | GLM-5.3 * | 2064.1 | 17.04 | 8.42 | 3.23 |
| 6 | gpt-5.6-sol | 1963.4 | 16.78 | 11.68 | 3.41 |
| 7 | ox-alpha * | 1960.6 | 16.89 | 9.83 | 3.05 |
| 8 | claude-fable-5 | 1934.6 | 16.81 | 10.28 | 3.92 |
| 9 | muse-spark-1.1 | 1916.1 | 16.54 | 12.11 | 3.49 |
| 10 | claude-opus-4-7 | 1907.1 | 16.57 | 11.09 | 4.00 |
| 11 | muse-spark-1.3 * | 1905.5 | 16.70 | 10.73 | 3.66 |
| 12 | gpt-5.6-terra | 1850.3 | 16.56 | 12.40 | 3.02 |
| 13 | gpt-5.5 | 1843.5 | 17.01 | 13.10 | 2.48 |
| 14 | Qwen3.8-2.4T-A95B * | 1840.9 | 16.72 | 12.26 | 3.50 |
| 15 | gpt-5.4 | 1835.6 | 16.89 | 12.20 | 2.71 |
| 16 | claude-opus-4-8 | 1835.2 | 16.66 | 13.16 | 3.68 |
| 17 | muse-spark-1.2 * | 1835.2 | 16.44 | 12.98 | 3.31 |
| 18 | gpt-5.6-luna | 1825.8 | 16.58 | 11.80 | 4.00 |
| 19 | claude-sonnet-4-6 | 1804.2 | 16.50 | 9.90 | 4.06 |
| 23 | GLM-5.2 | 1752.8 | 16.44 | 13.11 | 3.94 |
| 24 | gemini-3.8-flash * | 1749.7 | 16.54 | 22.60 | 3.12 |
| 25 | Kimi-K2.6 | 1721.3 | 16.67 | 13.30 | 3.77 |
| 26 | Qwen3.8-27B * | 1668.4 | 15.50 | 12.12 | 4.19 |
| 27 | Kimi-K2-Instruct | 1662.7 | 16.40 | 15.51 | 3.40 |
| 32 | GLM-5.1 | 1589.2 | 16.26 | 23.56 | 3.76 |
| 33 | grok-4.5 | 1576.0 | 16.25 | 17.73 | 3.52 |
| 35 | grok-4.20-beta | 1570.7 | 14.51 | 15.85 | 4.50 |
| 36 | DeepSeek-V4-Flash | 1555.7 | 16.29 | 20.92 | 4.30 |
| 37 | DeepSeek-V4-Pro | 1552.1 | 16.45 | 19.66 | 3.21 |
| 39 | DeepSeek-V3.2 | 1511.2 | 16.28 | 23.19 | 4.06 |
| 40 | DeepSeek-R1 (baseline) | 1500.0 | 15.68 | 31.21 | 4.62 |
| 48 | DeepSeek-V4-Flash-0731 | 1438.4 | 15.98 | 24.59 | 5.45 |
| 53 | gemini-2.5-pro-preview-06-05 | 1418.9 | 16.16 | 28.73 | 4.90 |
Not on this leaderboard (important): qwen3.8-max, Qwen3-Max, grok-4.6, grok-4.3 did not participate in this evaluation.
Key readings: GLM-5.3 (2064) is a full 311 Elo points above GLM-5.2 (1752)—same base model, purely separated by post-training. Zhipu really went all-in on this generation of post-training. Kimi K3 (2070) is the only non-US model in the top 10. DeepSeek V4 series scores low on literary style (1550 tier), confirming the community criticism of “overly dense writing”—on the EQ-Bench radar chart, V4-Pro’s weakest dimension is “Avoids Purple Prose,” while its strongest is “Creativity.”
2.2 LMArena Creative Writing Category (human blind votes, 2026-09-11)
Top 12 dominated by Claude and Gemini: 1. claude-fable-5 · 2. claude-opus-4-6-high · 3–4. gemini-3.7/3.8-flash-high · 5. claude-opus-4-7-high · 6. claude-fable-5.1-max · 7. gemini-3-pro · … Highest Chinese: glm-5.3-max #17, qwen3.8-max #18, kimi-k3-max #25.
Note the divergence between Arena and EQ-Bench: Gemini Flash high-thinking tier excels in Arena’s quick blind votes, but its slop is as high as 22.6 in EQ-Bench’s detailed evaluation—short chats may be pleasing, but they don’t sustain long-form reading.
2.3 Chinese “AI Flavor” Special (community consensus from real testing)
Cross-validated from four independent Chinese sources:
- Lightest AI flavor: Doubao (colloquial/internet feel), Claude (literary quality)
- Deepest reasoning: DeepSeek, Kimi
- No model outputs clean Chinese without some polishing—the difference is only in how heavy the base AI flavor is
Dimension 3: Long-context Recall (Weight 25%)
For serialized long-form, the model must remember foreshadowing you planted 50 chapters ago.
3.1 MRCR v2 Multi-needle Recall (by OpenAI, closest to “cross-referencing your story bible”)
| Scenario | Model | Score |
|---|---|---|
| 8 needles @128K | Claude Opus 4.6 | 93.0% (leading) |
| 8 needles @128K | Claude Sonnet 4.6 / Gemini 3.1 Pro | 84.9% |
| 8 needles @128K | GPT-5.5 | 74.0% |
| 8 needles @1M | Claude Opus 4.6 | 76–78% |
| 8 needles @1M | Gemini 3 Pro | 26.3% (cliff drop of 50 points from 128K→1M) |
| 4 needles @256K | GPT-5.2 | 98% (first to approach perfection at this scale) |
| 8 needles @256K | Qwen3.8-Max | 92.9% (official) |
| 8 needles @256K | Qwen3.7-Max | 86.7% |
| MRCR 1M | DeepSeek-V4-Pro-Max | 83.5% (official) |
3.2 Fiction.LiveBench (120K-token novel reading comprehension, most relevant to “remembering the first 100 chapters”)
2026-05 snapshot (earlier than H2 2026 flagships, for格局 reference only):
| Rank | Model | Score |
|---|---|---|
| 1 | o3 | 100.0 |
| 2 | GPT-5.2 | 96.9 |
| 2 | Grok 4 | 96.9 |
| 4 | Gemini 2.5 Pro | 90.6 |
| 5 | Qwen3 235B | 68.8 |
| 10 | MiniMax M2 | 59.4 |
| 13 | Claude 3.7 Sonnet | 53.1 |
| 16 | Kimi K2 Instruct | 40.6 |
| 19 | Claude Opus 4.5 | 37.5 |
| 21 | DeepSeek R1 | 33.3 |
⚠️ Note: Claude series scores anomalously low on this leaderboard (likely due to reasoning-model methodology advantage), opposite to the MRCR conclusion—the takeaway from cross-referencing both is: OpenAI/xAI excel at “novel text location and recall,” while Claude excels at “multi-thread cross-recall.” Also, the leaderboard site is currently inaccessible, and no H2 2026 flagships appear on it.
3.3 LongBench v2 (8K–2M character long-form comprehension)
| Rank | Model | Score |
|---|---|---|
| 1 | Qwen3.8-Max | 66.3 (official + third-party consistent) |
| 2 | Claude Opus 4.5 | 64.4 |
| 3 | Qwen3.5-397B-A17B | 63.2 |
| — | DeepSeek-V4-Pro (official self-report) | 51.5 |
| — | DeepSeek-V4.1-Flash (official self-report) | 45.2 |
3.4 Vendor-stated vs. actually reliable context (2026 community consensus from real testing)
| Model | Stated | Actually reliable range |
|---|---|---|
| Claude Opus 5 / Fable | 1M | ~200–400K |
| GPT-6 Astra | 1.05M | Single-needle up to 1M at 96% |
| Gemini 3.1 Pro | 1M (3 Pro claims 10M) | Excellent ≤128K, cliff after 256K |
| DeepSeek V4 full series | 1M | Single-needle @1M competitive, downstream tasks degraded |
| Grok 4 Fast | 2M | Few independent tests |
| Kimi K3 | 1M | Community实测 ultra-long continuation reputation unmatched |
Industry consensus: the reliable range is 50–70% of stated context. For chapters under 100K characters plus a story bible, top models are sufficient; for total context exceeding 300K characters, Claude (multi-needle) and GPT/Grok (single-needle location) are most stable.
Dimension 4: Story Logic & Consistency (Weight 25%)
This dimension has no public third-party quantitative leaderboard—all claims about “million-word non-degradation” or “character consistency” come from writing platform marketing articles (one platform claims “300K-character setting contradiction rate 1.3 vs. ChatGPT 4.7,” with no third-party reproduction, untrustworthy). Reliable evidence comes from same-prompt community tests:
Chinese Community Same-Prompt Tests (2026-09, traditional Chinese blogger tested 12 models on identical web novel outline)
| Model | Test Result |
|---|---|
| Claude Opus 4.8 Max / 4.7 | Strongest at foreshadowing weaving and chapter-level layout; 4.7 has the most natural prose |
| DeepSeek V4 | Doesn’t miss a single foreshadowing thread, but “like espresso—lacks breathing room” |
| Doubao | Best Chinese pun-based chapter naming, only 16 seconds of thinking, but plot drifts after 30K characters |
| GPT-5.5 Pro | 7-minute research-mode thinking, actually over-interprets instructions |
| Gemini 3.x | Delicate worldbuilding, built-in ideation parsing |
| Grok 4.x | Text + image in one shot, only one that ran end-to-end, but Chinese text is mediocre |
Web Novel Author Circle Division-of-Labor Consensus (Tahou platform 2026-09 test of 10 models)
Claude handles prose polishing and emotional scenes; DeepSeek handles logic deduction and setting audits; Kimi handles ultra-long prior-context continuation; Doubao handles fragmented inspiration.
Official Writing Data (DeepSeek V4 Technical Report §5.4.1, arXiv 2606.19348)
The only vendor that published official data treating “Chinese writing” as a core scenario:
| Comparison | Instruction Following Win Rate | Writing Quality Win Rate |
|---|---|---|
| DS V4-Pro vs Gemini-3.1-Pro (creative writing) | 60.0% | 77.5% |
| DS V4-Pro vs Claude Opus 4.5 (complex multi-turn instruction writing) | 45.9% vs 52.0% | — |
DeepSeek’s own official data admits: on the hardest constraint-heavy writing tasks, Claude Opus still leads. This triangulates with leaderboard data and community reputation.
Dimension 5: General Reasoning (Weight 5%)
Negligible differentiation for novel writing, quick overview:
| Leaderboard | #1 | Notable |
|---|---|---|
| GPQA Diamond | GPT-6 Astra 96.0 | Saturated, 24 models ≥90% |
| AIME 2026 | GLM-5.2 series 0.992 | Note: all 26 models are vendor self-reports, zero independent verification |
| MathArena Competition Overall | GPT-6 Astra 90.7% by a mile | Claude Opus 5 second at 72.1% |
| HLE | Claude Fable 5.1 0.650 | DeepSeek-V4-Pro-0813 0.600 (open-source #1) |
| Artificial Analysis Intelligence Index | Fable 5.1 = GPT-6 Astra = 53 | GLM-5.3 45, Kimi K3 44, Qwen3.8-Max 40 |
| LMArena Overall Elo | Fable 5.1 = GPT-6 Astra ≈ 1520 | Kimi K3 1506, Qwen3.8-Max 1506, GLM-5.3 1505, DeepSeek-V4.1-Flash 1503 |
Model Variant Clarifications: Three version numbers you’re likely to mix up
❌ “DeepSeek-V4-0831” Does Not Exist
The 0831 date belongs to DeepSeek-V4-Flash-Vision-Exp (multimodal experimental)—announced August 21, but the HuggingFace weights repo wasn’t created until August 31. “0831” is the repo creation timestamp, not a text model snapshot. DeepSeek’s official changelog only has two versions in August: 08-13 (V4-Pro GA) and 08-21 (Vision-Exp).
✅ DeepSeek V4 Full Series Real Version Timeline
| Version | Release Date | Positioning |
|---|---|---|
| V4 / V4-Pro / V4-Flash (preview) | 2026-04-24 | First-gen preview |
| V4-Flash-0731 | 2026-07-31 | Flash official release |
| V4-Pro-0813 | 2026-08-13 | Pro official release (the one with HLE 0.600 open-source #1) |
| V4-Flash-Vision-Exp | Announced 2026-08-21 | Multimodal experiment |
| V4.1-Flash | 2026-09-10 (yesterday) | New architecture Flash, GPQA 90.9 / officially claims full V4 Pro supersession |
⚠️ Important Timeliness: V4 Pro Downgraded September 14
DeepSeek official pricing page notice: Starting 2026-09-14 12:00 (Beijing Time), deepseek-v4-pro API requests will be routed to V4.1-Flash and billed at Flash rates, because V4.1-Flash has fully surpassed V4 Pro. V4.1 Pro has not yet been released. If you’re currently using v4-pro, watch for output behavior changes in three days.
❌ “qwen3.8-max-0902” Does Not Exist
Alibaba Cloud’s official documentation only has dated snapshots up to qwen3.7-max-2026-05-20 / qwen3.7-max-2026-06-08; qwen3.8-max is a dateless rolling alias. Its open-source counterpart is Qwen3.8-2.4T-A95B (HF official: “Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B”). Any reference to “0902” scores online is unverifiable.
Qwen3.8 / GLM-5.x Version Timeline
| Model | Released | Context | Output Cap | Notes |
|---|---|---|---|---|
| qwen3.8-max | 2026-08 | 1M (native 262K) | Reasoning 262K / Prose 131K (recommended) | = 2.4T-A95B official closed-source + vision + non-thinking mode |
| Qwen3.8-2.4T-A95B | 2026-08-12 | 262K native, extensible to 1M | Same | Open-source weights, 2.4T total params / 95B active MoE |
| Qwen3.8-27B | 2026-08-14 | 262K, extensible to 1M | Same | 27B dense compact powerhouse, native multimodal |
| GLM-5.2 | 2026-06-16 | 1M | 128K | 744B-A40B MoE, MIT open-source |
| GLM-5.3 | 2026-08-25 | 1M | 128K | Same base as 5.2, pure post-training improvement |
| GLM-5.3-Flash | Late 2026-08 | 1M | 128K | New base 320B/18B, first native multimodal GLM-5, ~1/10 the price of 5.3 |
(Note: glm-5.2-fast-preview cannot be found in Zhipu’s official docs—likely a temporary alias from a middleware provider, unverifiable through official channels.)
Final Selection Guide
Single-model Champion: Claude Opus 5
#1 in long-form, top 3 in literary style, strongest multi-needle recall, ceiling of Chinese community reputation—no weak spots across the five weighted dimensions. Use it for key chapters if budget allows.
Multi-model Division of Labor for Chinese Serialized Long-form
| Stage | Recommended Model | Rationale |
|---|---|---|
| Outline / worldbuilding brainstorm | Gemini 3.1 Pro / GLM-5.3 | Delicate setting, built-in ideation parsing |
| Key chapter prose | Claude Opus 5 | Long-form 86.3 + lowest slop |
| Mass-produced chapter prose | Kimi K3 / GLM-5.3 | EQ 2070/2064, far lower cost than Claude |
| One-shot ultra-long single chapter | DeepSeek V4 Pro / V4.1-Flash | 384K output ceiling |
| Setting audit / foreshadowing check | DeepSeek V4 Pro | Top foreshadowing |
