Featured image of post September 2026 Complete Guide to AI Novel-Writing Models: Full-Scale Multi-Dimensional Benchmark (Including Real-World Test Data for DeepSeek V4.1 / GLM-5.3 / Kimi K3 / Qwen3.8)

September 2026 Complete Guide to AI Novel-Writing Models: Full-Scale Multi-Dimensional Benchmark (Including Real-World Test Data for DeepSeek V4.1 / GLM-5.3 / Kimi K3 / Qwen3.8)

Based on the latest September 2026 data from leaderboards such as EQ-Bench, WritingBench, LongBench v2, Fiction.LiveBench, and MRCR, a comprehensive multi-dimensional evaluation of novel-writing capabilities across 30+ flagship models, with complete data tables attached.

When picking an AI model for novel writing, there’s an ocean of opinions online—“Claude is the writing ceiling,” “DeepSeek has the strongest foreshadowing density,” “Kimi’s ultra-long-context continuation is in a league of its own.” Which of these are backed by real testing and which are marketing fluff? This article pulls all publicly available ranking data from September 2026, cross-evaluates across five dimensions that actually matter for novel writing, and gives you a practical selection guide you can follow.

Evaluation weights are set according to the real demands of serialized long-form fiction: Long-form Generation 30% · Long-context Recall 25% · Story Logic & Consistency 25% · Literary Style/Voice 15% · General Reasoning 5% (Coding is excluded—basically irrelevant for novel writing).


Bottom Line: Overall Rankings

RankModelOverall ScoreOne-sentence rationale
🥇Claude Opus 592#1 in long-form writing (86.3, lowest slop at 5.6), Creative Writing Elo 2120, multi-needle long-context recall 93%@128K—the only model with no weak spots
🥈GPT-6 Astra88Ceiling for general capability (GPQA 96), Creative Writing Elo 2163 best overall (tentative sample), but more long-form slop/repetition
🥉Claude Fable 5 / 5.187Writing style on par with Opus 5 (Elo 2152), #2 on EQ-Bench Emotional Intelligence, but heavy quota consumption
4Kimi K382Strongest in Chinese: EQ Creative Writing 2070 (non-US #1), EQ4 Emotional榜 #3, 1M-context continuation reputation unmatched
5GLM-5.381Dark horse: EQ Creative Writing 2064, Long-form 81.8 ties GPT-5.6, AIME 2026 0.992
6GPT-5.6 Sol79Strong at writing but heavy “AI flavor” (slop 16.9), community consensus on readability gap
7DeepSeek V4 Pro78384K single-output longest in the field + top reputation for Chinese foreshadowing; weak in dense writing (slop 19.7)
8Muse Spark 1.3 (Meta)77King of cost-effectiveness, long-form 82.8 ties GPT-6 Astra
9Qwen3.8-Max76#1 in LongBench v2 long-form comprehension (66.3), MRCR 256K 8-needle 92.9
10Gemini 3.1 Pro74Delicate worldbuilding, but recall drops 50 points after 128K, weak Chinese language sense

Overall score = five dimensions weighted and normalized, for ranking reference only—not precise measurement. Those marked with * are EQ-Bench tentative samples; ranks may shift.

One-line answer: If budget allows, Claude Opus 5; for pure Chinese serialized production, the combination of Kimi K3 + DeepSeek V4 Pro + GLM-5.3 delivers ~80% of the effect at a fraction of the cost.


Dimension 1: Long-form Generation Capability (Weight 30%)

1.1 EQ-Bench Creative Writing Longform (the most directly relevant “write chapters” benchmark)

This is currently the only public leaderboard specifically testing “long chapter writing” (judge is Claude Sonnet 4.6, 0–100 scale, with two penalty items: slop “AI tone” and repetition). Data source creative_writing_longform.js, 2026-09-07 snapshot, 134 models total.

RankModelOverallAvg Chapter Length (tokens)Slop↓Repetition↓
1claude-opus-586.362645.645.0
2claude-fable-5-1 *85.357777.595.4
3claude-fable-583.062958.314.4
4gpt-6-astra *82.858459.076.3
4muse-spark-1.3 *82.8625310.734.8
6claude-opus-4-781.855529.064.6
6GLM-5.3 *81.859287.094.6
8gpt-5.6-sol81.7688111.986.3
9muse-spark-1.2 *81.5725811.434.5
10claude-opus-4-880.854609.393.8
11claude-sonnet-4-6 *79.9689310.565.5
11ox-alpha * (hidden model)79.963027.434.0
13kimi-k379.672969.674.9
14Kimi-K2.678.5664918.924.6
15gpt-5.478.3819212.454.8
16claude-sonnet-578.3513813.535.6
17gpt-5.578.2881216.895.3
18gpt-5.6-terra78.0748215.417.4
19GLM-5.277.9531616.514.5
20claude-opus-4-6 *77.7618916.864.7
21gemini-3.8-flash *76.8706127.695.2
22DeepSeek-V4-Pro75.6751621.833.9
24Kimi-K2.574.9631517.284.8
27GLM-5.173.5627124.944.4
31GLM-570.9806524.044.4
35grok-4.20-beta68.5837027.415.2
37gemini-3.1-pro-preview68.2743338.124.9
40DeepSeek-V3.266.8649641.244.9
41Qwen3-Max-2025-09-2466.2440349.197.7
42DeepSeek-V4-Flash66.0550519.597.0
44DeepSeek-V4-Flash-073161.2622131.447.6
46DeepSeek-R159.5403556.446.4
51Qwen3.8-27B *53.3629346.348.9

Note who’s missing from this leaderboard: Qwen3.8-Max, Qwen3.8-2.4T-A95B, grok-4.5/4.6 did not participate in longform evaluation—anyone quoting longform scores for these models is making it up.

1.2 Single-output token ceiling (determines whether a chapter can be written in one shot)

ModelMax OutputContext Window
DeepSeek V4 Full Series (Pro/Flash/V4.1)384K (longest in field)1M
Claude Opus 5 / Fable 5.1 / Sonnet 5128K1M
GPT-6 Astra128K1.05M
GLM-5.2 / 5.3128K1M
Qwen3.8-2.4T-A95B (open-source weights)Reasoning 262K / Prose 131K (recommended config)262K native, extensible to 1M
Kimi K3Not officially disclosed, community reputation: “seamless continuation after feeding 2M Chinese characters of prior context”1M

DeepSeek’s 384K output ceiling means you could theoretically write 200K Chinese characters in one go—though you’d rarely use it that way. The real advantage is writing a complete 10–20K chapter in one shot without truncation and resumption.


Dimension 2: Literary Style / Voice (Weight 15%)

2.1 EQ-Bench Creative Writing v3 (main Creative Writing leaderboard, Elo-based)

2026-09-07 snapshot, 133 models. * = tentative sample.

RankModelEloWriting Score /20Slop↓Repetition↓
1gpt-6-astra *2163.916.808.413.54
2claude-fable-5-1 *2152.716.958.163.64
3claude-opus-52120.617.076.594.31
4kimi-k32070.616.859.703.73
5GLM-5.3 *2064.117.048.423.23
6gpt-5.6-sol1963.416.7811.683.41
7ox-alpha *1960.616.899.833.05
8claude-fable-51934.616.8110.283.92
9muse-spark-1.11916.116.5412.113.49
10claude-opus-4-71907.116.5711.094.00
11muse-spark-1.3 *1905.516.7010.733.66
12gpt-5.6-terra1850.316.5612.403.02
13gpt-5.51843.517.0113.102.48
14Qwen3.8-2.4T-A95B *1840.916.7212.263.50
15gpt-5.41835.616.8912.202.71
16claude-opus-4-81835.216.6613.163.68
17muse-spark-1.2 *1835.216.4412.983.31
18gpt-5.6-luna1825.816.5811.804.00
19claude-sonnet-4-61804.216.509.904.06
23GLM-5.21752.816.4413.113.94
24gemini-3.8-flash *1749.716.5422.603.12
25Kimi-K2.61721.316.6713.303.77
26Qwen3.8-27B *1668.415.5012.124.19
27Kimi-K2-Instruct1662.716.4015.513.40
32GLM-5.11589.216.2623.563.76
33grok-4.51576.016.2517.733.52
35grok-4.20-beta1570.714.5115.854.50
36DeepSeek-V4-Flash1555.716.2920.924.30
37DeepSeek-V4-Pro1552.116.4519.663.21
39DeepSeek-V3.21511.216.2823.194.06
40DeepSeek-R1 (baseline)1500.015.6831.214.62
48DeepSeek-V4-Flash-07311438.415.9824.595.45
53gemini-2.5-pro-preview-06-051418.916.1628.734.90

Not on this leaderboard (important): qwen3.8-max, Qwen3-Max, grok-4.6, grok-4.3 did not participate in this evaluation.

Key readings: GLM-5.3 (2064) is a full 311 Elo points above GLM-5.2 (1752)—same base model, purely separated by post-training. Zhipu really went all-in on this generation of post-training. Kimi K3 (2070) is the only non-US model in the top 10. DeepSeek V4 series scores low on literary style (1550 tier), confirming the community criticism of “overly dense writing”—on the EQ-Bench radar chart, V4-Pro’s weakest dimension is “Avoids Purple Prose,” while its strongest is “Creativity.”

2.2 LMArena Creative Writing Category (human blind votes, 2026-09-11)

Top 12 dominated by Claude and Gemini: 1. claude-fable-5 · 2. claude-opus-4-6-high · 3–4. gemini-3.7/3.8-flash-high · 5. claude-opus-4-7-high · 6. claude-fable-5.1-max · 7. gemini-3-pro · … Highest Chinese: glm-5.3-max #17, qwen3.8-max #18, kimi-k3-max #25.

Note the divergence between Arena and EQ-Bench: Gemini Flash high-thinking tier excels in Arena’s quick blind votes, but its slop is as high as 22.6 in EQ-Bench’s detailed evaluation—short chats may be pleasing, but they don’t sustain long-form reading.

2.3 Chinese “AI Flavor” Special (community consensus from real testing)

Cross-validated from four independent Chinese sources:

  • Lightest AI flavor: Doubao (colloquial/internet feel), Claude (literary quality)
  • Deepest reasoning: DeepSeek, Kimi
  • No model outputs clean Chinese without some polishing—the difference is only in how heavy the base AI flavor is

Dimension 3: Long-context Recall (Weight 25%)

For serialized long-form, the model must remember foreshadowing you planted 50 chapters ago.

3.1 MRCR v2 Multi-needle Recall (by OpenAI, closest to “cross-referencing your story bible”)

ScenarioModelScore
8 needles @128KClaude Opus 4.693.0% (leading)
8 needles @128KClaude Sonnet 4.6 / Gemini 3.1 Pro84.9%
8 needles @128KGPT-5.574.0%
8 needles @1MClaude Opus 4.676–78%
8 needles @1MGemini 3 Pro26.3% (cliff drop of 50 points from 128K→1M)
4 needles @256KGPT-5.298% (first to approach perfection at this scale)
8 needles @256KQwen3.8-Max92.9% (official)
8 needles @256KQwen3.7-Max86.7%
MRCR 1MDeepSeek-V4-Pro-Max83.5% (official)

3.2 Fiction.LiveBench (120K-token novel reading comprehension, most relevant to “remembering the first 100 chapters”)

2026-05 snapshot (earlier than H2 2026 flagships, for格局 reference only):

RankModelScore
1o3100.0
2GPT-5.296.9
2Grok 496.9
4Gemini 2.5 Pro90.6
5Qwen3 235B68.8
10MiniMax M259.4
13Claude 3.7 Sonnet53.1
16Kimi K2 Instruct40.6
19Claude Opus 4.537.5
21DeepSeek R133.3

⚠️ Note: Claude series scores anomalously low on this leaderboard (likely due to reasoning-model methodology advantage), opposite to the MRCR conclusion—the takeaway from cross-referencing both is: OpenAI/xAI excel at “novel text location and recall,” while Claude excels at “multi-thread cross-recall.” Also, the leaderboard site is currently inaccessible, and no H2 2026 flagships appear on it.

3.3 LongBench v2 (8K–2M character long-form comprehension)

RankModelScore
1Qwen3.8-Max66.3 (official + third-party consistent)
2Claude Opus 4.564.4
3Qwen3.5-397B-A17B63.2
—DeepSeek-V4-Pro (official self-report)51.5
—DeepSeek-V4.1-Flash (official self-report)45.2

3.4 Vendor-stated vs. actually reliable context (2026 community consensus from real testing)

ModelStatedActually reliable range
Claude Opus 5 / Fable1M~200–400K
GPT-6 Astra1.05MSingle-needle up to 1M at 96%
Gemini 3.1 Pro1M (3 Pro claims 10M)Excellent ≤128K, cliff after 256K
DeepSeek V4 full series1MSingle-needle @1M competitive, downstream tasks degraded
Grok 4 Fast2MFew independent tests
Kimi K31MCommunity实测 ultra-long continuation reputation unmatched

Industry consensus: the reliable range is 50–70% of stated context. For chapters under 100K characters plus a story bible, top models are sufficient; for total context exceeding 300K characters, Claude (multi-needle) and GPT/Grok (single-needle location) are most stable.


Dimension 4: Story Logic & Consistency (Weight 25%)

This dimension has no public third-party quantitative leaderboard—all claims about “million-word non-degradation” or “character consistency” come from writing platform marketing articles (one platform claims “300K-character setting contradiction rate 1.3 vs. ChatGPT 4.7,” with no third-party reproduction, untrustworthy). Reliable evidence comes from same-prompt community tests:

Chinese Community Same-Prompt Tests (2026-09, traditional Chinese blogger tested 12 models on identical web novel outline)

ModelTest Result
Claude Opus 4.8 Max / 4.7Strongest at foreshadowing weaving and chapter-level layout; 4.7 has the most natural prose
DeepSeek V4Doesn’t miss a single foreshadowing thread, but “like espresso—lacks breathing room”
DoubaoBest Chinese pun-based chapter naming, only 16 seconds of thinking, but plot drifts after 30K characters
GPT-5.5 Pro7-minute research-mode thinking, actually over-interprets instructions
Gemini 3.xDelicate worldbuilding, built-in ideation parsing
Grok 4.xText + image in one shot, only one that ran end-to-end, but Chinese text is mediocre

Web Novel Author Circle Division-of-Labor Consensus (Tahou platform 2026-09 test of 10 models)

Claude handles prose polishing and emotional scenes; DeepSeek handles logic deduction and setting audits; Kimi handles ultra-long prior-context continuation; Doubao handles fragmented inspiration.

Official Writing Data (DeepSeek V4 Technical Report §5.4.1, arXiv 2606.19348)

The only vendor that published official data treating “Chinese writing” as a core scenario:

ComparisonInstruction Following Win RateWriting Quality Win Rate
DS V4-Pro vs Gemini-3.1-Pro (creative writing)60.0%77.5%
DS V4-Pro vs Claude Opus 4.5 (complex multi-turn instruction writing)45.9% vs 52.0%—

DeepSeek’s own official data admits: on the hardest constraint-heavy writing tasks, Claude Opus still leads. This triangulates with leaderboard data and community reputation.


Dimension 5: General Reasoning (Weight 5%)

Negligible differentiation for novel writing, quick overview:

Leaderboard#1Notable
GPQA DiamondGPT-6 Astra 96.0Saturated, 24 models ≥90%
AIME 2026GLM-5.2 series 0.992Note: all 26 models are vendor self-reports, zero independent verification
MathArena Competition OverallGPT-6 Astra 90.7% by a mileClaude Opus 5 second at 72.1%
HLEClaude Fable 5.1 0.650DeepSeek-V4-Pro-0813 0.600 (open-source #1)
Artificial Analysis Intelligence IndexFable 5.1 = GPT-6 Astra = 53GLM-5.3 45, Kimi K3 44, Qwen3.8-Max 40
LMArena Overall EloFable 5.1 = GPT-6 Astra ≈ 1520Kimi K3 1506, Qwen3.8-Max 1506, GLM-5.3 1505, DeepSeek-V4.1-Flash 1503

Model Variant Clarifications: Three version numbers you’re likely to mix up

❌ “DeepSeek-V4-0831” Does Not Exist

The 0831 date belongs to DeepSeek-V4-Flash-Vision-Exp (multimodal experimental)—announced August 21, but the HuggingFace weights repo wasn’t created until August 31. “0831” is the repo creation timestamp, not a text model snapshot. DeepSeek’s official changelog only has two versions in August: 08-13 (V4-Pro GA) and 08-21 (Vision-Exp).

✅ DeepSeek V4 Full Series Real Version Timeline

VersionRelease DatePositioning
V4 / V4-Pro / V4-Flash (preview)2026-04-24First-gen preview
V4-Flash-07312026-07-31Flash official release
V4-Pro-08132026-08-13Pro official release (the one with HLE 0.600 open-source #1)
V4-Flash-Vision-ExpAnnounced 2026-08-21Multimodal experiment
V4.1-Flash2026-09-10 (yesterday)New architecture Flash, GPQA 90.9 / officially claims full V4 Pro supersession

⚠️ Important Timeliness: V4 Pro Downgraded September 14

DeepSeek official pricing page notice: Starting 2026-09-14 12:00 (Beijing Time), deepseek-v4-pro API requests will be routed to V4.1-Flash and billed at Flash rates, because V4.1-Flash has fully surpassed V4 Pro. V4.1 Pro has not yet been released. If you’re currently using v4-pro, watch for output behavior changes in three days.

❌ “qwen3.8-max-0902” Does Not Exist

Alibaba Cloud’s official documentation only has dated snapshots up to qwen3.7-max-2026-05-20 / qwen3.7-max-2026-06-08; qwen3.8-max is a dateless rolling alias. Its open-source counterpart is Qwen3.8-2.4T-A95B (HF official: “Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B”). Any reference to “0902” scores online is unverifiable.

Qwen3.8 / GLM-5.x Version Timeline

ModelReleasedContextOutput CapNotes
qwen3.8-max2026-081M (native 262K)Reasoning 262K / Prose 131K (recommended)= 2.4T-A95B official closed-source + vision + non-thinking mode
Qwen3.8-2.4T-A95B2026-08-12262K native, extensible to 1MSameOpen-source weights, 2.4T total params / 95B active MoE
Qwen3.8-27B2026-08-14262K, extensible to 1MSame27B dense compact powerhouse, native multimodal
GLM-5.22026-06-161M128K744B-A40B MoE, MIT open-source
GLM-5.32026-08-251M128KSame base as 5.2, pure post-training improvement
GLM-5.3-FlashLate 2026-081M128KNew base 320B/18B, first native multimodal GLM-5, ~1/10 the price of 5.3

(Note: glm-5.2-fast-preview cannot be found in Zhipu’s official docs—likely a temporary alias from a middleware provider, unverifiable through official channels.)


Final Selection Guide

Single-model Champion: Claude Opus 5

#1 in long-form, top 3 in literary style, strongest multi-needle recall, ceiling of Chinese community reputation—no weak spots across the five weighted dimensions. Use it for key chapters if budget allows.

Multi-model Division of Labor for Chinese Serialized Long-form

StageRecommended ModelRationale
Outline / worldbuilding brainstormGemini 3.1 Pro / GLM-5.3Delicate setting, built-in ideation parsing
Key chapter proseClaude Opus 5Long-form 86.3 + lowest slop
Mass-produced chapter proseKimi K3 / GLM-5.3EQ 2070/2064, far lower cost than Claude
One-shot ultra-long single chapterDeepSeek V4 Pro / V4.1-Flash384K output ceiling
Setting audit / foreshadowing checkDeepSeek V4 ProTop foreshadowing