Supermodel Arms Race: Gemini 4 Pro Suspected Leak, Anthropic Accelerates Compute

On September 18, Google’s next-generation flagship large language model, Gemini 4 Pro (internal codename Argon), was reportedly leaked online under the name ‘gemini-3.8-flash’. Though Google has not officially confirmed whether this version equals Gemini 4 Pro, leaked benchmark tests suggest its performance significantly outperforms peers. According to screenshots circulating online, the model achieved 88%+ on DeepSWE v1.1, 95.3% on Terminal-bench 2.1, and 86.8% on OSWorld 2.0—outpacing OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1 in all three agent benchmarks. This represents the key surprise: performance metrics surfaced while the model was still under secrecy, bypassing standard release protocols.
However, the test screenshot lacksfull metadata, conditions, or raw data, making independent verification impossible. Google has not commented. The leak underscores the intensifying velocity race in large model development—iteration cycles are compressing as evaluation shifts from single parameters to comprehensive agent capability across diverse tasks.
Compute Arms Race Heats Up: Anthropic Targets 5 GW by Year-End, 10 GW in 2027

Compute infrastructure is now the decisive battlefield. According to The New York Times citing sources, Anthropic plans to scale its available compute power to ~5 gigawatts (GW) by the end of 2026, then ~10 GW by late 2027. To contextualize: Anthropic’s compute stood at ~1.5 GW as of 2025 year-end, while OpenAI reached ~2 GW. If executed, Anthropic’s capacity would grow over threefold in one year, dramatically closing the gap.
Both firms now follow nearly identical expansion paths—OpenAI targets 5 GW in 2026 and 10 GW in 2027. This alignment reveals that top AI labs are racing toward the same compute milestones, driven by exponentially rising training costs. Training a single trillion-parameter model now exceeds millions of dollars; compute control directly dictates iteration speed and technological leadership.
Data Ethics controversy Deepens: Microsoft exec Calls AI Training ‘Largest Labor Theft’
Even as model performance surges, data legitimacy issues are facing high-level scrutiny. Internal Microsoft documents obtained by The New York Times show应用 science director Brent Hecht labeling current AI training practices as ’the largest labor theft in human history’ in a January 2023 internal memo. He stated that ‘almost nobody intended their work to be used this way, nor received compensation’.
Hecht also introduced the ‘doom loop’ concept for newsmedia: AI products erode news site traffic and revenue → news outlets reduce content production → AI models lose reliable information sources → self-sabotage. He emphasized in January 2024 materials that ‘a terminal product threatening the economic foundation of its core supplier is highly unusual—yet this is precisely the position we’ve created for the LLM business’.
Security and Product Updates: OpenAI Community Breached, Chatterfly Opens Beta

On security, researchers leveraged Claude Opus 5 to breach OpenAI’s internal systems within 72 hours. The attack chain began with a libheif image-processing buffer overflow in the Discourse-based community forum, escalated via a separate SSO flaw to gain employee account access, and ultimately verified write permissions to the GitHub codebase. Hacktron AI claimed Claude generated ARM64 exploits within hours and adapted them to x86-64 systems. OpenAI patched the issues within a day of receiving the July 25 report via Bugcrowd, paying a $6,500 bounty.
Product-wise, Tencent’s AI-native input tool Chatterfly quietly opened macOS/Windows beta access. Positioned as an ‘AI expression assistant’, it focuses on auto-rewriting spoken content into polished书面 text, supporting six scenarios including meeting notes, work reports, and coding prompts.语音 input processes locally by default with no cloud upload, plus user vocabulary and style adaptation. Unlike traditional input methods, Chatterfly remains desktop-only; iOS/Android versions are marked ‘coming soon’.
Pricing and Market: iPhone Duo Repair Costs Unconfirmed, ZCode Offers Quota Reset

Apple’s first foldable iPhone Duo carries a starting price of ¥15,999, with pre-orders opening October 16 and launch on October 23. Regarding the rumored ¥8,000 repair cost, Apple officially stated ‘specific pricing has not been announced’. The estimate stemmed from a third-party repair shop’s analysis of a damaged unit ( Tim’s reverse bending test), not an Apple official quote. Apple clarified that reverse bending constitutes人为 damage excluded from standard warranty coverage.
ZCode faced controversy over its ‘codebase indexing’ feature triggering unauthorized data uploads. The wiki generation function defaulted to enabling uploads, prompting internal mishandling.智谱 has fixed the issue, committed to open-sourcing the codebase for third-party review, and distributed one-time weekly quota resets to all users.
Write-Up
The large model competition is evolving from pure parameter benchmarks into a holistic contest spanning compute infrastructure, security resilience, and ethical data practices. As performance metrics keep breaking ceilings, the legitimacy of training data sources, the sustainability of escalating costs, and the protection of end-user rights have become the critical pivot points determining whether the technological flywheel remains stable—or spins out of control.
