Featured image of post OceanBase Tops Global Data Agent Benchmark, First to Cross 90% Accuracy

OceanBase Tops Global Data Agent Benchmark, First to Cross 90% Accuracy

OceanBase's Scout, built on GLM-5.2, achieves 90.62% accuracy, ranking first on the DAB leaderboard.

Breakthrough: Chinese Data Agent Tops Global Benchmark

Breakthrough: Chinese Data Agent Tops Global Benchmark
Breakthrough: Chinese Data Agent Tops Global Benchmark|News screenshot

OceanBase’s Data Agent solution (internal code: Scout) has topped the international Data Agent Benchmark (DAB) with 90.62% accuracy, becoming the first participating solution to cross the 90% accuracy threshold. The underlying technology will be integrated into OceanBase DataPilot.

Key technical facts:

  • Benchmark result: 90.62% accuracy, ranked #1
  • Underlying model: Chinese large language model GLM-5.2
  • Database foundation: Chinese database OceanBase
  • Solution code name: Scout
  • Integration target: OceanBase DataPilot product

Technical Architecture: Closing the Loop from Data Understanding to Answer Verification

DAB, developed by UC Berkeley’s EPIC Data Lab and Hasura PromptQL, evaluates solutions across multiple domains including internet/lifestyle, financial stocks, biomedicine, intellectual property, enterprise operations, government management, and media/entertainment. It supports PostgreSQL, MongoDB, SQLite, and DuckDB.

Unlike traditional Text-to-SQL benchmarks that only assess AI’s ability to convert natural language to SQL, DAB tests end-to-end task completion in real data environments. A complete Data Agent workflow involves:

  • Data understanding: Identifying field semantics and relational structures
  • Data selection and planning: Architecting analysis paths based on task complexity
  • Query execution: Performing data filtering, joining, and aggregation
  • Result validation: Verifying computation through evidence tracing; adjusting and re-running when errors are detected

OceanBase’s approach implements a closed-loop system: “data understanding → planning → execution → validation → correction”.

Key Surprise: Chinese Model Stack Outperforms Foreign LLMs

The most unexpected finding is that the solution built on the domestic GLM-5.2 model outperformed multiple Data Agent approaches based on GPT, Claude Opus, and Claude Fable.

This demonstrates that Data Agent performance depends not on individual model provenance, but on the synergy among three layers: the model (reasoning), the Agent (planning/execution), and the data system (discovery, computation, verification). All three must work cohesively.

Industry Impact: Database Evolving into AI Data Platform

With AI Agents becoming new data consumers, database roles are shifting fundamentally. Historically, databases handled storage, computation, and retrieval. Today, they must also help AI understand, organize, and analyze data. OceanBase’s DAB success validates its strategic pivot from “delivering data to AI” to “enabling AI to use data effectively.”

The AI-era data foundation is transforming from passive response to active collaboration—from storage engine to data work engine.

Practical Recommendations

  • Who should adopt now: Enterprises needing complex analysis workflows with robust error correction capabilities should monitor DataPilot’s productization progress
  • Who should wait: Teams requiring only basic Text-to-SQL for simple queries may find existing solutions sufficient; this breakthrough targets mid-to-high complexity use cases demanding end-to-end automation

Final Thoughts

While a single benchmark cannot represent real-world performance, DAB clearly signals a technology shift: Data Agent value lies not in model parameter scale, but in system-level coordination precision and reliability. Domestic tech stacks demonstrate significant potential in this emerging domain.