Featured image of post What the Financial Industry Actually Tracks, What's Still Done by Humans, and Where Individuals Can Break In: An In-Depth Study of 114 AI Agents

What the Financial Industry Actually Tracks, What's Still Done by Humans, and Where Individuals Can Break In: An In-Depth Study of 114 AI Agents

Ran a 97-minute adversarial validation deep research with 114 AI sub-agents: A comprehensive map of the information objects and data sources the financial industry must track (巨潮/深证信, Choice, SEC EDGAR, Massive, FMP, Quiver), the areas still stuck in manual repetitive work, and the opportunities and compliance minefields for individual developers looking to break in.

I previously ran a round of research and concluded that finance is one of the industries with the deepest AI penetration. Deep penetration means a solid data foundation, many players, and potentially more opportunities. This time I broke the question into three parts and ran them through a multi-agent deep-research pipeline for 97 minutes: What does the financial industry need to track in real time? Which tracking and processing tasks are still manual repetitive labor? Where can individual developers or small teams carve in?

Let me lay out the methodology upfront, because every key conclusion in this piece survived a three-vote adversarial verification process (three independent verification agents each attempted to knock down the same conclusion; a conclusion had to survive at least two votes to stand). The ones that got knocked down, I’ve disclosed honestly. 114 sub-agents, 1,752 tool calls in total; all vendor pricing is a snapshot as of 2026-09-16.

Overseas Financial Data API Pricing Tiers
Monthly fee tier comparison across three overseas data APIs (official website snapshot, 2026-09-16)

I. What the Financial Industry Tracks, and Where the Data Comes From

Tracking targets fall into five categories: announcements and disclosures (earnings reports, prospectuses, shareholder shareholding changes, lock-up expirations), regulatory actions (penalties, disciplinary measures, inquiry letters), holdings and trading movements (fund holdings, insider trading, 13F filings), ratings and opinions (analyst ratings, earnings call transcripts), and sentiment and credit risk. Each category has a completely different landscape in China versus overseas markets.

China market: Official channels are more open than you’d expect, but the processing layer is already occupied.

The authoritative source for announcements is 巨潮资讯网 (cninfo.com.cn) — designated by the CSRC, operated by Shenzhen Securities Information Co., Ltd. (a wholly-owned subsidiary of the Shenzhen Stock Exchange), covering the Shanghai, Shenzhen, and Beijing exchanges plus a Hong Kong mirror, funds, and bonds. The key intelligence here is its official API platform, 深证信数据服务平台 (webapi.cninfo.com.cn): the announcement query endpoint is live and usable (in this round of deep research, I pulled 7,694 Hong Kong announcements and 4,256 Shenzhen announcements in the first half of September). You register a token and call it directly — no need to scrape web pages.

深证信 itself is an API marketplace. Out of 2,466 endpoints, it already offers AI-processed products like “stock intelligent summaries” and “full-text announcement extraction” — but note that 2,445 of those 2,466 endpoints are flagged as “internal interfaces,” and intelligent-summary products are delivered via custom engagement rather than self-serve purchase. In other words, “announcement → AI summary” is not a blank space in China — the official authority has already commercialized it. What an individual can self-serve purchase is far narrower than what the catalog suggests.

The one-stop commercial source is 东方财富 Choice (quantapi.eastmoney.com, full SDK suite covering Python/MATLAB/R/C++/C#/Java). What really made me shelve the “sentiment monitoring startup” idea: Choice has already productized its sentiment early-warning system — covering 2 million+ enterprises, with three-level penetration across enterprise–group–related parties, multi-channel auto-push, and linkage to credit ratings and bond repayment progress. The general-purpose sentiment monitoring space is done — the giants have already built it out.

The regulatory side, on the other hand, is a pain point: The Shanghai Stock Exchange scatters regulatory measures across five-plus independent sections — listed company supervision dynamics, issuance and listing review, bond supervision, trading supervision, and member supervision — with no unified “penalty notice” entry point on the official website. Tracking accountability for intermediaries (regulatory warnings against sponsor representatives, signing accountants, signing lawyers) requires manually monitoring multiple pages. This type of information is continuously produced at high frequency (the latest entry is dated 2026-09-09); the dispersed structure itself is a source of genuine demand for aggregation tools.

Overseas market: Red ocean, but pricing tiers are clear.

SEC EDGAR is free and authoritative, but has two hard gates: a global 10 requests/second rate limit (aggregated across all your machines), and a mandatory User-Agent header declaring company information (omitting it gets you a straight 403). These two constraints are the dividing line between free official sources and paid commercial APIs.

On the market data API front, there’s a news-worthy intelligence item: Polygon.io has renamed itself to Massive.com (site-wide 301 redirect; the official SDK self-describes as “formerly Polygon.io”). The personal tier has four levels at $0 / $29 / $79 / $199; all paid tiers offer unlimited calls, with the differences being latency and historical depth.

FMP (Financial Modeling Prep) has API-ified nearly every tracking target you can think of: 8-K filings, insider trading, congressional trading, 13F institutional holdings, earnings call transcripts, analyst ratings, press releases. The free tier gives you 250 calls/day for zero-cost validation. But earnings call transcripts (covering 8,200 companies, 200,000+ full texts) are locked behind the $99/month Ultimate tier — this conclusion passed by a 2-1 vote, and vendor pricing could change at any time.

On the alternative-data front, Quiver Quantitative anchors its personal-tier pricing at $30–$75/month across 15+ datasets (congressional trading, lobbying, government contracts, patents, executive compensation); commercial use is negotiated separately.

深证信 API Structure
Of 深证信’s 2,466 endpoints, 2,445 are internal interfaces; the individual self-serve scope is far smaller than what the catalog suggests

II. Which Tasks Are Still Manual Repetitive Labor

Let me be upfront: the judgments in this section rely mainly on indirect evidence (the dispersed structure of official sections, inferences drawn from Choice’s automated products, and the absence of first-hand interviews with institutional workflows). The confidence level is inherently lower than Part I.

Three confirmed manual labor hotspots:

  1. Regulatory action aggregation. Five sections are scattered with no unified entry point; people doing risk control or media work can only manually rotate through them. This is directly visible from the Shanghai Stock Exchange website structure.
  2. Manual monitoring of long-tail information. What giant products cover is the “core list” — listed companies and bond issuers. The long tail beyond those 2 million enterprises (e.g., local financial regulatory bureaus monitoring microfinance companies in their jurisdiction, or vertical industry players tracking their own supply chain companies) is most likely still manual.
  3. Final sign-off. The workflow where AI generates a summary and a human reviews and signs off on it is still pervasive inside institutions — AI does the initial screening, humans do the judgment. This means there’s room for a collaboration tool layer around “AI output → human gatekeeping,” but purely automated, sell-the-conclusion products won’t fly on the institutional side in China.

III. Where Individual Developers Can Carve In

Factor in competitive density and whitespace, and my ranking is:

Opportunity 1: Regulatory penalty aggregation pipeline (China, most certain whitespace). CSRC penalty database, Shanghai/Shenzhen Stock Exchange disciplinary actions, NFRA penalties, and AMAC disciplinary actions — all dispersed, with no aggregation product. Build it as a “regulatory radar” with daily push updates and sell subscriptions. Risk: this round of deep research confirmed that the Shanghai Stock Exchange’s actual scraping difficulty is higher than surface-level judgment suggests (see the knocked-down list below). Before building, you must do hands-on anti-scraping testing site by site. And the compliance ceiling is unknown (see Opportunity 4).

Opportunity 2: Vertical industry sentiment monitoring (the only path to differentiation). General-purpose monitoring is done — Choice has built it out. But “only monitor one specific vertical industry, with accurate industry-terminology recognition, pushing to small and mid-size clients who don’t use Wind” still has cracks. The essence is being the long-tail complement to the giants.

Opportunity 3: Overseas vertical aggregation and processing layer. Data APIs themselves are a red ocean (FMP/Massive/Quiver have clear tiering), but the combination of “EDGAR free source + AI summary + vertical scenario” still has room — for instance, serving only one type of investor, or processing only one type of filing. Data cost floor ranges from $0 (EDGAR) to $29/month (Massive Starter).

Opportunity 4 (actually a minefield): Compliance. Does providing A-share/Hong Kong stock data to clients require a data provider registration or licensing? What are the authorization boundaries for reselling Choice/深证信 data? These two questions did not get reliable answers in this round of deep research (Grok declined to answer compliance-related questions, and no clear regulations were found in primary sources either). My stance: Before building any financial data product面向 the public in China, you must spend legal consulting budget to resolve this question first — it directly determines the feasibility ranking of the three opportunities above. Overseas vendors generally distinguish personal vs. commercial licensing; resale boundaries also need to be confirmed vendor by vendor.

IV. Claims Knocked Down by Adversarial Verification (Honestly Disclosed)

The most valuable output of this deep research is often “what got falsified”:

  • “The Shanghai Stock Exchange regulatory sections are server-side rendered HTML, scrapable without JS, and have low crawling difficulty” — 0 votes for, 3 votes against, knocked down. Actual scraping difficulty is higher than it appears (anti-scraping/dynamic rendering). If you want to build an aggregation pipeline, test first — don’t trust “official sections are all easy to scrape.”
  • “FMP’s full bundle can be had for under $100/month” — 1 vote for, 2 votes against, knocked down. Transcripts are only on the $99/month Ultimate tier, billed annually (≈$1,188/year). “Full bundle under $100” doesn’t hold.

A few explicitly unverified gaps remain: structured access costs for HKEX disclosure易 (hkexnews.hk), the CSRC penalty database (csrc.gov.cn), NFRA, and AMAC, as well as the credibility of China-side AI research assistant competitors (问财 and 妙想 are relatively trustworthy; Wind’s “问数” and 慧博智研 were not verified against official pages) — these are topics for the next round of deep research.

Closing

My overall takeaway from this run: the “data acquisition layer” and “general-purpose processing layer” of finance AI are saturated in both markets. The genuine cracks are in aggregating dispersed information and verticalized processing — precisely where individual developers have an advantage in stamina. In China, the biggest uncertainty isn’t technology — it’s compliance; overseas, the biggest cost isn’t the data fee — it’s figuring out which vertical to target.

The complete research process (including all source links and per-conclusion verification records) has been archived. Reach out if you want the raw report.