Actionable insights — China LLM primer
The repeatable analysis behind the calls: not which AI names GS picked, but the frameworks it used to score them — written so they can be rerun on any AI-model company, Chinese or Western.
How to read this page: each insight is a method — the rubric GS built to rank AI-model companies and the signals it watches. The boxed line shows how it played out in this primer.
1. The three-axis rubric for scoring an AI-model company
The repeatable method
- Score pricing power on three observables: time-to-market (does each flagship release meaningfully close the gap to SOTA vs its own prior generation?), LMArena score (blind, large-scale real-usage reviews — deliberately preferred over static benchmarks, which can be gamed), and realized blended price per 1M tokens (sustained pricing through iterations = power; cuts = a volume strategy).
- Score cost advantage — the structural cost-to-serve that sets the floor for price and margin: throughput (tokens/second per GPU), cache hit rate (cached tokens are near-free to serve but still billed — margin-accretive), parameter size and activation ratio (fewer active parameters = fewer FLOPs per token), and inference gross margin where disclosed.
- Score financial strength — capacity to fund frontier R&D before profitability: cash on hand, net cash as % of assets, and the valuation multiple (P/ARR for unprofitable independents, P/E for mega-caps).
- Overlay token-scale and market-share progression — the winner profile is largest ARR scale × gross-margin advantage + financial strength, not any single axis.
Here: the rubric ranks Zhipu (2513.HK) and DeepSeek strongest in foundation models and ByteDance in multi-modal, while surfacing 0100.HK MiniMax as the Buy — top-decile cost efficiency at ~13x P/ARR vs peers several times higher.
Watch for
- Any AI-model company pitched on benchmark scores alone: rerun the three axes — real-usage Arena rank, cost-to-serve mechanics, and cash runway — before accepting the story.
2. The ARR-quadrant map — decompose revenue into tokens × price before comparing players
The repeatable method
- Treat AI-model ARR as token volume × realized price per token, and plot every player on that grid rather than comparing headline ARRs.
- Only two quadrants maximize ARR: premium pricing on frontier intelligence/multi-modal (where pricing power is real), or attractive pricing on very large token volumes (the agentic low end). The middle — mediocre price on mediocre volume — is where companies bleed.
- Check which tier the segment is in: expect the premium coding tier to consolidate around the best models, and the low-end agentic tier to stay a fragmented price war for as long as players carry cash buffers willing to run zero/negative gross margins.
- Adjust reported ARR for the open-source leak: free third-party hosting means disclosed ARR can understate true deployment multiple-fold — and a license shift to open-weight/Community-License (revenue share on commercial use) is the catalyst that converts that hidden deployment into revenue.
Here: GLM5.2/Qwen3.7-Max hold ~$1 per 1M tokens (5x the low end) in the premium quadrant; MiniMax M3 wins the volume quadrant; Zhipu's $1bn ARR target understates worldwide GLM usage because platforms like Alibaba's Bailian host it fee-free.
Watch for
- License changes (MIT → Community License / open-weight) at Chinese model companies — GS expects them broadly, and each one is a step-change in monetizable ARR without new demand needing to appear.
3. The token-maxxing test — judge AI adoption by cost per task, not token volume
The repeatable method
- When token-consumption growth is cited as proof of AI adoption, test the quality of that consumption: compare token growth to output growth (the Jellyfish study's tell — heavy users at 10x tokens for only 2x output means waste, not productivity).
- Watch for enterprise-behavior inflections that mark the regime change: monthly caps per tool, scrapped internal token leaderboards, defaults downgraded to cheap flash models with SOTA reserved for top-value tasks.
- Switch metrics with the market: cost per task, Daily Active Agents (DAA) and Agentic Work Units (AWU) replace price-per-token and raw volume once buyers turn ROI-first.
- Position for the mix shift, not against it: an ROI-first world routes typical workloads to value-for-money models (structurally favoring cheap Chinese open-source) while frontier models keep only the highest-value work — margin implications on both sides.
Here: the pivot from token-maxxing to ROI-first is GS's core demand argument for Chinese models going global — SMEs and enterprises managing token costs adopt "good enough" models at 10–25% of US SOTA pricing, driving the 25x token-growth forecast.
Watch for
- Enterprise disclosures shifting from token-volume boasts to cost-per-task / agent-count metrics — the same transition that ended "growth at any cost" in SaaS; it re-prices who captures AI spend.
Methods distilled from a proprietary Goldman Sachs Global Investment Research primer (PDF linked above), for personal study. Not investment advice. Source material © Goldman Sachs Global Investment Research.