Goldman Sachs — LLM primer: China's open-source models at a critical point of intelligence for global proliferation
"From DeepSeek's moment last year (on cost efficiency) to Zhipu's GLM moment this year (on model intelligence)" — GS Asia tech research (Keung, Sheridan & team), 48-page primer
One-line take: GS's Asia team argues China's open-source/open-weight AI models have reached a "critical point of intelligence" versus global proprietary models — good enough for agentic tasks and specific coding scenarios at 10–25% of US SOTA pricing — and that ramping domestic-enterprise plus global-SME adoption creates a positive data flywheel of further improvement. The primer introduces a Competitive Positioning framework (pricing power × cost advantage × financial strength) and picks Zhipu (Knowledge Atlas, initiated Neutral at a $110bn valuation) and DeepSeek (private) as strongest-positioned in foundation models, ByteDance (private) in multi-modal; MiniMax is the team's Buy-rated way in. China AI-model revenue is forecast to grow from Rmb35bn (2026E) to Rmb879bn by 2030E on ~25x domestic token growth, with "going global" the key upside and US/Western market access, frontier-model export restrictions, and high-end-compute access the three swing risks. Proprietary 48-page sell-side primer — the summary here is our own; the PDF is the archived source. The eli5 blocks below carry the theses.
1. Stocks & names mentioned
| Ticker | Name | Research | View | What GS said | At |
| 0100.HK | MiniMax Group | STK | Positive | Buy reiterated, 12-month TP HK$860 (DCF). Stands out on cost efficiency under the framework; M3 sits in the favored ARR-maximizing quadrant (attractive pricing + high token volumes); 60–70% of revenue already overseas; trades at ~13x P/2026E year-end ARR vs peers several times higher — risk-reward skewed up. Next drivers: the imminent H3 video model and M3 coding updates (Jul–Aug), a larger-parameter M3 in 2H26. | read ↗ |
| BABA | Alibaba Group | QT · SA · STK · FA | Positive | Buy, SOTP TP US$186/HK$180. The Qwen family pioneered China's open-source approach (now closed-source for the top Qwen-Max tier for monetization); Qwen3.7 Max holds pricing power at ~$1 per 1M blended tokens; Qoder coding harness + Alibaba Cloud's Bailian MaaS platform are distribution; capex-to-cloud-revenue conversion benchmarked against Amazon/Google. | read ↗ |
| TCEHY | Tencent Holdings | QT · SA · STK | Positive | Buy, 12-m SOTP TP HK$700. One of the mega-cap AI players with balance-sheet strength to sustain the low-end price war; its Workbuddy agentic/coding harness is a named 2H26 signpost as model companies race to capture real-life coding data. | read ↗ |
| 1024.HK | Kuaishou Technology | STK | Positive | Buy, TP HK$68 (15x avg 2026–27E EPS). Its Kling video-generation model (US$18bn post-money implied valuation) is a key player in the multi-modal segment GS expects to keep healthy pricing and gross margins — unlike foundation text — with demand outpacing capacity into 2H26. | read ↗ |
| 1810.HK | Xiaomi Corp. | STK | Positive | Buy, TP HK$40 (SOTP). Covered as one of the mega-cap China AI players in the Competitive Positioning framework — profiled mainly on financial strength (net-cash estimates alongside Alibaba/Tencent) rather than a flagship frontier model. | read ↗ |
| 2513.HK | Knowledge Atlas Technology (Zhipu) | STK | Neutral | Initiated Neutral at a US$110bn valuation — "China's top enterprise/coding AI model." Its GLM5.2 is the intelligence moment of the report's title: ~$1 per 1M blended tokens (5x low-end China models, 10–25% of US SOTA), strongest-positioned in foundation models alongside DeepSeek; GS thinks its $1bn year-end-2026 ARR target understates worldwide GLM deployment multiple-fold since third parties host the open-source model free. | read ↗ |
| DeepSeek | DeepSeek (private, China) | — | Neutral | Strongest-positioned in foundation models with Zhipu. V4 official launch set for mid-July with new peak/off-peak API pricing (peak at 2x); the DSpark speculative-decoding framework made V4 serving 57–85% faster with no quality loss; 1.6T-parameter V4 Pro runs at a fraction of US SOTA size and cost. | read ↗ |
| ByteDance | ByteDance (private) | — | Neutral | Leads multi-modal under the framework. Closed-source Seed model + the Doubao app give it ~70% of China's domestic token share alone; per LatePost/36Kr its Seedance video model runs ~70% gross margins at a US$2bn+ ARR run-rate. | read ↗ |
| 3690.HK | Meituan | STK | Neutral | Case study: its LongCat 2.0 (Jun 30) is China's first 1.6-trillion-parameter open-source MoE model trained and deployed entirely on a 50,000-card domestic compute cluster (reportedly Huawei Atlas-950 SuperPods) — GS calls it a milestone proving Chinese silicon can handle frontier-class training, not just inference, driving a more self-sufficient China AI outlook. | read ↗ |
| MSFT | Microsoft | QT · SA · STK · FA | Neutral | Cited for its CEO's WSJ remarks that Microsoft is considering hosting versions of DeepSeek on Copilot as an optional, cost-effective model — operated inside Azure so customer data stays in its cloud — giving customers cheaper choices alongside US proprietary models. | read ↗ |
| GOOGL | Alphabet | QT · SA · STK · FA | Neutral | Its Gemini Enterprise Agent Platform (via Model Garden) already offers a broad selection of Chinese AI models (DeepSeek, MiniMax, Moonshot, GLM, Qwen) fully managed inside the US cloud ecosystem — the open-source distribution channel underpinning GS's "going global" leg. | read ↗ |
| AMZN | Amazon | QT · SA · STK · FA | Neutral | AWS Bedrock likewise hosts Chinese models fully managed in the US cloud; GS also flags open-weight revenue-sharing on hyperscaler platforms (AWS Bedrock, Alibaba Bailian) as a potential high-margin monetization path for Chinese model companies. | read ↗ |
2. Talking points
The headline thesis — from a cost moment to an intelligence moment
- Last year was DeepSeek's moment (cost efficiency); this year is Zhipu's GLM moment (model intelligence): Chinese open-source/open-weight models are reaching a "critical point" of intelligence vs global proprietary models — "good enough" for agentic tasks and specific coding scenarios.
- The flywheel claim: per LatePost, AI-generated code has reached as high as 90% at some China mega-caps (from 20–30% in 2H25); real-world coding adoption feeds reinforcement learning on actual user data, reducing reliance on distillation — GS credits this for GLM5.2's step-up over GLM5.1 within months and expects further step improvements over the next 6–12 months.
How they do it cheap — small parameters, MoE, and DSpark
- Chinese models run 200bn–1.6T total parameters (2–10% of leading SOTA models, a consequence of constrained access to high-end compute) with MoE / Sparse Attention architectures activating only 3–5% of parameters — driving much lower training and inference costs.
- Flagship sizes: DeepSeek V4 Pro 1.6T, Zhipu GLM5.2 0.7T, MiniMax M3 0.4T. DeepSeek's DSpark speculative-decoding framework (Jun 27) made V4-Flash serving 60–85% faster and V4 Pro 57–78% faster without changing weights or output quality.
- Result: top Chinese models price at ~US$1 per 1M blended tokens — 10–25% of US SOTA at US$4–8 — while still earning 10–20% gross margins (GSe), lower than global SOTA peers due to weaker pricing power.
A two-tier market and the two ARR-maximizing quadrants
- The market is bifurcating: frontier-performance / multi-modal models keep pricing power (GLM5.2, Qwen3.7 Max at ~$1, about 5x low-end China models); the low-end agentic segment ($0.06–0.2 per 1M tokens) is in a price war — yet agentic AI is driving explosive volume demand exactly there.
- GS maps players on token volume × pricing and identifies two favored "ARR-maximizing" quadrants: premium pricing on frontier intelligence, or attractive pricing on huge token volumes (MiniMax's M3 sits in the latter).
- Expect the coding segment to consolidate around the best models (premium pricing, like the US) while the agentic low-end stays fragmented for longer — mega-cap and well-funded independents can subsidize zero/negative gross margins for a while.
Why open source — and how the money eventually shows up
- Open source maximizes flexibility (where models are trained and deployed, inside and outside China), adoption and community feedback — and offers a worldwide alternative when the best US proprietary models carry stricter access and premium prices.
- Disclosed ARRs understate reality: third-party clouds can host open-source models free even for commercial use (Alibaba's Bailian hosts GLM5.2 with no fee to Zhipu), so Zhipu's $1bn year-end-2026 ARR target likely understates worldwide GLM deployment multiple-fold; some Chinese models are even re-branded by global companies with no revenue back.
- The monetization path GS expects: a shift from pure open source (MIT license, free for all) to open weight with a Community License — commercial terms / revenue-share for commercial use, as MiniMax's M-series already does — so inference gross profits can eventually cover training costs. Take-rate deals on hyperscaler platforms (AWS Bedrock, Bailian) could be highly margin-accretive since the model company bears no inference cost.
The TAM math — 25x token growth, Rmb879bn by 2030
- GS estimates China AI models' aggregate API + subscription revenue grows from Rmb35bn (2026E) to Rmb879bn by 2030E; implied Chinese daily token consumption rises from 350T (2026E) to 4,600T by 2030E (~25x domestic growth).
- Domestic share today: at ~140T daily national tokens in March 2026 (several hundred trillion by June), open-source/open-weight models hold roughly 30% token share — ByteDance alone (closed-source, Doubao-driven) holds ~70%.
- International is the key upside: GS's US team separately models agentic AI driving 24x global token growth by 2030 (to ~4 quadrillion tokens/day), with Chinese models gaining global (ex-China) token share as SMEs conclude they are "good enough" at far lower cost.
The enterprise pivot — from "token-maxxing" to ROI-first
- GS sees a paradigm shift from "token-maxxing" (late 2025–early 2026, when high token consumption was equated with productivity) to an ROI-first model: a Jellyfish study found heavy enterprise AI users consumed 10x more tokens for only 2x more output, and several US internet companies burned a year's AI budget in four months or had to scrap token leaderboards that incentivized low-value agent runs.
- The new metrics: clear task boundaries, Daily Active Agents (DAA), Agentic Work Units (AWU), back-end automation and cost per task — not price per token. Enterprises are downgrading default models to cheap flash tiers and reserving SOTA models for the highest-value tasks (coding), a mix shift that structurally favors value-for-money Chinese models.
US hyperscalers as a distribution channel
- Alphabet's Gemini Enterprise Agent Platform (Model Garden) and AWS Bedrock already host a broad selection of Chinese models (DeepSeek, MiniMax, Moonshot, GLM, Qwen), fully managed within the US cloud ecosystem.
- Microsoft's CEO said at a WSJ interview that Microsoft is considering hosting DeepSeek versions on Copilot as an optional cost-effective model — run inside Azure so customer data never leaves — a notable legitimization of Chinese models inside US enterprise software.
The Competitive Positioning framework — and the winners
- Long-term winners = largest ARR scale (token scale × pricing power) + gross-margin advantage (training/inference efficiency) + financial strength (balance sheet, access to compute). Scored on quantifiable metrics: time-to-market, LMArena score (blind real-usage reviews, not static benchmarks), blended pricing; throughput/GPU, cache hit rate, parameter/activation ratio, inference GPM; cash on hand, net cash as % of assets, valuation multiple (P/ARR for independents, P/E for mega-caps).
- Verdict: Zhipu (Knowledge Atlas) and DeepSeek strongest-positioned in text foundation models; ByteDance leads multi-modal. Independent AI model companies in aggregate represent over US$200bn of implied valuations.
- Financial strength matters because the low-end price war will persist: cash-rich players can subsidize zero/negative gross-margin pricing near-term — GS keeps API pricing "suppressed" at $0.1–0.2 per 1M tokens through 2H26.
Meituan's LongCat 2.0 — the self-sufficiency milestone
- Released Jun 30, 2026: China's first 1.6T-parameter open-source MoE model trained and deployed entirely on a 50,000-card domestic cluster (reportedly Huawei Atlas-950 SuperPods), with a native 1M-token context window (LongCat Sparse Attention) activating ~48bn parameters per token.
- GS's read: prior flagships used domestic chips for inference; proving end-to-end pre-training of a trillion-parameter model on Chinese silicon overcomes the critical memory/distributed-stability bottlenecks — a fundamentally more self-sustainable China AI outlook, less reliant on foreign high-end chips.
Signposts into 2H26
- Harness/agentic apps as entry points — Zhipu's ZCode, Tencent's Workbuddy, Alibaba's Qoder — as model companies close the loop on capturing real-life coding/agentic data; enterprises adopt a multiple-models approach judged on cost per task.
- Multiple 2–5T-parameter Chinese model launches in 2H26; coding competition intensifies against GLM's leadership; multi-modal/visual understanding is the next upgrade for GLM/DeepSeek (MiniMax M3 already has it); DeepSeek V4 launches mid-July with 2x peak-hour pricing.
- Video generation keeps healthy pricing/margins (ByteDance SeeDance, Kuaishou Kling, MiniMax Hailuo/H3) with demand far outpacing capacity; press reports flag potential future restrictions on overseas access to China's most advanced models as frontier AI is treated as a national asset.
Key risks — both ways
- Downside: Western market-access limits (possibly mirroring TikTok's trajectory — local-jurisdiction computing/data rules), restrictions on high-end leased compute for training, entity-list designations, anti-distillation/IP claims, cash burn in a fragmented landscape, and SLM/new-architecture competition.
- Upside: a faster open-weight shift (take-rate revenue without inference costs) transforming model companies into high-gross-margin businesses; stronger-than-expected intelligence gains from the data flywheel; restrictions could even accelerate China's AI self-sufficiency across software/CPU/ASICs.
3. In plain English
2513.HK — Knowledge Atlas Technology (Zhipu) Neutral
Zhipu — newly public in Hong Kong as Knowledge Atlas — makes GLM, the Chinese AI model that Goldman calls China's best for enterprise work and coding. The report's whole premise is named after it: last year the world was shocked that Chinese models were cheap (DeepSeek); this year the shock is that they're genuinely smart (Zhipu's GLM). GLM5.2 charges about $1 per million tokens — a quarter or less of what top US models charge — and it's now good enough that big Chinese companies reportedly generate up to 90% of their code with AI, much of it on GLM.
The subtle but important point: because GLM is open-source, anyone (including Alibaba's cloud) can host it and sell access without paying Zhipu a cent. That means Zhipu's official $1bn revenue target badly understates how widely its model is actually used — and Goldman expects Zhipu and peers to move to an "open-weight" license that finally charges commercial users a cut. Goldman rates it the strongest-positioned foundation-model company in China alongside DeepSeek, but starts coverage at Neutral: at a $110bn valuation, a lot of that promise is already in the price.
0100.HK — MiniMax Group Positive
MiniMax is Goldman's actual Buy in this primer. Its M3 model is small (0.4T parameters), extremely cheap to run, and sits exactly where the explosive demand is: bargain-priced AI for "agentic" work — the automated software agents that small businesses and one-person companies run around the clock. Unusually for a Chinese AI firm, 60–70% of its revenue already comes from overseas, and it also owns a strong video-generation line (Hailuo, with the H3 model about to launch) where pricing and margins are much healthier than in text AI.
The valuation argument is simple: MiniMax trades at ~13x its expected end-2026 recurring revenue, while comparable AI companies in China and globally command multiples several times higher at a similar stage. Goldman's framework flags its weak spots honestly — it lacks the pricing power of Zhipu and the balance sheet of the giants — but thinks the risk-reward is skewed upward at this price, with a HK$860 target.
DeepSeek — DeepSeek (private) Neutral
DeepSeek remains the reference name for cheap frontier AI. Its V4 Pro model is 1.6 trillion parameters — a tenth or less the size of leading US models — and its engineering keeps compounding: a June framework called DSpark made its models respond 57–85% faster without changing their output at all. Goldman scores it, together with Zhipu, as the strongest-positioned foundation-model player in China.
Two telling details in the note: DeepSeek is about to launch the official V4 in mid-July, and it's introducing peak-hour pricing — charging double during Chinese business hours — because demand for its compute now exceeds supply. A price war at the bottom of the market and surge pricing at the top is exactly what "critical point of adoption" looks like. As a private company there's nothing to buy directly; it matters here as the pace-setter that keeps compressing what the world is willing to pay for AI.
ByteDance — ByteDance (private) Neutral
ByteDance is the odd one out in China's open-source wave: its Seed model is fully closed and proprietary, like OpenAI's. It doesn't need openness for distribution because it already owns the audience — its Doubao chatbot is China's #1, and ByteDance alone accounts for roughly 70% of all AI tokens consumed in China. In video generation, its SeeDance model reportedly runs a healthy ~70% gross margin at a $2bn+ revenue run-rate, in an segment where demand far exceeds available computing power.
Goldman's framework crowns it the leader in multi-modal AI (video, image, audio) — the part of the market that, unlike text, still holds pricing power. Private, so not investable directly, but it's the benchmark every listed Chinese AI player is measured against.
BABA — Alibaba Group Positive
Alibaba plays both sides of the Chinese AI boom. Its Qwen model family did more than any other to establish China's open-source approach — and now that the strategy has worked, Alibaba has quietly closed the source on its very best Qwen-Max models so it can charge properly for them (about $1 per million tokens, top-tier pricing for a Chinese model). Meanwhile its cloud arm monetizes everyone else's models too: the Bailian platform hosts competitors like Zhipu's GLM, and Alibaba's cloud is one of the places global enterprises go to run Chinese AI.
Goldman is Buy-rated with a $186 target, and benchmarks Alibaba's capex-to-cloud-revenue conversion against Amazon and Google — the framing being that Alibaba is China's closest equivalent to a US hyperscaler: a mega-cap with the balance sheet to keep buying compute through the price war, the flagship model, the cloud distribution, and the coding harness (Qoder) to capture the data flywheel.
Summary distilled from a proprietary Goldman Sachs Global Investment Research primer (PDF linked above) — quoted only in short fragments. For personal study — not investment advice. Source material © Goldman Sachs Global Investment Research.