← Analysis page  ·  WSJ Heard on the Street hub  ·  Research hub

Actionable insights — Does going all-in on AI actually pay?

The repeatable analysis behind the piece: not "buy Shopify," but how to test whether a company's AI adoption is producing real operating leverage — and how to tell an AI-disruption narrative that will stick from one that will unwind.
2026-SEP-08 · WSJ AI & Business · Asa Fitch · Read ↗ · full analysis · transcript
How to read this page: each insight is a method — the test, the diagnostic, the tell — distilled from the article so it can be rerun on the next company claiming an AI transformation, or the next sector the market decides AI is about to destroy. The boxed line shows how it played out here.

1. The revenue-vs-headcount test — the only AI claim that can't be faked

The repeatable method
  1. Ignore AI product announcements and "AI-powered" language in the release. Pull two series instead: revenue and period-end headcount, from the year AI adoption began to the latest reported year.
  2. Compute revenue per employee at both ends. Real operating leverage from AI shows up as revenue rising while headcount is flat or falling — not as a margin footnote management attributes to AI.
  3. Check the denominator honestly: a headcount fall that came from selling or closing a business, or from a plain cost-cutting layoff into a shrinking top line, is not AI leverage. The signature you want is growth on a smaller base.
  4. Then ask what mechanism produced it — a company that can name the specific policy (below) has a repeatable process; one that can't has had a good year.
Here: SHOP went from ~$7bn revenue in 2023 to a projected $15bn+ this year while headcount fell from ~11,600 (end-2022) to 7,600 (end-2025) — revenue roughly doubling on a workforce about a third smaller. That pairing, not the Sidekick launch, is what the article treats as evidence the bet is working.
Watch for

2. Distinguish an AI mandate from AI tools — look for the hiring gate

The repeatable method
  1. Sort companies into two buckets: those that provide employees AI tools to experiment with (the default, and near-worthless as a differentiator), and those that have made AI use a condition of how work is done.
  2. Look for the three hard mechanisms that separate them, all of which are observable from memos, filings and interviews: (a) AI use scored in performance evaluations; (b) AI-first prototyping on new projects; (c) a hiring gate — teams must demonstrate the job cannot be done with AI before a headcount request is approved.
  3. The hiring gate is the highest-signal item: it is the only one that mechanically converts AI capability into the headcount line, which is where the leverage in test 1 comes from.
  4. Treat a mandate without those mechanisms as marketing, and a mandate whose mechanisms are in place as a reason to expect the revenue-per-employee series to bend.
Here: Lütke's early-2025 memo made AI a "baseline expectation" at SHOP and produced exactly those three mechanisms — reviews, AI-dominated prototyping, and a prove-you-can't-do-it-with-AI test before hiring. Controversial at the time, when prevailing wisdom was to hand out tools and let people experiment.
Watch for

3. Make management disclose an AI-attributed operating metric — and then track its rate

The repeatable method
  1. On the earnings call, find the metric that is specific to the AI channel rather than a total — AI-driven traffic, AI-originated orders, agent-sourced sessions — and its year-over-year rate.
  2. Prefer a metric that is small-but-tripling to one that is large-but-vague: an early channel growing at a triple-digit rate is a leading indicator; "AI is embedded across the platform" is not measurable and cannot be falsified.
  3. Log the number each quarter and watch the second derivative. A new channel decelerating from 3x to 1.4x while still described in the same language is the tell that the narrative has outrun the business.
  4. Separately ask what share of the total it is — a tripling channel that is 1% of orders changes the story only if the rate persists for several more quarters.
Here: Shopify President Harley Finkelstein said on last month's analyst call that AI-driven traffic and orders to Shopify stores tripled year-over-year in Q2 — the one hard, dated, AI-specific operating number in the piece, and the thing to re-check next quarter.
Watch for

4. In a platform shift, ask who becomes the supply the new interface must call

The repeatable method
  1. When the interface to an industry changes — here, from a human searching to an AI agent recommending — do not just ask who builds the new interface. Ask what the new interface needs to consume.
  2. Identify the party that owns the aggregated, structured supply the agents cannot easily rebuild (inventory, catalog, pricing, fulfillment) and check whether it is making itself easy to integrate with every model developer rather than betting on one.
  3. Neutrality is the test: a supplier wired into OpenAI, Anthropic and others is positioned to win regardless of which model wins — the picks-and-shovels position in an agent world.
  4. Then verify with test 3 that the position is producing measurable volume, not just a press release.
Here: SHOP is putting its merchants' product catalog where OpenAI, Anthropic and other developers can easily tap it for recommendations — positioning itself as the supply layer of "agentic commerce" rather than trying to own the assistant.
Watch for

5. Test an AI-disruption narrative against the maintenance half of the job

The repeatable method
  1. When the market de-rates an incumbent because "AI can now do what they sell," decompose the job into creation and ongoing maintenance (updating, patching, integrating, adapting to changing requirements).
  2. Estimate what share of the customer's spend and the incumbent's revenue sits in each half. AI is strongest at the first draft and weakest at continuous upkeep.
  3. If the maintenance half dominates — as it does in enterprise software — the disruption thesis is at best years out, and the de-rating is a sentiment overshoot to fade, not a terminal repricing.
  4. Wait for the confirming evidence before sizing up: one or two earnings reports that fail to show the predicted damage is the trigger, because that is when the re-rating starts and the narrative traders leave.
Here: January's panic that AI coding tools would replace corporate software at a fraction of the cost hit CRM and NOW. The article's verdict — it "hasn't happened, and isn't likely to happen," because AI codes well but maintains and updates poorly — plus a couple of strong earnings reports, has a software recovery under way with room to run.
Watch for

6. Split an "AI is helping our prices" story into mix and pass-through

The repeatable method
  1. When average selling prices rise in a shrinking unit market and AI is credited, decompose the increase into mix/premium (a genuinely better product customers pay more for) and input-cost pass-through (a higher bill of materials handed to the buyer).
  2. Pass-through carries no margin and depends on continued pricing power; mix does. Only the mix component supports a re-rating.
  3. Check whether the same AI wave is on both sides of the P&L — driving the premium feature and bidding up the component the product needs. If so, the AI story is a wash until units recover.
Here: PC prices are up and projected to climb, but the article names two causes at once — HPQ and DELL selling on-device-AI machines at a premium (mix), and memory costs "through the roof" because AI data centers ate the supply (pass-through) — into falling unit volumes.
Watch for

Methods distilled from the WSJ AI & Business newsletter (a structured digest is saved in transcript.txt) for personal study. Not investment advice. © The Wall Street Journal / Dow Jones for source material.