Interactive Fermi model
How many tokens will analytics consume?
A scenario model for steady-state global LLM usage in data analytics. Adjust the market, workload share, and token price to see the infrastructure spend implied by your assumptions.
Public anchors
The updated model replaces the original Google run rate with May–July 2026 disclosures. These are directional anchors, not a complete census of global usage.
Sources: Google I/O 2026, Alphabet Q2 2026, and Microsoft FY26 Q3.
Workload mix
Today’s observed developer-heavy mix is separated from the 2029 steady-state hypothesis. Data analytics is the highlighted share.
Build your case
The defaults reproduce the original 2029 base case. Move any assumption to update the analytics token volume and infrastructure spend.
130Q analytics tokens per year
Forecast envelope
The original global token scenarios, retained as the model’s uncertainty range.
| Year | Bear case | Base case | Bull case | Key anchor |
|---|---|---|---|---|
| 2025 | 8–12 Q/yr | 15–20 Q/yr | 25 Q/yr | Early API and hyperscaler disclosures |
| 2026E | 35–50 Q/yr | 60–90 Q/yr | 130 Q/yr | Google alone annualizes above 38Q |
| 2027E | 80–110 Q/yr | 120–180 Q/yr | 280 Q/yr | Agent rollout and inference capacity growth |
| 2028E | 160–220 Q/yr | 280–420 Q/yr | 700 Q/yr | Falling unit costs unlock more workloads |
| 2029E | 280–380 Q/yr | 500–800 Q/yr | 1,500+ Q/yr | Steady-state enterprise agent adoption |
Why 20%
The estimate triangulates workforce size, SQL penetration, and the extra token intensity of agentic analytics loops.
Workforce ratio
Roughly 12–18M data workers versus 28.7M software engineers, with data workflows estimated at 50–80% of SWE token intensity. That implies 10–26% of all tokens.
SQL penetration
SQL reaches 54% of professional developers, while Python is deeply used for data work. The proxy implies roughly 25–35% of coding tokens come from analytics.
Agentic multiplier
Plan → inspect schema → generate SQL → execute → validate → iterate creates repeated context-heavy calls, estimated at 5–15× a single-pass interaction.
Demand expansion
Natural-language analytics removes the SQL barrier and can expand the active user base by 3–5×, while enterprise schemas add 10K–50K context tokens per query.
Assumptions & caveats
The model is deliberately adjustable because both usage and token efficiency remain unusually uncertain.
Adoption lag
Data analytics trails software engineering adoption today, but the model assumes the gap narrows materially by 2029.
High context intensity
Enterprise schemas, tool calls, result inspection, and iteration make analytical sessions structurally heavier than casual chat.
Specialized-model risk
Smaller NL-to-SQL models, caching, and reusable query plans could compress token consumption without shrinking the software opportunity.
Measurement gap
Public disclosures mix API use, consumer surfaces, internal products, and different token-counting conventions. Treat every top-down total as approximate.