Agentic AI's Compute Shortage: A Forecast Without a Receipt

WooPanda
Altcoins
Gavin Baker's forecast is precise: 500,000 agentic AI users today, 100 million tomorrow, and insufficient compute for both. That is not a prediction; it is a thesis without a balance sheet. The claim travels well in headline form, but under ledger review it collapses into a single unverifiable dependency: the assumption that agentic workloads will scale 200x while compute efficiency stands still. Hype evaporates; receipts remain. This piece parses the claim, tests it against known technical constraints, and asks a question Baker's narrative does not answer: where exactly is the shortage? Based on my audit experience, a forecast that cannot specify its own bottleneck is not a forecast. It is a marketing memo. Let me establish context. The original quote, circulated as a flash news item, rests on a simple chain: agentic AI has 500,000 users today, will have 100 million tomorrow, and the compute needed for both already exceeds supply. The implied solution, floated by some in the same conversation, is orbital compute — data centers in space, powered by solar arrays, cooled by radiative surfaces, and connected to Earth via ground stations. The sequence sounds logical. Agentic AI requires more inference per task. Inference requires more compute. More compute requires more capacity. And if terrestrial capacity is exhausted, move the data center upward. The problem is that each link in that chain is qualitative, not quantitative. Baker gives no token-per-agent metric. No inference-cost curve. No utilization rate for existing data centers. No timeline for the alleged 100 million users. No comparison between current GPU deployment and projected demand. The word "tomorrow" is doing a lot of work in a forecast that cannot be audited. Core teardown starts with the only verifiable part of the claim: agentic workloads consume more compute per task than conversational AI. That part is true. An agent does not answer a prompt and stop. It plans, calls tools, reads responses, iterates, retries, and parses partial failures. A single browser-operating task can trigger dozens of LLM invocations. A coding agent can generate hundreds of candidate patches before testing one. In my own analysis of deployed agent frameworks, a typical multi-step task consumes one to two orders of magnitude more tokens than a simple Q&A session. This is not an edge case; it is the architecture. The compute curve shifts from a single spike to a sustained plateau. That structural difference supports the broad claim that agentic AI will stress inference capacity. But broad support is not sufficient for a 200x user forecast. Let me test the user number. 500,000 users today, if accurate, is already a niche population. The products that count as agentic — browser-use assistants, autonomous coding tools, and early "computer use" demos — are not mass-market applications. They are developer tools and pilot deployments. Reaching 100 million users means not only a product-market fit miracle but also a behavioral shift: non-technical users trusting autonomous agents with money, credentials, and irreversible actions. That shift requires solving trust, security, and accountability problems that have no current solution. Token economics matter, but user trust is the silent gating factor. The forecast assumes away the hardest part of the adoption curve. Now examine the compute shortage claim. Baker says there is not enough compute for today's 500,000 users. If true, that is a glaring inefficiency, not a capacity limit. Current AI data centers run at varying utilization rates, with significant idle capacity during off-peak hours. Before declaring a global compute shortage, one must demonstrate that existing capacity is saturated. The public evidence shows the opposite: major cloud providers continue to add regions, GPU clusters sit partially idle between training runs, and inference engines are improving throughput per chip at roughly 20-30% per generation. The shortage narrative is convenient for vendors selling accelerators and orbital infrastructure, but it is not a measured fact. Volatility is not risk; opacity is. When a forecast depends on unmeasured utilization, the forecast itself becomes an opacity play. Orbital compute deserves special attention because it is the most concrete, most fragile part of the narrative. The engineering constraints are brutal. Spacecraft radiate heat only via thermal radiation; there is no convection or conduction in vacuum. High-density GPUs produce hundreds of watts per chip, and a rack of GPUs produces tens of kilowatts. That heat must be rejected into deep space via radiator panels. The radiator mass scales with heat load, and mass scales with launch cost. Launch economics have improved, but still, a single orbital data center would require dozens of heavy-lift launches for the computing payload alone, plus additional launches for radiators, power systems, and redundant links. Maintenance in orbit is a second-order problem. On Earth, a technician can swap a failed GPU in minutes. In space, a failed unit becomes orbiting e-waste. Ground-station bandwidth adds another ceiling: orbital platforms are not connected by dark fiber but by radio links, with latency determined by altitude and throughput limited by link budget. A low-Earth orbit constellation could theoretically serve niche users, but a global 100-million-user agentic workload is absurd. The narrative treats orbital compute as a scalable alternative to terrestrial data centers. It is not. It is a long-duration research experiment. The deeper flaw, however, is not engineering. It is incentive structure. Baker's forecast arrives in a bull market for AI infrastructure, where every capacity constraint is translated into a demand signal for more capital expenditure. This is the same pattern I saw in the 2017 ICO era: a narrative of scarcity justified tokens, not products. In 2020, liquidity mining APY disguised user acquisition subsidies. In 2025, "compute shortage" disguises a need to justify multi-billion-dollar data center investments. Ledger balances do not lie; they only wait. When the next quarterly earnings call arrives, the question will not be whether agentic AI users grow, but whether inference revenue grows faster than hardware depreciation. Now the contrarian angle. The bulls are not entirely wrong. There is a real, measurable shift in inference demand driven by agents. My own audits of tool-calling workflows show token consumption per productive output increasing every quarter. Some workloads are genuinely compute-bound, particularly those involving long-context reasoning and iterative code execution. The agents that do real work — executing multi-step trades, generating audit reports, monitoring on-chain positions — consume far more compute than chat. And the market for such agents, while small today, is growing from a non-zero base. The direction of the claim is defensible. The magnitude is not. What the bulls miss is adaptation. Compute supply is not a static ledger. Chip design is shifting toward inference-optimized architectures, with far lower energy per token. Model compression — quantization, distillation, and speculative decoding — is reducing the compute required per agent step. Distributed inference across personal devices is still a niche, but it is advancing faster than orbital launch manifest schedules. If user growth to 100 million takes five years rather than one, terrestrial efficiency gains will absorb much of the demand. The "not enough compute" claim assumes the denominator stays fixed while the numerator grows. In every previous hardware cycle, the denominator improved faster than the numerator expanded. This cycle does not show structural divergence from that pattern. The more interesting contrarian point is that the bottleneck for agentic AI may not be compute at all. It may be reliability. An agent that fails 10% of tasks at 500,000 users produces a manageable support load. At 100 million users, the same failure rate produces 10 million unresolved tasks per day. That is a customer-service catastrophe, not a GPU shortage. The compute narrative conveniently masks the more uncomfortable truth: deployed agents are not reliable enough for mass adoption. Focusing on orbital compute distracts from the mundane, unglamorous work of building evaluation suites, adversarial testing, and state rollback mechanisms. The industry would rather spend billions on silicon than spend months on robustness. That is not an engineering choice; it is a narrative choice. Accountability is the real takeaway. If Gavin Baker has proprietary data on agentic user counts, token consumption, and capacity utilization, he should publish the methodology. A forecast that cannot be replicated is not a forecast. It is a positioning statement. I have spent years auditing projects where founders claimed impossible growth and impossible throughput; the metric that mattered was always the same: verifiable receipts. User counts can be gamed. Compute shortages can be manufactured by underutilization. Orbital data centers can be rehearsed in press releases. But the ledger balances do not lie; they only wait. When the next wave of AI infrastructure financing is closed, the question for every limited partner should be simple: show me the utilization curve. This is not a call to abandon agentic AI. The technology is real. The compute demand is real. The direction is real. But a 200x user forecast with zero methodological disclosure is exactly the kind of narrative that makes institutional investors suspicious. In a bull market, euphoria masks technical flaws. The audit eye cuts through. Let the agents run, let the users grow, and let the chips burn — but let the data be public. For now, the only verifiable fact is that 500,000 users, 100 million users, and orbital compute were mentioned in the same sentence without a single honest ratio. That is not analysis. That is vapor. And as I have written before, volatility is not risk; opacity is.