⚠ AI Levels the Cold StartModerate threat

Alphabet (Google) (GOOGL) — threat to the moat

A model trained on the whole web arrives already knowing — pretraining softens the wall that made the flywheel decisive.

The data flywheel's power rests on the cold-start problem — the idea that a rival begins blind, without the two decades of click behavior that make Google's answers feel prescient. The danger is that modern AI weakens that assumption, because a large language model trained on the whole of the web arrives already knowing a great deal about how to answer questions, without ever having watched a billion users click. Pretraining is, in effect, a way to start warm — and early 2025 showed how warm, when DeepSeek's R1, trained largely on public data at a fraction of frontier budgets, matched the leaders' benchmarks1.

How far the best U.S. models led Chinese ones (points)End of 2023End of 2024MMLU17.50.3MMMU13.58.1MATH24.31.6HumanEval31.63.7Percentage-point gaps on four benchmarks. Stanford HAI, AI Index 2025
Labs working far from Google's query logs closed most of the gap in a single year — the warm start, measured.

This matters because it attacks the specific mechanism that made search un-copyable. If a capable answer engine can be built from a foundation model plus the public web, then the proprietary advantage of Google's accumulated query log — the thing no rival could reproduce — counts for less than it did, and a well-funded newcomer can offer answers that feel competitive from its first day rather than flat and generic.

The shield here is that behavioral data still matters enormously for the things pretraining cannot supply: what people click right now, which fresh results satisfy, what the latest intent looks like, and how to rank the living, changing web. Google also has that click history and the models, so if the game becomes model-plus-data, it holds both hands while a challenger holds only one.

A moderate worry. AI genuinely erodes the cold-start moat that made the flywheel so protective, and a newcomer can now start far warmer than before — but the real-time behavioral signal Google alone possesses remains valuable, and Google brings both the data and the models to the new contest. The flywheel is less exclusive than it was, not stopped.

References
  1. ReportedDeepSeek-R1 (Jan 2025): frontier-level results from largely public data.
    DeepSeek-R1 (January 2025) — an open-weight model trained largely on public data at a fraction of frontier budgets, matching leading models on key benchmarks — January 2025 · publ. January 2025 · source ↗
Sources
Generated September 16, 2026