1) training data (common crawl as one example web data source), and
2) live data optionally retrieved at runtime.
My comment was about 2), and that part runs via a search engine (Bing?), at least if you look at how ChatGPT does it.
So, why would you use Google as a tool or search target when you can, in some combination, go direct to the website (or whatever the target data endpoint is) yourself as an LLM provider to retrieve the most recent data or rely on your own "hot cache" of that data that was crawled recently but said data is not stale enough warranting a live web crawl to retrieve and present to the user or AI agent? Is this capability to perform retrieval from a data source in real time not similar to an AI agent?
Broadly speaking, I'm just spitballing on the concept of "You must use a search engine for an LLM to return 'live-ish' results" as I think we're directionally headed to where that isn't the case.
I'm not sure frontier labs do it, but imo fetching live data requires fetching both the search engine (1) and the website/page (2). The search engine gives you search results + content snippets (potentially stale), the website/page gives you the actual content fresh from the source.
The tradeoff I see here is between liveness/staleness and cost (hitting an index if of course cheaper and faster than querying live websites again).