The AI search era: how being found is changing, and what we are building for it
People increasingly ask AI systems instead of scanning links. How those systems choose their sources is documented, and most firms have never read it. What is changing, and what to do about it.
For twenty years, being found by a new client meant one thing: appearing in a list of links. A person typed a question into a search box, scanned a page of results, and clicked. Every assumption about how professional firms present themselves online, from the structure of their websites to the wording of their pages, grew up around that single behaviour.
That behaviour is changing. Not everywhere, not for everyone, and not overnight. But a growing share of questions that once went to a search box now go to an AI assistant, and a growing share of search results pages now begin with an AI-composed answer rather than a list. When someone asks a question about a legal problem, there is an increasing chance that the first thing they read was assembled by a machine from sources it selected, and that they may never see a list of links at all.
This piece sets out what is actually changing, how the AI systems involved decide which sources to read and cite, and what a professional firm can sensibly do about it. It reflects the research we do at PageMax, where this question is the foundation of the product.
What is actually changing
Three shifts are happening at once, and they compound each other.
First, search engines themselves are answering. Google began rolling out AI Overviews, answers composed at the top of the results page, from its 2024 developer conference onward, and has been extending them across markets and query types since (Google's announcement). For queries where an Overview appears, the answer sits above every traditional result. The links are still there, but they are no longer the first thing a person reads.
Second, assistants are becoming a front door of their own. Hundreds of millions of people now use conversational AI products weekly, and a meaningful portion of their questions are exactly the kind that used to start at a search box: what am I entitled to, how does this process work, do I need professional help, who should I talk to. When the answer to that last question includes named firms, it matters enormously which firms the system knows about.
Third, the click itself is weakening. When a person gets a complete answer, whether from an assistant or an AI summary, they often have no reason to visit the source. Web publishers and measurement firms have been documenting this pattern for several years. The practical consequence for a firm is that being read by machines and being cited in their answers is becoming a distinct channel from being visited by people, and it needs to be understood on its own terms.
Being read by machines and being cited in their answers is becoming a distinct channel from being visited by people.
How AI systems choose their sources
The selection process is less mysterious than it sounds, because the companies involved document a good deal of it. The single most useful fact is that the major AI providers do not operate one crawler. They operate several, with different jobs, and a website can welcome some while refusing others.
OpenAI, for example, documents three distinct bots. One gathers content that may be used to train future models. One builds the search index that its products use when they need to find current information. And one fetches a specific page on demand, at the moment a user asks a question that the page might answer. Anthropic and Perplexity publish similar documentation describing the same broad division of labour, and Google documents its crawlers in comparable detail, including the token that controls whether content may be used for AI training separately from whether it appears in search.
The division matters because the three jobs have very different value to the firm whose content is being read. Training ingestion contributes to a model's general knowledge, with no citation and no traffic. Search indexing and on-demand fetching are different: they are how an assistant finds, quotes, and links a specific firm's page when a specific person asks a relevant question. That is a citation, a recommendation of sorts, delivered at the exact moment someone needs it.
For the content itself, the systems reward a quality that traditional practice often neglected: passages that answer a question completely, in one self-contained place. An assistant quoting a source does not quote a page, it quotes a passage. Content written so that its key answers survive being lifted out alone, with the direct answer first and the qualifications after, is structurally easier to cite. This is not a trick. It is simply clear writing, which happens to be what both people and machines prefer.
The part most firms cannot see
Here is the uncomfortable finding from our own work. Whether an AI system can read a firm's website at all is decided in two places, and most firms only know about one of them.
The first is the site's robots file, the published policy that tells crawlers what they may read. Firms, or their web people, control this and can inspect it.
The second is the hosting layer, and this is where firms go invisible without knowing. Hosting providers and security layers increasingly ship default rules that treat AI crawlers as unwanted bots and drop their connections outright. The firm's stated policy says welcome; the infrastructure says nothing at all, because the request never reaches the site. When we moved our first client's website onto our own infrastructure, we found exactly this: the previous hosting had been silently refusing the AI crawlers while serving human visitors normally. Nothing in any setting the firm could see said so.
This default darkness is about to become more common, not less. Cloudflare, which sits in front of a substantial fraction of the web, declared its position in 2025 (Content Independence Day) and has continued to move its defaults toward blocking AI crawlers unless a site owner decides otherwise, with a further defaults change announced for September 2026. The direction is clear: the web is becoming closed to AI systems by default, and visibility to them is becoming something a site owner must deliberately choose.
The web is becoming closed to AI systems by default. Visibility to them is becoming something a site owner must deliberately choose.
A choice, not an accident
It is worth saying plainly that openness to AI systems is not automatically the right choice, and blanket openness is not what we advocate. Training ingestion, in particular, is a trade each publisher should weigh: it feeds systems that may never send anything back. There is active litigation between major publishers and AI companies over exactly this question, and reasonable people disagree.
But the choice should be made, not defaulted into. A firm that blocks everything loses the citations along with the training. A firm that allows everything donates its content without asking what it gets in return. The deliberate position sits in between: welcome the crawlers that produce citations and answers, decide consciously about the ones that only ingest, and verify, with actual evidence rather than assumption, that the infrastructure is doing what the policy says.
Very few firms are in a position to have this conversation today, because the moving parts live in documentation that solicitors have no reason to read. That is not a criticism of solicitors. It is a gap in what their providers do for them.
What a firm can check today
None of this requires a firm to become technical. It requires three questions, asked of whoever looks after the website, with actual answers rather than reassurance.
First: what does our robots file say about AI crawlers, specifically? Not bots in general. The named crawlers from the AI companies, the ones whose documentation we linked above. If the person answering cannot name them, that is itself the answer.
Second: does our hosting or security layer block any of those crawlers at the connection level, regardless of what the robots file says? This is the invisible layer described earlier. The honest test is not to read settings but to fetch the site the way each crawler does and record what happens. Any competent provider can run that test in minutes; ask to see the results, crawler by crawler.
Third: when did anyone last verify this? A setting checked once, years ago, says nothing about today. Hosting providers change defaults, security products update rules, and the crawler landscape itself is being renamed and reorganised as the companies evolve their products. Verification has a shelf life, which is why we treat it as a routine rather than an audit.
On the content side, the practical bar is also plainer than the industry makes it sound. Pages that answer one question thoroughly beat pages that gesture at many. Passages that begin with the direct answer beat passages that build suspense. Claims that carry a source beat claims that ask for trust. Dates matter: content that is visibly maintained signals a firm that is paying attention, to machines and people alike. Nothing in that list is new advice. What is new is that machines now enforce it at scale, quoting the pages that comply and passing over the ones that do not.
What we are building at PageMax
PageMax exists because we believe the answer to all of this is a system, not a checklist. Our platform runs a law firm's entire online presence: the website, the writing, the enquiries, the Google presence, and the visibility questions this piece describes, with one governing rule that everything the system proposes is approved by a named person at the firm before it happens.
On the specific question of AI-era visibility, the platform takes the deliberate position described above and makes it real. The sites we run are verifiably readable by the AI systems that cite their sources, and we check that this remains true rather than assuming it. The content the system writes is grounded in primary sources and structured so that its answers can be quoted whole, because that is what both readers and machines reward. Every factual claim in a published article carries a citation that a person can check. And the firm's own performance, including how AI systems interact with its site, is recorded as evidence rather than asserted as marketing.
The shift from links to answers is still in its early chapters, and anyone claiming certainty about where it lands is selling something. What can be said with confidence is that the mechanics of being found are changing faster than most professional firms' arrangements for being found, and that the gap between the two is where clients quietly go elsewhere.
We think the firms that do well in this era will be the ones whose presence is looked after continuously, by something that reads the documentation so they never have to. That is the product we are building, and this blog is where we will keep showing our work.