Server rendered HTML
Many agents run no JavaScript. What only comes into being in the browser does not exist for them.
A growing share of requests comes neither from people nor from classic crawlers, but from agents looking something up on behalf of a person.
Categories
The classic crawler indexes regularly and broadly. It collects for later and takes no interest in any single question.
The training collector gathers content for building models. Its visit takes effect months later and its impact can barely be measured.
The agent on assignment fetches exactly one page right now, because somebody wants to know something right now. Its visit is the most valuable and the least patient: anything slow to load, or appearing only after scripts run, drops out of the answer.
In robots.txt those three can be treated separately, and that is exactly what we recommend. A blanket rule treats the visit that brings you customers the same as the one that trains on you.
Requirements
Many agents run no JavaScript. What only comes into being in the browser does not exist for them.
Time limits are tight. A slow page gets dropped rather than waited for.
Price, availability, opening hours. An agent extracts, it does not interpret charitably.
Bot defences, consent banners and login walls end the visit. Often unnoticed.
Practice
The simplest test is a fetch without a browser. If the essential details are missing from your page's raw text, they are missing for the agent too.
The second test runs in the server log. The known identifiers of the AI providers turn up there, and their ratio to human requests is a number almost nobody looks at.
The third test is the answer itself: ask questions, look at the sources, check whether the cited detail is right. That is the work we do systematically in the audit.
Common questions
Next step
One conversation, 30 minutes, no sales pressure. You leave with an honest read on whether this is worth the effort for you.