Your estate agency may have years of experience, detailed market knowledge and an extensive transaction history. But can the AI systems your potential customers use actually access the information that demonstrates that expertise?
LLM crawlers are automated systems that help AI platforms discover and retrieve information from websites. Understanding which crawlers can access your website, what they are looking for and what might be blocking them is an increasingly important part of managing your agency's digital footprint.
For estate agencies, the first step is understanding how these systems work and ensuring that important public information about the business is technically accessible.
What are LLM crawlers?
LLM stands for Large Language Model, the technology behind AI platforms such as ChatGPT and Claude.
LLM crawlers are automated systems used by AI providers to discover and collect information from websites. They operate in a similar way to traditional search engine crawlers, although their purposes can differ.
Some collect information that may contribute to AI model training. Others support AI-powered search, helping platforms find relevant information when users ask questions.
There are also systems that retrieve individual web pages in response to a specific user's request.
These distinctions are important because website owners can often manage access for different purposes independently.
An estate agency might, for example, choose to make its website available to AI search crawlers while restricting access for model-training crawlers.
The objective is to understand what each system does and make informed decisions about how your agency's information is accessed.
Which LLM crawlers should estate agencies know about?
The major AI providers operate different crawling and retrieval systems. Understanding their names and purposes is a useful starting point for any website audit.
| AI provider | System | Primary purpose |
|---|---|---|
| OpenAI | OAI-SearchBot | Discovers websites for ChatGPT searc |
| OpenAI | GPTBot | Collects information that may be used for model training |
| OpenAI | ChatGPT-User | Retrieves information for certain user-initiated requests |
| Anthropic | Claude-SearchBot | Supports Claude's web search functionality |
| Anthropic | ClaudeBot | Collects information that may contribute to model training |
| Anthropic | Claude-User | Retrieves content in response to user requests |
| Perplexity | PerplexityBot | Discovers information for Perplexity search |
| Perplexity | Perplexity-User | Retrieves pages in response to user requests |
| Googlebot | Crawls websites for Google Search, including its AI search features |
These systems and their functions are documented by OpenAI, Anthropic, Perplexity and Google.
Google also provides Google-Extended. This is not a separate crawler, but a control that allows website owners to manage certain uses of their content for Gemini model training and grounding. Its settings do not affect inclusion in Google Search.
One important distinction is that user-triggered retrieval systems do not always follow the same access rules as automated crawlers. For example, OpenAI states that robots.txt rules may not apply to certain requests made through ChatGPT-User.
For estate agencies, this means there is no single setting that controls every form of AI-related access.
Can these crawlers access your estate agency's website?
An agency's website might work perfectly well for human visitors while presenting difficulties for automated systems.
Consider an estate agency specialising in family homes in Walthamstow.
Its website contains detailed area guides, agent profiles, property listings and local market reports. Each helps communicate the agency's experience and knowledge of its market.
However, an important market report might be blocked by the website's security system. An agent profile might depend on technology that a particular crawler cannot process. A technical setting might prevent relevant search crawlers from accessing entire sections of the website.
The agency has published valuable information, but some automated systems may be unable to retrieve it.
These issues are not necessarily obvious during normal website maintenance.
That is why crawler accessibility deserves specific attention.
The five areas your website developer should check
1. Crawler permissions
Your website's robots.txt file communicates which pages and sections compliant crawlers are permitted to access.
Check that your settings reflect which AI systems you want to permit. An accidental restriction could prevent relevant search crawlers from retrieving important content.
2. Website security
Security systems protect your website against malicious automated traffic, but they can sometimes block legitimate crawlers.
Check whether recognised AI systems are being challenged or denied access, and verify their identities before adjusting security rules.
3. Important website content
Can automated systems retrieve the information that demonstrates your agency's expertise?
Agent profiles, area guides, service pages and market reports are sensible starting points.
Information that depends heavily on JavaScript should be tested rather than assumed to be accessible to every crawler.
4. Website structure
Can machines identify what your agency specialises in, which locations it serves and who provides its services?
Clear website architecture, accessible text and appropriate structured data can help communicate this information.
5. Technical errors
Are important pages returning errors, redirecting incorrectly or delivering incomplete content?
These problems can prevent crawlers from retrieving information even when the agency intends to make it publicly accessible.
Addressing them helps establish a more reliable technical foundation.
How can you tell whether LLM crawlers are visiting your website?
Traditional website analytics focus primarily on human visitors and their behaviour.
To understand crawler activity, you generally need access to your website's server logs or equivalent records from your hosting or security provider.
These can help establish which automated systems are requesting information, which pages they are visiting and whether those requests are successful.
However, identifying a crawler by its name alone is not sufficient.
User-agent names can be imitated, so website administrators should use the verification methods published by the relevant providers, where available.
Both OpenAI and Google publish information that can help website owners verify crawler traffic.
Once legitimate traffic has been identified, patterns can be investigated.
For example, an agency might discover that its homepage is accessible but its market reports are consistently blocked. That gives its website developer a specific problem to address.
The objective is not simply to increase crawler traffic. It is to establish whether relevant systems can successfully access important information.
What information should estate agencies make accessible?
Resolving technical barriers is only part of the process.
Estate agencies should also consider whether the information available on their websites accurately reflects their capabilities.
For example, an agency may have extensive experience selling prime residential property in a particular area, but its website might provide little information about that specialism.
Making the website technically accessible will not address that information gap.
The agency also needs to communicate its expertise clearly, supported by appropriate evidence.
This may involve improving agent profiles, explaining specialist services in greater detail, publishing relevant market knowledge or creating clearer relationships between the agency's people, locations and expertise.
Structured, machine-readable information can provide additional context where appropriate.
However, the aim is not to publish every piece of information the business possesses.
Confidential client records, private transaction information and commercially sensitive CRM data should remain protected.
The priority is to make accurate, relevant public evidence of your agency's expertise accessible and clearly organised.
How Schutle approaches LLM crawler accessibility
At Schutle, we view crawler accessibility as part of an agency's wider digital infrastructure.
Understanding which systems can access your website helps identify technical restrictions and establish whether important information is available to relevant automated systems.
But accessibility and information quality must be assessed separately.
An agency may have a technically accessible website that fails to communicate its genuine expertise. Another may have detailed, credible information that is difficult for automated systems to retrieve.
Each situation requires a different response.
Schutle's approach focuses on identifying these gaps and strengthening the digital footprint that represents an agency's services, people, expertise and market knowledge.
Crawler accessibility provides an important technical foundation. The wider objective is ensuring that the information available about the business accurately demonstrates what it does and where it has genuine expertise.
Is your agency's website accessible to AI crawlers?
Understanding which systems can access your website is a useful starting point for identifying technical barriers and strengthening your agency's digital footprint.
Schutle helps estate agencies understand how their businesses are represented across AI platforms and identify opportunities to improve that representation.
