goguides.com trust favicon wikipedia.org trust favicon mozilla.org trust favicon
GoGuides publishes independent web evidence, provenance, and machine-readable records.
Log In Navigation
Crawler control

Does robots.txt control how AI uses my website?

Robots.txt is primarily a crawl-access convention for compliant automated clients. It is useful, but it is not a privacy system, an identity system, a universal indexing control, or a complete content-use license.

The short answer

A standards-compliant crawler can consult robots.txt before fetching a URL and choose whether to crawl based on the rules that apply to its declared user-agent. That is an access preference. It does not prove what every automated system will do, and it does not control copies of content obtained elsewhere.

Four different problems

Crawl access

Robots.txt can tell compliant crawlers which public paths they are asked not to fetch.

Indexing and display

Search or answer products may use separate controls and policies for indexing, snippets, citations, or display. Crawl blocking is not identical to removal from every index.

Privacy and security

Private information should be protected with authentication and access control. Listing a path in robots.txt does not make that path private.

Licensing and reuse

Permissions for copying, retrieval, training, commercial use, attribution, or redistribution are legal and contractual questions that robots.txt alone does not settle universally.

Typical examples

User-agent: ExampleBot
Disallow: /private-section/

User-agent: *
Disallow: /temporary/

# An empty Disallow allows crawling:
User-agent: *
Disallow:

The exact behavior still depends on whether the client identifies itself accurately and chooses to honor the convention.

Common mistakes

Using robots.txt as privacy protection

Do not expose secrets and rely on crawler politeness. Use real access controls.

Assuming a User-Agent proves identity

A string can be copied. Network or cryptographic evidence is a separate question.

Equating crawling with indexing

Fetch behavior, indexing, retrieval, ranking, citation, and model-development use are different events.

Equating access with permission for every use

Being technically reachable is not the same as resolving every licensing or source-use question.

Where GoGuides fits

GoGuides treats crawler access as one observable dimension among many. Its public evidence layer can also preserve URL state, redirects, selected fingerprints, structured data, links, AI-policy indicators, chronology, ownership-verification context, and supported machine-identity evidence.

GoGuides evidence does not override a publisher's robots rules, and robots rules do not replace provenance, identity, factual verification, or a consuming system's own source-use policy.