Crawl access
Robots.txt can tell compliant crawlers which public paths they are asked not to fetch.
Robots.txt is primarily a crawl-access convention for compliant automated clients. It is useful, but it is not a privacy system, an identity system, a universal indexing control, or a complete content-use license.
A standards-compliant crawler can consult robots.txt before fetching a URL and choose whether to crawl based on the rules that apply to its declared user-agent. That is an access preference. It does not prove what every automated system will do, and it does not control copies of content obtained elsewhere.
Robots.txt can tell compliant crawlers which public paths they are asked not to fetch.
Search or answer products may use separate controls and policies for indexing, snippets, citations, or display. Crawl blocking is not identical to removal from every index.
Private information should be protected with authentication and access control. Listing a path in robots.txt does not make that path private.
Permissions for copying, retrieval, training, commercial use, attribution, or redistribution are legal and contractual questions that robots.txt alone does not settle universally.
User-agent: ExampleBot Disallow: /private-section/ User-agent: * Disallow: /temporary/ # An empty Disallow allows crawling: User-agent: * Disallow:
The exact behavior still depends on whether the client identifies itself accurately and chooses to honor the convention.
Do not expose secrets and rely on crawler politeness. Use real access controls.
A string can be copied. Network or cryptographic evidence is a separate question.
Fetch behavior, indexing, retrieval, ranking, citation, and model-development use are different events.
Being technically reachable is not the same as resolving every licensing or source-use question.
GoGuides treats crawler access as one observable dimension among many. Its public evidence layer can also preserve URL state, redirects, selected fingerprints, structured data, links, AI-policy indicators, chronology, ownership-verification context, and supported machine-identity evidence.