Malgré l'absence apparente de standard, le consensus est que le robots.txt s'applique aux moteurs d'indexation récursifs :
D'après robotstxt.org : Web Robots (also called "Wanderers" or "Spiders") are Web client programs that automatically traverse the Web's hypertext structure by retrieving a document, and recursively retrieving all documents that are referenced.
D'après google : crawler: A crawler is a service or agent that crawls websites. Generally speaking, a crawler automatically and recursively accesses known URLs of a host that exposes content which can be accessed with standard web-browsers. As new URLs are found (through various means, such as from links on existing, crawled pages or from Sitemap files), these are also crawled in the same way.
Il me semble donc que l'indexation de la tribune n'est pas concernée.
# C'est possible
Posté par papatte3 . En réponse à l’entrée du suivi robots.txt pour les outils d'archivage Web. Évalué à 1 (+0/-0).
Malgré l'absence apparente de standard, le consensus est que le robots.txt s'applique aux moteurs d'indexation récursifs :
D'après robotstxt.org : Web Robots (also called "Wanderers" or "Spiders") are Web client programs that automatically traverse the Web's hypertext structure by retrieving a document, and recursively retrieving all documents that are referenced.
D'après google : crawler: A crawler is a service or agent that crawls websites. Generally speaking, a crawler automatically and recursively accesses known URLs of a host that exposes content which can be accessed with standard web-browsers. As new URLs are found (through various means, such as from links on existing, crawled pages or from Sitemap files), these are also crawled in the same way.
Il me semble donc que l'indexation de la tribune n'est pas concernée.