crawlsignals · How we read · About, corrections and removal · Privacy · records.json

Instagram www.instagram.com · social

What it says about automated access

robots.txt:

robots.txt https://www.instagram.com/robots.txt read 2026-10-03

AgentRules it follows (RFC 9309)Over the whole file
*the * groupdisallowed at / (line 288)
Googlebotits own group11 Disallow rules; / allowed; not bound by 1 of the * Disallow rules
Bingbotits own group11 Disallow rules; / allowed; not bound by 1 of the * Disallow rules
GPTBotits own groupdisallowed at / (line 24)
ClaudeBotits own groupdisallowed at / (line 18)
Google-Extendedits own groupdisallowed at / (line 21)
Applebot-Extendedits own groupdisallowed at / (line 12)
PerplexityBotits own groupdisallowed at / (line 27)

With no group of their own, these follow the * group: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, CCBot, Bytespider, meta-externalagent.

Under RFC 9309 an agent with its own group follows only that group, so the * rules do not apply to it.

Official way in

Not found by us: no API, feed or export.

Documents

Automated Data Collection Terms terms · named in robots.txt, line 7 of www.facebook.com · not read by us: the robots.txt of www.facebook.com prohibits crawling in its comments, so our crawler reads nothing else there

Nothing here is legal advice. A mistake? Write to hello@crawlsignals.com (corrections).