crawlsignals · How we read · About, corrections and removal · Privacy · records.json

arXiv arxiv.org · reference

What it says about automated access

robots.txt:

robots.txt https://arxiv.org/robots.txt read 2026-10-03

AgentRules it follows (RFC 9309)Over the whole file
*the * group28 Disallow rules; / allowed
Googlebotits own group26 Disallow rules; / allowed; not bound by 2 of the * Disallow rules
Bingbotits own group26 Disallow rules; / allowed; not bound by 2 of the * Disallow rules

With no group of their own, these follow the * group: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, Applebot-Extended, PerplexityBot, CCBot, Bytespider, meta-externalagent.

Under RFC 9309 an agent with its own group follows only that group, so the * rules do not apply to it.

Official way in

Not found by us: no API, feed or export.

Documents

Not found by us: terms of use.

robots help page policy · not read by us: robots.txt's comments prohibit crawling, so our crawler reads nothing else from this host

Nothing here is legal advice. A mistake? Write to hello@crawlsignals.com (corrections).