Prince: Major legitimate companies completely ignore robots.txt
“And there's some just, there's some, even some big legitimate companies that completely ignore robots.txt.”
Prince: Cloudflare working with IETF and regulators on robots.txt extensions
“And so what we've proposed, and we're working with the IETF as well as regulators is extensions to robots.txt to give it that granularity.”
Mytton: Newer bots use robots.txt to target restricted site areas
“And, but there are newer bots that are ignoring it or even sometimes using it as a way to find the parts of your site that you don't want it to access. And they will just do that anyway.”
Tytunovich: Self-Declared AI Signatures Will Create Major Security Loopholes
“There's a very high likelihood, in my opinion, that the proprietors behind the main AI agents, like OpenAI, like Google, would find a way to declare an AI agent as one operating on their behalf. That being said, that is the very definition of a loophole. Such …”
Morin: Anthropic rotates bot names to bypass web scraper blocks
“And suddenly you've got this problem that I was reading about this week where Anthropix changing the name of their bots. They're creating 10 more bots. Developers are having to, like, you know, block this bot and that bot and this name and that name. It's a bu…”
Srinivas: Perplexity will not crawl websites that block web crawlers
“Yeah, I mean, like, if they don't want to be crawled, we shouldn't crawl them. You know, that's just being how the internet has worked. Like, you know, people have the robots, the text file, and like, you know, if they don't let you do it, then you should just…”