For site operators

You found us in your logs.

404-not-found reads a documentation site the way an AI coding agent would, and reports back where such an agent gets stuck. We only run when somebody pastes a URL and asks us to. We are not a search crawler and we do not index anything.

How to identify us

Every request carries this User-Agent:

Mozilla/5.0 (compatible; 404-not-found Gauntlet/1.0; +https://404-not-found.nanocorp.app/agent)

What we request

  • One read of the URL we were given, then up to 10 pages linked from it, plus a short list of common documentation paths.
  • robots.txt, llms.txt, agents.md and an OpenAPI document, if they exist.
  • Around 30 GET requests in total, once, over about 3 minutes. We never POST, never submit a form, never log in, and never follow a link that changes state.

How to stop us

Disallow the token below in your robots.txt. We read it before anything else and we honour it — if it disallows us at the root, we stop and refund nothing because we never charge for a run we cannot do.

User-agent: 404-not-found
Disallow: /

How to allow us

If a Gauntlet was blocked and you want it to run, allow the token at the root, or allow the docs path only:

User-agent: 404-not-found
Allow: /docs/
Disallow: /

A WAF or bot-management rule may also be refusing us with a 403 before robots.txt matters. In that case allow the User-Agent above, or allow it for your docs hostname only.

What we do with the text

  • Page text is read once, used to produce one report, and treated as untrusted data throughout. We never execute anything we read.
  • A report is private to whoever ran it. It becomes public only if that person explicitly publishes it.
  • We do not train models on your pages and we do not resell the text.

Something still wrong?

Send us the hostname and a timestamp and we will tell you exactly what we requested.

404-not-found@nanocorp.app