DocumentationDoorFall SearchWebsite Sources & Crawling

Test website for crawling

What the pre-add crawler test checks and how to interpret its result.

Website URL

Enter the normal public homepage, for example https://example.com.

Click Test website. This does not save the website, queue URLs or change the index.

Suitable

DoorFall reached the website, robots.txt allows the tested URL, normal HTML was returned and the homepage contained enough indexable text.

A Suitable result does not guarantee that every page can be crawled. Individual URLs may have different robots rules, permissions, size or content.

Use caution

The website responded, but DoorFall found something that may reduce crawl quality. Examples include a temporary timeout, HTTP 429 rate limiting, a server error, unusual content type or insufficient searchable text.

Retry temporary problems before deciding that a site is unsuitable.

Not suitable

DoorFall found a clear blocker such as robots.txt denying access, HTTP 401/403, or a missing homepage.

TIPDoorFall respects the website owner's crawler rules. Do not try to work around robots.txt or access restrictions.
Last updated August 21, 2026