Free tool
Technical SEO Audit
Crawl a site in batches and list the technical problems found, page by page, with severity and evidence.
The check could not be completed
Audit results
What this audit could not check
Issues found
| Issue | Severity | Page | Detail |
|---|
Pages crawled
| Page | Status | Depth | Issues |
|---|
What This Tool Does
The audit crawls one domain, page by page, and reports the technical problems it can prove from the responses it received: pages that did not answer, redirect chains, broken internal links, duplicate and missing metadata, canonical conflicts, pages blocked from indexing, thin pages, oversized documents, structured data that will not parse, and pages that almost nothing links to.
It runs in small batches so a slow site cannot exhaust a PHP request. Discovered pages, checked pages, errors and the current state are shown while it works, and you can stop it at any point and keep the results collected so far.
How to Use It
Enter the address the crawl should start from, choose a page limit and a depth, and start the audit. If you have a single sitemap file, add it: the audit can then tell you which crawled pages are missing from it.
- Start small. A limit of ten pages and a depth of two will usually surface template-level problems, and every template-level problem is repeated on every page built from it.
- Filter by severity first, then by issue type. One issue type repeated across twenty pages is one fix, not twenty.
- Export the CSV before you start fixing things. It gives you a before-and-after record that the tool itself does not keep unless you save the report.
- Rerun the audit after a deployment rather than after every edit. The value is in the comparison.
How the Tool Calculates or Retrieves Data
Each page is fetched once, from this website's server, over HTTP or HTTPS only. Addresses are validated before the request and revalidated after every redirect, private and loopback destinations are refused, response sizes are capped, and the crawler stays on the starting domain.
Addresses are normalised before they are queued: fragments are dropped, relative links are resolved, a base tag is honoured, and tracking parameters that do not change the page are ignored so the same page is not crawled repeatedly. Calendar, cart, logout, admin and endless-pagination patterns are skipped, because they generate addresses forever.
Every issue is derived from a response the crawler actually received. Where a check cannot be completed — sitemap membership without a complete sitemap, for instance — it is listed under what the audit could not check rather than reported as a pass or a failure.
Understanding Your Results
Errors are responses and directives that are wrong as they stand: a page that did not answer, a broken internal link, a canonical pointing at another domain, a page you did not mean to block. Warnings are departures from convention that are often deliberate. Notices are observations worth knowing.
The few-internal-links and crawl-depth issues describe the site the crawler saw, within the limit you set. A page can look isolated simply because the crawl stopped before reaching the section that links to it, which is why the limit is shown alongside the results.
Redirect issues are worth reading in order. One redirect is normal; a chain of three is a link that has been rewritten three times and never updated at the source.
Frequently Asked Questions
Because a shared host will end a PHP request long before twenty-five pages have been fetched from a slow server. Each batch is a separate request, which also means the progress you see is real rather than an animation, and cancelling actually stops the work.
No, and it says so rather than guessing. A crawl only finds what is linked. Adding a single sitemap file lets it check the reverse — crawled pages that the sitemap omits — which is a different and answerable question.
The count is of links seen during this crawl, within the page limit you set. Raise the limit or start the crawl from the section that links to it, and the count will change.
By default, yes. You can turn that off for a site you control, which is useful when a staging environment blocks everything, but it should not be turned off for a site that is not yours.
It expires on its own. Job state is temporary, belongs to the account that started it, and is removed automatically; nothing is left running in the background.