Docket / Fix it / Content audit

How to do a content audit

If you searched for an Ahrefs content audit, the first useful thing to know is that Ahrefs does not sell one. There is no product of that name. What exists is a job — going through the pages you have published and deciding which to keep, improve, merge or remove — and several tools that each do a slice of it. Knowing which slice saves you paying for the wrong one.

What a content audit actually answers

Four questions, and they are separable:

The last one is worth doing first, because a brilliant page that a crawler cannot read is a worse problem than a mediocre one that it can, and it is far cheaper to fix.

What Ahrefs' real tools do for this

Content Explorer is described by Ahrefs as a way to "find top-performing content, link prospects, and media opportunities in your niche" — content ideas on a topic. It is an index of what other people have published. That is genuinely useful for deciding what to write next, and it is the tool whose name most resembles the phrase people search for. It is not looking at your site.

Site Audit is the one that crawls your own pages — Ahrefs' own summary is "audit & optimize your website". It is a technical crawler: status codes, duplicate titles, broken links, the mechanical fourth question above.

Site Explorer holds their backlink and organic-traffic data, which is where "which of my pages actually earn anything" comes from.

So the job is split across at least two products and the judgement is still yours. Nobody is hiding a Content Audit button — the work simply does not live in one place. If you are weighing the crawling half specifically, that comparison is on Docket vs Ahrefs Site Audit, with dated prices read from their pricing page.

The order that wastes least time

1. Get the list. Crawl your own site and export every indexable URL. Not the sitemap — the sitemap is what you claim you have. The crawl is what a crawler can actually reach, and the difference between those two lists is itself a finding.

2. Separate thin from invisible. These look identical in a spreadsheet and have opposite remedies. A genuinely thin page needs writing or removing. A page whose body arrives via JavaScript needs server-rendering — writing more would not help, and deleting it would be a mistake. Docket labels a word count as a count of the HTML as served when a page carries a framework's hydration marker, precisely because the two get confused; the JavaScript SEO audit procedure is how you tell them apart on a single URL.

3. Find the pages competing with each other. Duplicate and near-duplicate titles are the cheapest signal — see fixing duplicate title tags. Where two pages genuinely serve one intent, merge and redirect rather than rewriting both.

4. Check what is claiming to be canonical. A page that points its canonical elsewhere has opted out of ranking, sometimes by accident. Canonical tags are a hint, not an instruction, and the common failure is a template applying one site-wide.

5. Only now, judge quality. With the mechanical problems cleared you are reading a much shorter list, and you are reading it about pages that can actually rank.

Index pages are not thin pages

A category listing, an A–Z, a tag or author page — its job is to point elsewhere, and a low word count is normal. Every automated content audit flags these, and the advice it gives is wrong three ways: do not write copy on it, do not merge it, and do not delete it, because other pages link to it. Judge it on whether the list is worth having and whether anything links to it. This is a limit worth knowing about whatever tool you use, and Docket says so in the finding itself rather than leaving you to discover it.

What Docket does for the mechanical half

Docket runs 97 checks across 13 areas on your machine, and the content lane covers the fourth question: pages with almost no body text, pages in the thin band above them, archive pages that list almost nothing, utility pages that should not be indexed at all, titles duplicated across pages, and pages whose content is not in the HTML the server sends. It reports which of those it measured on the served HTML rather than on a rendered page, because that distinction changes the remedy.

What it does not do: tell you whether a page is good. It has no index of what competitors published, no keyword volumes, and no opinion about what people search for. If the question you actually have is "what should I write next", Content Explorer or a keyword tool is the right purchase and Docket is not. The full list of what is and is not covered is what Docket checks, and the wider procedure is the technical SEO audit.

Doing it more than once

A content audit is usually run as a one-off and then not again until the next reorganisation, which is why the same problems reappear. The mechanical half is worth re-running on a schedule — site monitoring is that, and it is a different job from a crawl in that it tells you what changed rather than what is true today. If you want the server's own record of which pages a crawler actually fetched, that is log file analysis.

And if the reason for the audit is that traffic arrives and does nothing, the content is probably not the constraint — a conversion audit asks a different question of the same pages.

Download Docket

Common questions

Does Ahrefs have a content audit tool?

Not under that name. Checked on their pricing and Site Audit pages, Ahrefs' named tools are Site Explorer, Keywords Explorer, Site Audit, Rank Tracker and Content Explorer. The work people mean by “content audit” is split across Site Audit for the technical side and Content Explorer for what others have published.

What is the difference between Content Explorer and Site Audit?

Content Explorer indexes what other people have published, for finding ideas and link prospects in a niche. Site Audit crawls your own website. If the question is about your existing pages, Site Audit is the relevant half.

What should a content audit start with?

A complete list of your indexable pages, taken from a crawl rather than from your sitemap. The sitemap says what you claim to have published; the crawl says what a crawler can reach, and the gap between the two is a finding in itself.

Why is a short page not automatically a problem?

Because index pages are short on purpose. A category listing, an A–Z or a tag page exists to point elsewhere, and the usual advice — write more, merge it, delete it — is wrong for all three. Judge it on whether the list is worth having and whether anything links to it.

Can a tool tell me whether my content is good?

No, and one that claims to is scoring a proxy. Tools are reliable on the mechanical questions — is it indexable, is it duplicated, is it in the HTML — and those are worth clearing first, because they are cheap and they gate everything else. The judgement about depth stays yours.