After fifteen years of the same audit, an old idea of mine stopped waiting

The SEO tool interfaces haven't changed much in fifteen years.

You run an audit. You get a list. Pages with missing titles. Pages with duplicate meta descriptions. Pages returning 404s. Pages that load too slowly. Crawl depth too deep. Internal links pointing nowhere. The list is long. It's also largely useless on its own.

Not because the data is wrong — the data is usually correct. But because a list of pages with errors is the beginning of a conversation, not the end of one.

Here's the truth the list can't tell you: whether an error matters. Whether fixing it will move anything. Whether the site has bigger structural problems that make the individual errors irrelevant. Whether the business is heading in a direction that changes which problems are actually worth solving.

A page with a missing title on a product category that's being deprecated next quarter is not a priority. A page with technically correct metadata but content that no longer reflects what the company sells is a much bigger problem that doesn't appear on any error list. A crawl depth issue on a site where the deepest pages are the highest-converting ones is a nuance that the software has no way to surface.

None of that is a bug. The audit tooling category was built around a specific assumption: that technical correctness is the goal. It's a reasonable proxy, most of the time. But it's a proxy, not the thing itself — a crawler has no way to know which errors sit on pages that matter to the business and which don't. That's not a gap a better crawler closes. It's the ceiling of what the category was ever designed to do.

Agencies know that ceiling exists, and for years the industry has mostly chosen to build the engagement below it rather than push past it. A deck full of crawl errors is easier to sell than a considered argument about the client's business. It looks thorough. It fills the deliverable-shaped hole in the retainer. It maps cleanly onto software that already produced the list — no analyst had to think hard, no strategist had to understand the sector, no one had to spend a day reading the site the way a customer would.

The engagement can be automated all the way to the invoice. Which is exactly the problem. An audit that a piece of software could deliver on its own doesn't need the agency in the middle. What clients think they're paying for is judgement. What they're getting is a filtered export — one they could have pulled themselves with a licence and a login.

The harder questions are about the business. What are they trying to do? What does the competitive landscape look like? What does the data actually show about user behaviour, not just crawl behaviour? Where are the structural decisions — architecture, content strategy, internal linking logic — that are shaping performance in ways that a technical audit will never see?

Answering those questions requires something the existing tools don't provide: content intelligence that understands what pages are actually about, not just how they're formatted. A system that can read a site the way an experienced analyst reads it — with context, with pattern recognition, with the ability to connect technical signals to business meaning.

That's the gap. I've been looking at it for years. In December I restarted building into it, removing some dust from a very old project of mine.

The idea itself isn't new. I first built it in 2013 — more than a sketch, actually. Something close to finished, nearly ready to go to market. A series of unfortunate family circumstances meant it stopped being a priority, and it did what unfinished projects do: it waited.

What's changed is all three. The NLP and machine learning tooling that would have required a research team in 2013 now runs on a laptop overnight. The years of doing this work at scale have built the domain knowledge to know what actually matters. And somewhere in the space between those two things, the time appeared.

What's on screen now isn't a smarter error list. It clusters a site's content by what it's actually about, flags pages quietly competing for the same keyword territory before a client notices traffic splitting between them, and reads the link structure by topic instead of by folder. None of that shows up on a technical crawl, because none of it is technically wrong.

It still won't tell me a product line is being retired next quarter, or that a page ranks for a term the business stopped caring about last year. That judgement stays mine to make — it always will. What's different is that I'm making it from a shortlist the system built, not a spreadsheet of 404s.

What I built these past months was, above all, a determination not to leave something started long ago incomplete. But whether it becomes a commercial product, at this stage, is something I still don't know.