Retiring an Old Site on the Same Domain

Retiring an old website is easy when it lives somewhere else. You switch it off. The awkward version is the one I have: an old static site and a new WordPress install sharing one hostname, one document root and one certificate, where switching anything off means being extremely precise about which address means which page.

The redirect side of that job, the part where every old URL gets a 301 or a 410 and the leftover files come out of the web root, is written up separately in the post on moving a site to WordPress without leaving ghosts behind. This post is about the half that comes after: making sure that each surviving page has exactly one address, that the address is spelled the same way everywhere, and that the file telling search engines where to look is pointing at the right list. On a domain like mine, which contains a letter that does not exist in ASCII, that turns out to be more interesting than it sounds.

My domain has two spellings, and both are correct

This site sits on an internationalised domain name: it contains the letter ñ. Your browser shows you the pretty version with the accent. Underneath, DNS does not speak Unicode, so the real hostname is the punycode form, the one that begins with xn-- and looks like something went wrong during a file transfer.

Both forms are legitimate and they refer to the same domain. That is precisely the problem, because a page with two legitimate spellings is a page that can be written down two different ways in a canonical tag, a sitemap, an internal link, a social preview or a verification record. Some of those places accept the accented version and quietly convert it. Some accept it and do not convert it. Some reject it with an error that looks like a DNS failure and is not.

My rule, arrived at the slow way: everything machine-readable uses punycode, without exception. Canonical tags, sitemap entries, the Search Console property, the hreflang I do not have yet, the redirect destinations, the certificate. The accented version is for prose and for the logo. Checking which form my certificate actually carries takes one command:

echo | openssl s_client -servername example.com -connect example.com:443 2>/dev/null | openssl x509 -noout -subject -issuer -enddate

Mine returns a subject with the xn-- hostname and a Let’s Encrypt issuer, which is what it should be. If you are earlier in this story and still choosing a domain, the buying side of the decision is in the post about the domain trap nobody warns you about.

Four ways to write one page, before you add the accent

Every site has this problem, not just mine. A single page can be requested as http:// or https://, with www. or without, with a trailing slash or without, and on a site built from static files, as the directory or as the file inside it. Multiply that by two spellings of my hostname and one article has a genuinely silly number of valid addresses.

The fix is not clever, it is just tedious: one form wins and everything else answers with a redirect to it. Mine is HTTPS, no www, trailing slash, punycode host. Verifying it is four requests:

for u in http://example.com/ http://www.example.com/ https://www.example.com/ https://example.com/; do
  curl -sI -o /dev/null -w "%{http_code} $u -> %{redirect_url}\n" "$u"
done

When I ran that against this domain, the first three answered 301 and landed on the fourth, which answered 200. That was one of the few things about this site that was already correct, and I only know it because I checked rather than assumed. The static export left a separate flavour of the same problem behind, serving both a directory and a text file for each route, which is one of the reasons the redirect list in the migration post had to match several variants per URL.

Which sitemap is the real one

Here is the chain I found on my own domain, in the order a crawler would walk it. robots.txt advertised /sitemap.xml. That file was a sitemap index containing exactly one entry: /sitemap-0.xml. That file listed 27 URLs, every one of them from the old static site, every one of them answering 200.

Meanwhile, modern WordPress had been generating its own sitemap all along at /wp-sitemap.xml, and nothing on the site referenced it. So the only list of pages I was publishing was a list of the site I had replaced. Not blocked, not broken, not penalised: just quietly pointing at the wrong website.

Worth knowing if you use the built-in sitemap: it is an index, and it links to several sub-sitemaps. Mine contains one for posts, one for pages, one for categories and one for users. That last one is the one to look at, because it publishes your author archive, and on this site the author archive URL was built from the account login name. That is not a canonicalisation issue, it is a security one, and it is dealt with in the post on hardening WordPress after finding my own site wide open. But it is the kind of thing you only discover by opening the sitemap index and reading what is in it instead of assuming the plugin has your back.

The robots.txt file that advertised all this had its own defect, three invisible bytes at the top that I found with od -c. That story belongs to the migration post as well; the short version is that a file can look perfect in every editor and still start with something you did not intend.

The canonical you declare and the canonical Google picks

A canonical tag is a recommendation, not an instruction. You declare which address you consider authoritative; the search engine decides whether to agree. Most of the time it agrees. The times it does not are the times worth finding, and you find them in one place: the URL inspection tool in Search Console, which reports the canonical you declared and the canonical Google selected as two separate fields.

Reasons the two can disagree, roughly from the most common to the least:

  • Two pages that say nearly the same thing. If you publish several articles with the same skeleton and mostly the same content, you have told the engine they are interchangeable, and it may pick one to represent the group. Consolidating near-identical pages fixes this properly; a canonical tag does not.
  • Internal links that use a different form of the address than the canonical declares. Your own links are a signal, and if they disagree with your tag you are arguing with yourself in public.
  • Sitemap entries in a different form again, which is exactly what a domain with two spellings invites.
  • A redirect pointing at a URL that then canonicalises somewhere else, so the destination of the redirect is not the destination of the page.

Checking what you actually serve, rather than what the settings screen claims you serve, is one line:

curl -s https://example.com/some-article/ | grep -i -o '<link rel="canonical"[^>]*>'

I run that on a handful of pages after any theme or plugin change, along with the structured data check, because these are the two things that break silently when a template moves. Both are in the rotation of free tools I actually open.

The one-pass check I run now

Four things, in this order, whenever I have touched anything structural. It takes about five minutes, which is less time than it takes to wonder whether something is wrong.

  • The four address forms resolve to one, with a single 301 and no chains.
  • robots.txt starts where it should, allows the crawlers I want, and advertises exactly one sitemap: the current one.
  • The sitemap index lists only sub-sitemaps I recognise, and the URLs inside are in the punycode form and return 200.
  • A sample of canonicals matches the address I would want indexed, on a post, a page, the homepage and a category archive.

Then, and only then, Search Console: remove the sitemaps that no longer exist, submit the current one as a complete URL rather than a filename, and request indexing for the pages that changed. Google documents the sitemap side of this properly (Google Search Central on building sitemaps), and the screens I use are in the guide to the parts of Search Console that matter.

What I cannot claim yet

Everything above is verifiable from outside my server, today, with curl. What is not verifiable today is the result. Google recrawls on its own schedule, and until it does, the old addresses stay in the index exactly as wrong as they were. Nobody can promise you a timeline for that, and anyone offering one is guessing.

I also cannot tell you that any of this moves rankings, because I have no traffic worth analysing and I am not going to manufacture a case study out of a handful of clicks. What I can tell you is narrower and, I think, more useful: a domain that publishes a sitemap of a site that no longer exists is publishing a lie about itself, and fixing that is not an optimisation, it is basic hygiene.

The lesson I actually took from the whole exercise is that none of these files ever complain. A sitemap pointing at a dead site returns a perfectly valid XML document. A canonical tag with the wrong spelling of your hostname is syntactically flawless. The site keeps loading, the dashboard stays green, and the only way to find any of it is to ask the server what it is serving instead of asking the admin panel what it believes.