Website · THE NO-PANIC PLAN
Prepare a new website for search crawling and indexing
A robots.txt file controls crawler access to selected paths; it does not make private content secure or reliably keep a URL out of search results. A sitemap helps search engines discover canonical public URLs, but it is only a hint. This workflow checks your launch scope and public URLs, creates reviewed robots rules with Nirmion's Robots.txt Generator, assembles a canonical URL list with the Sitemap Generator, tests both files on the live origin, and monitors Search Console. Neither tool crawls your site or submits anything for you. Keep private pages behind authentication, and use an accessible noindex directive only when a page should remain crawlable but not appear in search.
MISSION Set up a new public website's crawl directives and canonical XML sitemap, then check access and indexing signals in Search Console without promising ranking or indexing.
Prepare a robots.txt file and canonical sitemap for launchTHE REAL-WORLD BIT
What happens outside this browser tab?
Define which pages should be public and indexable; verify representative pages and canonical URLs; create minimal crawler directives that do not block required resources; prepare a sitemap from a verified URL inventory; publish and fetch both files from the correct origin; submit and monitor them in Search Console; then fix access/indexing conflicts without treating crawl or indexing as guaranteed.
YOUR CHECKLIST, WITH FEWER DRAMATIC SIGHES
One step at a time.
Follow the order below. If a step names a Nirmion tool, its link is right there with it.
- 01
Decide which pages are public and eligible for search
Create a short inventory of launch URLs and mark each as public/indexable, public but intentionally noindex, or private. Confirm that public pages return their intended HTTP response, contain useful content, and use the preferred canonical URL. Keep confidential or account-only material protected by authentication; robots.txt is publicly visible and is not access control. For a page that must stay out of search while remaining accessible to crawlers, use an appropriate noindex directive and do not block that URL in robots.txt, otherwise a crawler may not see the directive. (Sources 1, 2, 3, 4)
- 02
Create the smallest reviewed robots.txt policy you need
If the site needs crawler directives, list only the paths that should be disallowed and verify that important public pages, CSS, JavaScript, and other resources remain accessible to the crawler. Use Nirmion's Robots.txt Generator to draft bounded rules and an optional same-origin sitemap reference. Review every directive against the exact production path and the site's deployment model before copying the output. robots.txt rules guide cooperative crawlers; they are not enforced access controls, and a disallowed URL may still be indexed if discovered elsewhere. Do not use the file to hide sensitive data or replace authentication. (Sources 1, 2)
- 03
Build a sitemap from verified canonical URLs
Export the launch inventory and include only absolute, canonical URLs that you want search engines to discover. Remove staging hosts, redirects, error pages, duplicate parameter variants, private pages, and URLs intentionally set to noindex. Nirmion's Sitemap Generator creates escaped XML from up to 500 supplied same-origin URLs; it does not crawl a site, discover missing pages, or decide which URLs should be included. For a larger site, use the site's CMS or a sitemap index workflow that follows Google's size and format guidance. (Sources 2, 5)
- 04
Publish and fetch both files from the production origin
Place robots.txt at the correct host root and publish the sitemap at its intended URL. Fetch each address over HTTPS from the production hostname, including any www/non-www or regional host variants that the site serves. Confirm a successful response, expected content type and exact contents, no accidental authentication prompt, and no rewrite to a staging file. Test important public URLs against the robots rules, then inspect representative pages to confirm the rendered page has the intended canonical and robots meta/header values. Do not assume generated text is live merely because a tool produced it. (Sources 1, 2, 4, 5)
- 05
Check for crawl and indexing conflicts
Compare each intended page's robots.txt access, page-level noindex directive, HTTP response, canonical target, and sitemap presence. A page blocked from crawling can prevent Google from reading its noindex directive; a URL in a sitemap does not override a noindex or guarantee indexing. Use the live URL Inspection view in Search Console to check what Google can fetch for representative important pages, and verify ownership of the correct property. Resolve contradictory signals before asking for recrawling, and test the actual rendered production URL instead of relying on a source template alone. (Sources 1, 3, 4, 6)
- 06
Submit the sitemap and monitor Search Console
Submit the production sitemap URL in the correct Search Console property and review its fetch status and reported errors. Use URL Inspection for a small number of important new or corrected URLs; use a sitemap to help discovery across many URLs. Requesting a crawl does not guarantee that a page will be indexed or appear in search, and repeating the request does not make crawling faster. Check the Page Indexing report and server logs over time, then fix the underlying access, response, canonical, or content issue indicated by evidence. (Sources 5, 6)
- 07
Keep launch rules current as the site changes
When pages are added, removed, moved, or reclassified, update the canonical URL inventory, sitemap, and crawler rules together. Review robots.txt after major deployments so old staging blocks do not accidentally stop crawling, and confirm removed content returns the intended status. Keep a dated copy of the deployed files and the owner who approved the rules. Treat crawl requests, sitemap acceptance, indexing reports, and search visibility as separate signals: neither generator guarantees coverage, traffic, rankings, or indexing. (Sources 1, 2, 5, 6)
THE HELPER CREW
Tools for the fiddly bits.
These are the currently published Nirmion tools matched to this guide. Open a tool page for its accepted inputs and limits.
RECEIPTS, PLEASE
Sources & review notes
Each source is linked to the steps it supports. Open it to check its scope and current guidance.
Source checked 2026-10-10
- Google Search Central: Introduction to robots.txt
- Google Search Central: Build and Submit a Sitemap
- Google Search Central: Block Search Indexing with noindex
- Google Search Central: Robots Meta Tag and X-Robots-Tag
- Google Search Central: Learn about Sitemaps
- Google Search Central: Ask Google to Recrawl Your Website