What the file is for

A crawler arriving at a site has to work out which pages carry assertions worth reading. On a large site that inference is expensive and often wrong - it will spend its budget on category listings and tag archives. The file exists to say: these are the pages that answer something, and here is what each one answers.

Why copying the sitemap defeats the purpose

A sitemap lists everything, which is exactly the problem the file is meant to solve. Four hundred URLs with no annotation conveys no more than crawling the site would. Twenty to fifty entries, each with a line saying what the page establishes, is the useful form.

What a good entry looks like

A link, then a clause describing the claim the page makes - not its title. "Pricing" is a title. "Published prices for all five tiers, including what each quota actually permits" is a claim. The second tells a crawler whether the page is worth fetching for the question in hand.

Group entries under headings that match how someone would ask: what the service is, what it costs, how it is delivered, who it suits, what the limits are. The limits section is the one most sites omit and the one most likely to be quoted, because honest constraints are the hardest thing for a model to find elsewhere.

Keep it honest and current

An entry pointing at a page that no longer exists is worse than no entry: it is a verifiable error on a file whose entire purpose is to be trusted. Generate the file during the build from the pages that actually exist, or review it on the same schedule as the pricing page.

Where it goes

At the site root, served as plain text, reachable without a redirect chain. Then fetch it from outside and confirm it returns 200 with the body you expect - the same silent blocks that stop page crawling stop this file too.