Allowing is a decision, not a default
A site that blocks AI crawlers can still be described by a model from second-hand sources, but it cannot be cited, and citation is what carries a link back to you. Deciding to allow them is a commercial choice with a real trade-off: your content becomes usable in answers you do not control.
Make the decision explicitly and write it down, because the person who later tightens a firewall rule needs to know the allowance was intentional.
Three layers can block, and only one is obvious
Robots is the visible layer and the easiest to get right. Below it, a CDN's bot-management ruleset will often challenge or drop unfamiliar agents by default - the request never reaches your server, so your logs show nothing and the robots file looks fine. Below that, rate limiting can let the first few pages through and then start refusing, which produces a partial crawl that is harder to spot than a total block.
Verify from outside
Fetch your own pages from a network that is not your office, using each crawler's user agent, and check the status code and the returned body. A 200 with a challenge page in the body is a block. Then fetch with scripts disabled: if the body comes back as an empty shell, the content exists only after client-side rendering and a crawler that does not execute scripts sees nothing.
Keep the check running
Access is not a one-off. CDN vendors update their bot rulesets, certificates expire, and a well-meaning change to a WAF rule can undo the allowance months later. A scheduled fetch that alerts on a status change costs almost nothing and catches the silent failures.