Because a whole-site crawl is indiscriminate. It ingests old blog posts, superseded pricing and a careers page, then answers a pricing question from a two-year-old announcement.
Website crawling trains the assistant on pages from your own site, on a schedule and within limits you set.
You choose paths to include and exclude, and the crawl recrawls on a schedule so answers follow edits. Each crawled page is a source that can be inspected and removed.
A blog archive with four years of announcements is excluded while documentation and pricing are included, so nothing answers from a superseded post.
It will not crawl what it cannot reach — pages behind authentication, blocked in robots, or rendered only after interaction are reported as skipped rather than guessed at.
The pages you include, at the depth you set, respecting robots rules and skipping anything behind a login.
On your schedule, and on demand after a change.
Yes, page by page, and remove any of them.
Yes, including on your own site.
Fourteen days, every module, no card. Or half an hour with someone who will run it on your own records and tell you where it does not help.