How Link Discovery Works
Crawlers start at a list of seed URLs, extract all outbound link href attributes from the parsed HTML, and add them to a queue to request and parse recursively.
Scope and Crawl Limits
To protect server resources, crawlers limit concurrency, request depths, and total page counts. A deep site structure (more than 5 clicks from the homepage) can prevent deeper pages from being discovered.
