BeaconBot
BeaconBot is the crawler behind Beacon’s SEO & AI-visibility audits. It only visits a site when someone asks it to — a Beacon user who added the site as a project, or a visitor running one of our free checks against it. It is not a search engine and does not build a public index.
User agent
Every request identifies itself with this User-Agent header (the headless-browser render used by the Render Gap report appends (rendered); the Lighthouse performance run uses Lighthouse’s own Chrome user agent):
BeaconBot/1.0 (+https://beacon.seo/bot)
What it fetches
/robots.txt,/llms.txtand your XML sitemap(s) — read first, before any page — plus/.well-known/security.txtfor the security report.- HTML pages of the audited site, following internal links breadth-first. Non-HTML resources (PDF, images, CSS, JS, fonts, archives, media) are skipped.
- Response headers (status, redirects,
X-Robots-Tag,Link, security headers) for the technical and security checks. - For the Render Gap and performance reports, one page is additionally loaded in headless Chromium so we can compare the JavaScript-rendered DOM with the raw HTML.
Beacon extracts on-page and technical SEO signals (titles, headings, schema, canonicals, word counts, link structure). It does not submit forms, log in, or execute anything beyond loading the page.
Crawl rate and limits
- Page budget: 40 to 400 pages per audit, depending on the plan of the Beacon user who requested it. A first (onboarding) scan is capped at 30 pages.
- Concurrency: pages are fetched in small batches of up to 5 concurrent requests.
- Timeout: 12 seconds per request; at most 5 redirects are followed per URL.
- Frequency: on demand, plus scheduled re-audits (weekly by default, daily at most) for projects on a paid plan that enabled monitoring.
robots.txt
BeaconBot honours robots.txt per RFC 9309: it obeys the group addressed to BeaconBot(matched case-insensitively), otherwise the * group. Disallow and Allow paths support * wildcards and $ anchors; the longest matching rule wins and Allow wins a tie. The crawler never follows a link into a disallowed path. The one URL a person explicitly submits (a project’s homepage, or the address typed into a free check) is fetched so the report can cover it; everything beyond it goes through robots.txt.
To block BeaconBot entirely:
User-agent: BeaconBot Disallow: /
To allow it while keeping a general block in place:
User-agent: BeaconBot Allow: / User-agent: * Disallow: /
Allow-listing BeaconBot
If a WAF or bot-protection layer (Cloudflare, Akamai, Fastly, PerimeterX, …) challenges or blocks our requests, the audit reports those pages as “blocked by bot protection” and cannot measure them. To let Beacon read your site, add a rule that allows requests whose User-Agent contains BeaconBot(for example a Cloudflare WAF custom rule with http.user_agent contains "BeaconBot" → Skip). Alternatively, connect Google Search Console or upload a sitemap in Beacon so we can audit your indexable URLs without fetching the live HTML.
Questions or abuse reports
If you believe BeaconBot is misbehaving on your site, email sales@ibeacon.ai with the URL and a log excerpt (timestamp, User-Agent). Beacon is operated by Youlinker SIA, Riga, Latvia.