SMF·SITEBULB
Can a prompt replace Sitebulb?
SEO & marketing — technical audits and on-page optimization
Exhibit tracking slip
Verdict
A crawler that fetches, parses and stores a site is a solved shape, and the catalogue already has an entry for the desktop-crawler category. Sitebulb's real product is the layer above: several hundred rules, each with a written explanation of why it matters, an evidence sample, and a priority ordering that stops an audit turning into a spreadsheet of 40,000 rows. You can write fifty good rules in a weekend. You cannot write the other few hundred, and the explanations are the part clients actually read.
Exhibit A — The prompt
Received on31.07.2026Build a site auditor whose output is a prioritised, explained report rather than a URL dump.
Guard first: before any crawl starts, require the user to confirm they own or have permission to crawl the target domain, and record that confirmation with the crawl.
Crawler: fetch with a descriptive user agent, obey robots.txt, respect a configurable concurrency and delay, and cap total URLs. Follow internal links only, record redirect chains in full rather than only the final destination, and store every response's status, headers, timing and body. Persist crawls in SQLite so a run can be reopened, compared and exported later.
Extraction per page: title, meta description, canonical, meta robots, H1 through H3, word count, internal and external links with anchor text, images with alt text and byte size, hreflang, structured data blocks, and the response's cache and compression headers.
Rule engine, which is the point. Rules live in a readable file, not in code. Each rule declares an id, a severity, a one-paragraph explanation of why it matters, a remediation note, and a predicate over the extracted page fields. Ship at least these, grouped:
- Indexability: noindex on a page receiving internal links, canonical pointing to a redirect, canonical pointing off-site, robots.txt blocking a page in the sitemap, redirect chains longer than two hops, redirect loops.
- Content: duplicate or missing titles, titles over the pixel budget, missing or duplicate meta descriptions, missing H1, multiple H1s, thin pages under a word threshold, near-duplicate body content by shingle similarity.
- Links: broken internal links, orphan pages present in the sitemap but not linked, pages more than four clicks from the home page, links to non-canonical versions, empty or generic anchor text.
- Performance and hygiene: uncompressed responses, images over a byte threshold, missing alt text, mixed content, missing hreflang return links.
Report: findings grouped by severity, each showing the count, up to twenty example URLs, the explanation, and the remediation note. Export a self-contained HTML file with everything inlined, plus CSV per rule. The HTML report is the deliverable, so it must open with no server and no network.
Crawl comparison: given two crawls of the same site, show which findings appeared, disappeared or changed count. That diff is what makes an audit tool worth running twice.
Out of scope: JavaScript rendering, crawling sites the user has not confirmed owning, backlink or keyword data, and any automatic change to a live site.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 the rule library — hundreds of checks refined against real client sites
- B.2 JavaScript rendering, which many sites now require to be crawlable at all
- B.3 the crawl visualisations that make site structure legible
- B.4 scheduled re-crawls with change detection between runs
Prior art
Exhibit C — Why people still pay: crawl scale, rule depth, and operational polish
Because an audit is a document you hand to someone else, and the value is in the explanation and the ordering, not in the crawl. That library is years of accumulated judgement about which problems actually cost traffic.
Questions
Can I import a Sitebulb crawl?
Sitebulb exports its data as CSV per report, which imports as findings but not as a re-usable crawl — the page-level detail this build's rules run against is not in those exports. In practice you re-crawl, which for a site under a few thousand pages takes minutes.
Does it handle JavaScript-rendered sites?
No, and that is the biggest practical limitation. Driving a headless browser per URL turns a five-minute crawl into an hour and brings a large maintenance burden. If your site only exposes its content after JavaScript runs, this build will report it as empty, which is at least an honest failure.
What does it cost to run?
Nothing. It runs locally against a SQLite file, and the only real limit is how long you are willing to wait. Sitebulb's Lite plan is per-month; a self-built crawler you use twice a year is free.
What is the one thing that does not survive the rebuild?
The explanations. Fifty rules you wrote yourself will find most of the important problems, but Sitebulb hands a client a paragraph on why each one costs them traffic, written by people who have argued the point with developers a thousand times. That is the deliverable, and it is not something you write in a weekend.
Related tools
Receipt