About our site audit
When you request a market snapshot, or when we research a business we are considering contacting, we fetch a handful of pages from the public website. This page explains exactly what that involves, because you should not have to guess.
What identifies us
Our requests carry the user agent ClearPathContentBot/1.0 and link back to this page. A few requests per site, a few seconds apart. Nothing sustained, nothing aggressive.
What we look at
- The homepage, and a blog or resources section if one is linked
/sitemap.xml, to count how many pages are published- Whether structured data is present, since AI answer engines depend on it
- Whether the site links to lead marketplaces such as Angi or Thumbtack
- Whether the business's own city appears in the page text
All of it is public information — the same things anyone sees opening the site in a browser.
What we do not do
- We do not attempt to access anything behind a login, paywall or form
- We do not collect personal data about individuals
- We do not submit forms, click buttons, or interact with the site in any way
- We do not resell, publish or share what we find with anyone else
- We do not crawl at volume — this is a handful of pages, not a spider
How to opt out
Add this to your robots.txt and we will not fetch your site:
User-agent: ClearPathContentBotDisallow: /
You can also just tell us. Reply to any email from us and we will stop, remove anything we hold on you, and not contact you again.
Why we do it
Because generic outreach wastes everyone's time. If we are going to email a business, the least we can do is look at their site first and say something true and specific about it — or work out that they are not a fit and leave them alone. Most of the businesses we audit never hear from us at all, because the audit tells us the program would not help them.