Skip to content

Search PixyScan

Guide contents

User guide5 min

Add a site

A site is one website you want PixyScan to check. To add one, give it a name, paste its URL and press Connect and scan now. The three optional sections below are for sites that need special handling. Most sites can leave them closed.

The two required fields#

PixyScan needs a name that you will recognise and a URL where the crawler should start.

  1. 1

    Press Add a site

    It is at the top of the workspace sidebar, and in the header of the workspace overview.

    Seeing an upgrade message instead of the form? You have used all the sites your tier allows: 1 on Free, 2 on Hobby, 5 on Basic and 25 on Pro. Upgrade to add more.

  2. 2

    Fill in Site name

    This name is only a label for you and your team, so use whatever you will recognise in a list, such as a brand, a client or an environment like “Acme staging”. You can change it later.

  3. 3

    Fill in Site URL

    The address must start with https://. The crawl begins at this address, so use the one your visitors actually land on, including www. if that is the version your site uses.

  4. 4

    Press Connect and scan now

    PixyScan creates the site and starts its first scan straight away. If you want to set crawl rules before anything is fetched, press Connect only instead; that creates the site without scanning it, and you can start the scan later.

app.pixyscan.com/w/…/s/new

The Add a site screen: Site name and Site URL fields, three collapsed sections for which URLs get crawled, how pages are fetched and signing in, and Cancel, Connect only and Connect and scan now buttons.
The three sections below Site URL start closed. Most sites can leave them as they are.

Which URLs get crawled#

This section controls where the crawler may go, how deep it goes and how many pages it may fetch.

Do I need this? Only if part of your site should be left out, like an admin area. Or if your site makes endless URLs, like product filters, a calendar or search results. Otherwise the defaults crawl every page the crawler can reach.

You can change these later on the Crawl tab of the site's Settings.

FieldDefaultWhat it does
Include URLs or patterns—Leave this empty to crawl every reachable URL. If you add a pattern, only matching URLs are crawled, and a full URL you add is also crawled directly.
Exclude patterns—URLs that match an exclude pattern are skipped. If a URL matches both an include and an exclude pattern, the exclude pattern wins.
When a URL matches—You must choose one once you add an exclude pattern. Pre skips a matching URL before it is fetched, so pages reachable only through it are never found. Post fetches the page so the crawler can follow its links, but leaves the page's own data out.
Max depthUnlimitedHow many links deep the crawler may follow from the start URL, up to 50. Sitemap URLs are only crawled when this is left blank.
Page budgetYour tier's limitThe most pages one scan may fetch. Leave it blank to crawl as many as your tier allows (500 per scan on Free, up to 15,000 on Pro).
Also crawl from sitemap.xmlOnThe crawler also reads your XML sitemap, so it finds pages that are listed there even if no other page links to them.
Respect robots.txtOnThe crawler obeys your robots.txt rules and nofollow instructions, as a search engine would. Turn it off only for a staging site that blocks everything on purpose.

Each crawled URL uses one URL credit

Your workspace gets URL credits each month: 500 on Free, 5,000 on Hobby, 20,000 on Basic and 125,000 on Pro. A scan that fetches 400 pages uses 400 credits. On a large site, set Page budget to cap how many credits each scan can spend. How credits work

How pages are fetched#

This section sets how the crawler reads each page, what name it gives your server and how fast it goes.

FieldDefaultWhat it does
JS renderingServer HTMLServer HTML reads the page your server sends, which is fast and right for most sites. Rendered JavaScript loads each page in a real browser and runs its scripts first. It is included from Pro up and takes roughly 7.6 times as long.
User agentDefault (pixyscan-bot)The name the crawler gives your server. You can also choose Desktop Chrome, Mobile Chrome, Googlebot (smartphone) or a Custom string, for example if your firewall only lets certain visitors through.
Rate limit healingOffWhen on, the crawler fetches pages one at a time and waits a set delay between them (from 0.1 to 10 seconds). Turn it on if your server blocks or slows down during scans; the form shows how much longer the scan will take.

Start with Server HTML

Server HTML shows you what a search engine receives before any JavaScript runs, which is often where problems hide. Switch to Rendered JavaScript only if your pages show their main content through JavaScript and the scan reports them as empty.

Signing in to this site#

Use this section for a staging or preview site that is protected by HTTP Basic Auth.

HTTP Basic Auth is the simple browser pop-up that asks for a username and password. Turn it on and enter the Username and Password. PixyScan sends them with every request and stores them encrypted.

Can it sign in through a login form? No. Basic Auth is the only sign-in the crawler supports.

Need a custom request header instead, such as an API key? Add it after the site is created, on the Crawl tab of the site's Settings.

What happens after you connect#

The form lists four stages under What happens next?, and the fourth starts with your second scan.

The same four stages the form lists under What happens next?
FieldStageWhat it does
Deep site crawl1The crawler discovers your pages and the files they use, and records the status code each one returns.
AI & search readiness2Every page is checked for how well search crawlers and AI answer engines can read and quote it.
Prioritised health score3Findings are sorted into Critical, Important and Standard issues, and the site gets a health score out of 100.
Change detection, from the second scan4Each later scan is compared with the previous one, so you see what a release changed instead of rereading the whole list.

The next page, Run a scan, explains what happens while a scan runs and where to find the results.

Does something here not match what you see in the app? Tell us