Site Crawler¶
Site Crawler is Experimental and off by default. Enable it in Settings → Features → Site Crawler to show it in your workspace sidebar.
Site Crawler saves desktop and mobile screenshots of a website, with descriptions of page elements and recorded navigation. Use it to inspect pages, compare captured states, and replay the interactions that were saved.
Current scope: a crawl creates reusable captures. Synthetic-shopper experiments do not yet run from these captures. Replaying a crawl does not visit the live site or place an order.
Start a crawl¶
- Select the property you want to work with in the property picker.
- Click Site Crawler in the sidebar.
- Click New crawl.
- Choose Default Property to use the property's current URL, or Variant to enter another public address in Variant URL. Choosing Variant creates a separate crawl; it does not create an experiment variant.
- Set Maximum pages and Time limit (minutes) using the limits below.
- Click Review crawl to check the credit reservation.
- Click Start crawl to reserve the credits and open the Base checkpoint. Every fresh crawl starts its own family.
| Setting | Default | Allowed values |
|---|---|---|
| Maximum pages | 100 | 1–500 unique page URLs, shared across desktop and mobile |
| Time limit (minutes) | 60 | 5–120 minutes, starting when capture begins |
| Devices | Desktop and mobile | Both are included in every crawl |
The crawler follows pages on the starting site after any initial redirect. It can record menus, dialogs, product choices and cart interactions. After adding an item to the cart, it opens the checkout page once and saves it as a final view without entering or submitting anything; payment and order steps are never opened, and purchase and personal-information submission controls are described without being used. That saved checkout view is what a 100-shopper experiment with a checkout goal verifies its conversion target against. A page or interaction may remain uncaptured if it is blocked, unavailable or outside the crawl's limits.
Expand a checkpoint¶
Open a finished checkpoint and select a saved view. The checkpoint selector shows the family's names, parent relationships, dates and statuses. You can branch from any completed, partial, stopped or failed checkpoint with a usable view after it has finished saving and returning unused credits.
- Select an unexplored button or link in the screenshot or interaction inspector, then choose Continue exploration to capture its result and continue along that branch.
- Choose Explore from this view to explore direct actions first, followed by deeper paths while limits allow. Already captured paths can lead to unexplored descendants.
- Choose Guided exploration and describe what to explore. The browser follows permitted observed actions, stops when the request is satisfied, and reports what it completed or could not do.
Give the new checkpoint a unique name, then review the limits and credits. Names are 1–80 characters after trimming spaces and must be unique across the whole family, ignoring capitalization. Base is reserved for the first crawl.
| Setting | Continuation behavior |
|---|---|
| Maximum new views | Defaults to 20; allowed range 1–500. Counts newly saved exploration views only. |
| Maximum pages | Inherits the parent's setting; 1–500 URLs. The starting URL counts. |
| Time limit (minutes) | Inherits the parent's setting; 5–120 minutes, including restoration. |
| Starting view/device | The selected saved view and its desktop or mobile device. |
The new checkpoint includes its parent's complete map plus the results it adds. Earlier checkpoints and sibling branches remain unchanged. Its progress shows the added views, explored URLs, remaining actions and charges beside the cumulative map.
The crawler must restore and verify the selected view on the live site before continuing. If the site, dialog, options or cart context cannot be restored reliably, exploration stops with an explanation. It never repeats an uncertain cart change to reconstruct a view. Start a fresh Base when you need a new snapshot of changed content.
Access and launch limits¶
Workspace members with permission to view the property can inspect its saved crawls. Owners, admins and members can start and stop crawls. Viewers cannot launch or stop them.
Only one crawl can be active in a workspace at a time. A demo, suspended or unapproved workspace cannot launch. If New crawl is unavailable, read the message on the history page; launch access may also be disabled while the feature is being prepared for your workspace.
Credits and refunds¶
A Base crawl reserves Maximum pages × 2 credits, covering one desktop and one mobile capture per URL. The default reservation is 200 credits. Reduce Maximum pages if you want a smaller reservation; reducing the time limit alone does not reduce it.
The final charge is 1 credit per successfully saved page and device, including its screenshot and completed component descriptions. Additional dialog, menu and other interaction states for that page and device are included. A screenshot with incomplete descriptions can remain visible without being charged.
A continuation reserves the smaller of Maximum pages and Maximum new views for its selected device. It charges one credit per URL/device delivering new, fully described exploration results, even when that URL was visited in a parent. Inherited screenshots, restoration-only work and duplicate outcomes are free.
Unused credits return automatically when the crawl finishes, fails or is stopped. For example, a crawl that reserves 20 credits and delivers six chargeable page/device captures charges six and returns 14. A crawl with no chargeable captures receives a full refund.
Crawls use the workspace's shared credit balance but do not count toward the monthly experiment limit. See Credits and tests for experiment pricing.
Follow progress¶
The results page adds screenshots as they are saved. It shows captured pages, pending pages, captured states, desktop/mobile state counts and the credit reservation. State counts include dialogs and interactions, so they can be higher than the page count.
The site size is not known in advance. Counts show the work discovered and saved so far, not a percentage of the entire website. You can leave the results page and return through Site Crawler; history groups checkpoints by Base family, newest first.
To stop an active crawl, click Stop crawl on its results page. Saved captures remain available while the crawl finishes cleanup and returns unused credits.
Inspect the map¶
- Select Desktop or Mobile. Desktop is selected first; the tabs show the states saved for each device.
- Enter a page title or URL in Find a page or dialog to narrow the map.
- Drag the map to pan, or use the zoom and fit controls to adjust the view.
- Select a screenshot to open its details in the inspector. Click Focus selected page to center it.
Each page is labeled Page title – URL. Dialogs and other interaction states have their own screenshots and state labels. A captured dialog can therefore appear beside the underlying page after it was closed.
Arrows show navigation and interaction links. Solid links represent captured actions; dashed links represent discovered navigation. A discovered link alone does not prove its destination can be replayed.
Component descriptions¶
- Enable Component viewer to show outlines around the recorded elements.
- Select an outline, or use the inspector's Components selector, to read its description.
- Read Page observations and Capture notes for the supporting observations and any limits.
Measured observations come from what the crawler recorded, such as an element's position at several scroll offsets. AI inference identifies an interpretation of the captured content. A scroll behavior marked unknown was not established by the observations. Descriptions do not alter the screenshot pixels.
Replay the captured site¶
- Select the page or dialog where you want to begin.
- Click Replay captured site.
- Click a recorded element in the screenshot, or choose an action under Navigation & interactions, to follow its saved destination.
- Click Back to return to the previous captured state.
- Click Back to map to leave replay.
Replay uses saved states only. If an action was not captured, it displays the recorded reason instead of opening a live page or inventing a result. A cart state follows only destinations recorded with that same cart context. Verified restoration links connect saved views to their newly restored browser context. Switching devices starts a separate replay history.
Partial or missing results¶
Partial, failed and stopped crawls keep the results they saved. Read the reason near the progress counts, then select an affected screenshot and read its Capture notes.
| What you see | What it means |
|---|---|
| A page limit or time limit reason | The crawl reached a configured boundary; some pages or interactions may remain uncaptured. |
| A cropped or incomplete page capture | Only the recorded portion of the document is available. |
| Incomplete component observations | The image was saved, but some descriptions could not be completed. |
| A blocked or unexplored action | There is no captured destination for that action. |
| No pages were captured | The crawl ended before a page state could be saved. |
Capturing a page does not mean every interaction on it was captured. Check both page coverage and the available actions before using the result. Continue from a usable view to expand missing coverage, or start a new Base to capture changed content. A checkpoint never replaces an earlier one.
Download the capture record¶
Click Download site-data.json on the results page to save the structured capture record. It describes pages, states, component bounds and observations, recorded actions, screenshot references and missing coverage. A download made while the crawl is active contains only the results saved so far; download again after it ends for the final record.
The JSON file is not a self-contained image archive. Screenshots remain private and require authorized access through Squoosh. Downloading or sharing the JSON does not make its referenced images public.