Skip to main content

How it works

The Web connector scrapes sites from a Base URL.
  • It only indexes files from the same domain that share the same base path.
  • It indexes pages reachable via hyperlinks from the base URL.
  • Page text is cleaned up, and metadata such as the page title is extracted.
As long as the page is reachable, no extra authorization is required.

Setting up

Setup runs through three steps: configure the connector, choose a space, then set sync options.

Step 1: Configure the connector

  1. Open AdminConnectorsAdd Connector.
  2. Select the Web connector.
  3. Enter a Connector Name — a descriptive name for this connector.
  4. Enter the Base URL to scrape.
  5. Choose a Scrape Method.
  6. Click Continue.
Web connector — name, base URL, and scrape method

Step 2: Select a space

Choose the space to index the documents from. This controls which space the connector’s indexed content belongs to. Web connector — select a space Use the search field to filter the list if you have many spaces, then click Continue. Use Previous to go back and change the URL or scrape method.

Step 3: Advanced Configuration

The final step controls how often UtopikAI syncs with the site and how far back it indexes. The defaults work for most setups. Web connector — Advanced Configuration Entering 0 or leaving Refresh Frequency blank means UtopikAI will never pull new documents for this connector. Entering 0 or leaving Prune Frequency blank disables pruning for this connector. Pruning checks every document against the source, so be cautious when increasing frequency. Use Reset to restore the default values. When you are done, click Create Connector to save the connector and begin indexing. You can return to an earlier step with Previous.