Skip to main content
Request parameters for Crawl API: limits, path filters, sitemap, domain scope, output formats, geo and webhook delivery

POST endpoint

https://data.decodo.com/v1/crawl

Input parameters

ParameterTypeDescription
urlstringThe URL to start crawling from.
limitintegerMaximum number of pages to scrape. Maximum value is 10000.
max_depthintegerMaximum number of link hops from the starting URL. Pages at this depth are scraped, but their links aren’t followed.
JS_retrybooleanwhen set to true, pages that failed to successfully scrape are retried with JavaScript rendering
sitemapstringPossible values:include, only, skip.
  • include - collect URLs from both the sitemap and page links.
  • only - crawl only URLs listed in the sitemap.
  • skip - ignore the sitemap and follow page links only.
Default value is include
select_pathsarrayOnly URLs whose path matches at least one pattern are crawled. Example: blog
exclude_pathsarrayURLs whose path matches any pattern are skipped. Example: case-studies
domain_filter
geostringSet the country to use when submitting the query. Example: United States
localestringThis will change the search page web interface language (not the results).
Example: – en-US – en-GB
webhook

Output

Starting a crawl returns its crawl_id immediately, along with links to its status and results. Use it to check status, fetch results, or cancel.