crawl_id and ready-to-use links for every endpoint below.
| Endpoint | Method | Purpose |
|---|---|---|
https://data.decodo.com/v1/crawl | POST | Starts a crawl. |
https://data.decodo.com/v1/crawl/{crawl_id}/status | GET | Crawl status. |
https://data.decodo.com/v1/crawl/{crawl_id} | GET | List of crawled URLs, their status and Task IDs for each URL. You can then fetch each URL content individually using Task ID. |
https://data.decodo.com/v1/crawl/{crawl_id}?type=content | GET | Scraped page content inline. |
https://data.decodo.com/v1/crawl/{crawl_id}/cancel | POST | Stop the crawl and keep collected results. |
Check crawl status
Endpoint: https://data.decodo.com/v1/crawl/{crawl_id}/status
This is the crawl_status link from the crawl response.
Get results
Results are stored for 24 hours after the crawl finishes.
- List of crawled URLs
Endpoint: https://data.decodo.com/v1/crawl/{crawl_id}
This is the results_tasks link. It returns every URL in the crawl with its own status, so you can see which pages are finished, still running, or failed.
https://data.decodo.com/v1/task/{task_id}/results endpoint
- Full content inline in the reponse
Endpoint: https://data.decodo.com/v1/crawl/{crawl_id}?type=content
This is the results_content link. It returns the scraped content of every page, in the format you chose with output. Results are split into pages of up to 10 MB - use the cur parameter to get the next one.
Webhooks
Addcallback_url to your crawl request and we’ll send a POST request to it once all pages are processed, so you don’t need to poll the status endpoint. To confirm a webhook really came from Decodo, add a passthrough value to your request. It’s returned unchanged in the webhook payload.