by watercrawl ยท v0.1.0
Watercrawl Datasource
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
A powerful Dify datasource plugin that integrates with Watercrawl to recursively crawl websites and extract clean, LLM-ready content for AI applications.
| Feature | Description |
|---|---|
| ๐ท๏ธ Recursive Crawling | Automatically discover and crawl linked pages |
| โก Fast Mode | Skip rendering for 3x faster crawling |
| ๐ Depth Control | Limit crawl depth to control scope |
| ๐ฏ Smart Extraction | Extract main content, skip navigation/ads |
| ๐ URL Filtering | Include/exclude specific URL patterns |
| ๐ Progress Tracking | Real-time crawl status monitoring |
| ๐ Clean Output | Markdown-formatted, LLM-optimized content |
| ๐ Self-Hosted | Use cloud service or your own instance |
Before using this plugin, you need:
# 1. Get your API key from Watercrawl dashboard
https://app.watercrawl.dev/dashboard/api-keys
# 2. Configure in Dify
Base URL: https://api.watercrawl.dev # (or leave empty)
API Key: wc_your_api_key_here
# 1. Deploy your Watercrawl instance
# Follow: https://docs.watercrawl.dev/self-hosted/overview
# 2. Configure in Dify
Base URL: https://your-watercrawl-instance.com
API Key: any-value # (required but can be arbitrary)
Extract content from one specific page:
Start URL: https://example.com/article
Maximum crawl depth: 1
Maximum pages to crawl: 1
Only main content: true
Crawl an entire website systematically:
Start URL: https://example.com
Maximum crawl depth: 2
Maximum pages to crawl: 50
URL patterns to include: # (optional)
URL patterns to exclude: admin/*, login/*
Focus on specific website sections:
Start URL: https://docs.example.com
Maximum crawl depth: 3
Maximum pages to crawl: 100
URL patterns to include: api/*, guides/*, tutorials/*
URL patterns to exclude: archive/*, changelog/*
| Parameter | Type | Default | Description |
|---|---|---|---|
| Start URL | string | required | Base URL to begin crawling |
| Ignore rendering | boolean | false | Skip page rendering for faster crawling |
| URL patterns to exclude | string | - | Comma-separated exclude patterns: blog/*, about/* |
| URL patterns to include | string | - | Comma-separated include patterns: docs/*, api/* |
| Maximum crawl depth | number | 1 | How deep to crawl (1 = start URL only) |
| Maximum pages to crawl | number | 1 | Total page limit for the crawl job |
| Only main content | boolean | false | Extract main content, skip nav/footer |
| Proxy Server Slug | string | - | Proxy server identifier (optional) |
Depth 1: [Start URL] โ Direct links
Depth 2: [Start URL] โ Direct links โ Links from those pages
Depth 3: [Start URL] โ Direct links โ 2nd level โ 3rd level pages
Include Only Specific Sections:
blog/*, docs/api/*, products/*/specs
Exclude Unwanted Areas:
admin/*, login/*, tag/*, category/*
Each crawled page returns structured data:
{
"source_url": "https://example.com/page",
"title": "Page Title",
"description": "Meta description",
"content": "# Clean Markdown Content\n\nParagraph text..."
}
Perfect for creating AI-powered knowledge bases:
Start URL: https://docs.example.com
Maximum crawl depth: 3
Maximum pages to crawl: 200
URL patterns to include: api/*, guides/*, tutorials/*
URL patterns to exclude: changelog/*, archive/*
Only main content: true
Extract blog posts and articles for analysis:
Start URL: https://blog.example.com
Maximum crawl depth: 2
Maximum pages to crawl: 100
URL patterns to include: posts/*, articles/*
URL patterns to exclude: tag/*, author/*, comments/*
Only main content: true
Gather product information systematically:
Start URL: https://shop.example.com/products
Maximum crawl depth: 2
Maximum pages to crawl: 50
URL patterns to include: products/*
URL patterns to exclude: cart/*, checkout/*, account/*
Only main content: true
Monitor competitor websites for changes:
Start URL: https://competitor.com
Maximum crawl depth: 2
URL patterns to include: products/*, pricing/*, features/*
URL patterns to exclude: blog/*, news/*, support/*
depth: 1, limit: 5)robots.txt before large crawlsignore_rendering: true for speed| Scenario | Typical Time | Recommendation |
|---|---|---|
| Single page | < 5 seconds | Use for quick extractions |
| Small site (< 20 pages) | 30-60 seconds | Good for testing patterns |
| Medium site (50-100 pages) | 2-5 minutes | Monitor progress |
| Large site (> 200 pages) | 5-15 minutes | Use during off-hours |
๐ก Tip: Deep crawling (depth > 3) can exponentially increase page count. Use cautiously!
"API key is required" error:
https://api.watercrawl.dev)wc-"Failed to crawl" error:
ignore_rendering: trueMissing expected pages:
Crawling too slow:
ignore_rendering for 3x speed boostMade with โค๏ธ by the Watercrawl Team
Website โข Documentation โข Dashboard