Watercrawl ยท emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginsWatercrawl
Plugin
Limited listing

Watercrawl

by watercrawl ยท v0.1.0

Watercrawl Datasource

1.8k installsUpdated Oct 5, 2025
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Data sources

Available inside your emploidai workspace after installation.

Category

datasource

Version

0.1.0watercrawl

Requirements

Maximum memory 256MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.9.0+

Permissions

  • Uses model capability

Dependencies

No additional dependencies

Resources

Privacy policy
emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

Watercrawl Datasource Plugin for Dify

Version License

A powerful Dify datasource plugin that integrates with Watercrawl to recursively crawl websites and extract clean, LLM-ready content for AI applications.

๐Ÿš€ Quick Start

  1. Get API Key: Sign up at Watercrawl and get your API key
  2. Install Plugin: Add the Watercrawl datasource in your Dify workspace
  3. Configure: Enter your API key and start crawling websites
  4. Extract: Get clean, structured content ready for your AI models

โœจ Features

FeatureDescription
๐Ÿ•ท๏ธ Recursive CrawlingAutomatically discover and crawl linked pages
โšก Fast ModeSkip rendering for 3x faster crawling
๐Ÿ“ Depth ControlLimit crawl depth to control scope
๐ŸŽฏ Smart ExtractionExtract main content, skip navigation/ads
๐Ÿ” URL FilteringInclude/exclude specific URL patterns
๐Ÿ“Š Progress TrackingReal-time crawl status monitoring
๐Ÿ“ Clean OutputMarkdown-formatted, LLM-optimized content
๐Ÿ  Self-HostedUse cloud service or your own instance

Setup

Prerequisites

Before using this plugin, you need:

  1. A Watercrawl API key (for cloud service) or self-hosted Watercrawl instance
  2. Target URLs ready for crawling
  3. Understanding of your crawling requirements (depth, limits, patterns)

โ˜๏ธ Cloud Setup (Recommended)

# 1. Get your API key from Watercrawl dashboard
https://app.watercrawl.dev/dashboard/api-keys

# 2. Configure in Dify
Base URL: https://api.watercrawl.dev  # (or leave empty)
API Key: wc_your_api_key_here

๐Ÿ  Self-Hosted Setup

# 1. Deploy your Watercrawl instance
# Follow: https://docs.watercrawl.dev/self-hosted/overview

# 2. Configure in Dify  
Base URL: https://your-watercrawl-instance.com
API Key: any-value  # (required but can be arbitrary)

๐Ÿ“– Usage Examples

๐Ÿ“„ Single Page Extraction

Extract content from one specific page:

Start URL: https://example.com/article
Maximum crawl depth: 1
Maximum pages to crawl: 1
Only main content: true

๐ŸŒ Full Website Crawl

Crawl an entire website systematically:

Start URL: https://example.com
Maximum crawl depth: 2
Maximum pages to crawl: 50
URL patterns to include: # (optional)
URL patterns to exclude: admin/*, login/*

๐ŸŽฏ Targeted Section Crawl

Focus on specific website sections:

Start URL: https://docs.example.com
Maximum crawl depth: 3
Maximum pages to crawl: 100
URL patterns to include: api/*, guides/*, tutorials/*
URL patterns to exclude: archive/*, changelog/*

โš™๏ธ Configuration Parameters

ParameterTypeDefaultDescription
Start URLstringrequiredBase URL to begin crawling
Ignore renderingbooleanfalseSkip page rendering for faster crawling
URL patterns to excludestring-Comma-separated exclude patterns: blog/*, about/*
URL patterns to includestring-Comma-separated include patterns: docs/*, api/*
Maximum crawl depthnumber1How deep to crawl (1 = start URL only)
Maximum pages to crawlnumber1Total page limit for the crawl job
Only main contentbooleanfalseExtract main content, skip nav/footer
Proxy Server Slugstring-Proxy server identifier (optional)

๐ŸŽฏ Understanding Crawl Depth

Depth 1: [Start URL] โ†’ Direct links
Depth 2: [Start URL] โ†’ Direct links โ†’ Links from those pages
Depth 3: [Start URL] โ†’ Direct links โ†’ 2nd level โ†’ 3rd level pages

๐Ÿ” URL Pattern Examples

Include Only Specific Sections:

blog/*, docs/api/*, products/*/specs

Exclude Unwanted Areas:

admin/*, login/*, tag/*, category/*

๐Ÿ“ค Output Format

Each crawled page returns structured data:

{
  "source_url": "https://example.com/page",
  "title": "Page Title", 
  "description": "Meta description",
  "content": "# Clean Markdown Content\n\nParagraph text..."
}

๐Ÿ”„ How It Works

  1. Job Creation โ†’ Submit crawl request to Watercrawl API
  2. Processing โ†’ Watercrawl crawls pages based on your parameters
  3. Monitoring โ†’ Plugin polls job status every 5 seconds
  4. Extraction โ†’ Content is cleaned and formatted as Markdown
  5. Delivery โ†’ Structured results returned to Dify

๐ŸŽฏ Common Use Cases

๐Ÿ“š Documentation Indexing

Perfect for creating AI-powered knowledge bases:

Start URL: https://docs.example.com
Maximum crawl depth: 3
Maximum pages to crawl: 200
URL patterns to include: api/*, guides/*, tutorials/*
URL patterns to exclude: changelog/*, archive/*
Only main content: true
๐Ÿ“ Content Analysis

Extract blog posts and articles for analysis:

Start URL: https://blog.example.com
Maximum crawl depth: 2
Maximum pages to crawl: 100
URL patterns to include: posts/*, articles/*
URL patterns to exclude: tag/*, author/*, comments/*
Only main content: true
๐Ÿ›๏ธ E-commerce Data

Gather product information systematically:

Start URL: https://shop.example.com/products
Maximum crawl depth: 2
Maximum pages to crawl: 50
URL patterns to include: products/*
URL patterns to exclude: cart/*, checkout/*, account/*
Only main content: true
๐Ÿ” Competitor Research

Monitor competitor websites for changes:

Start URL: https://competitor.com
Maximum crawl depth: 2
URL patterns to include: products/*, pricing/*, features/*
URL patterns to exclude: blog/*, news/*, support/*

๐Ÿ’ก Best Practices

  • ๐Ÿงช Start Small: Test with low limits first (depth: 1, limit: 5)
  • ๐ŸŽฏ Use Filters: Focus crawling with include/exclude patterns
  • ๐Ÿค– Respect Robots: Check robots.txt before large crawls
  • โšก Fast Mode: Use ignore_rendering: true for speed
  • ๐Ÿ“„ Main Content: Enable for cleaner, AI-ready content
  • ๐Ÿ“Š Monitor: Watch crawl progress for large jobs
  • ๐Ÿ”„ Iterate: Refine patterns based on results

โšก Performance Notes

ScenarioTypical TimeRecommendation
Single page< 5 secondsUse for quick extractions
Small site (< 20 pages)30-60 secondsGood for testing patterns
Medium site (50-100 pages)2-5 minutesMonitor progress
Large site (> 200 pages)5-15 minutesUse during off-hours

๐Ÿ’ก Tip: Deep crawling (depth > 3) can exponentially increase page count. Use cautiously!

๐Ÿ”ง Troubleshooting

โŒ Authentication Errors

"API key is required" error:

  • โœ… Verify API key in Dify plugin settings
  • โœ… Check Base URL is correct (https://api.watercrawl.dev)
  • โœ… Ensure API key starts with wc-
๐Ÿšซ Crawl Failures

"Failed to crawl" error:

  • โœ… Test URL accessibility in browser
  • โœ… Check if site blocks bots (robots.txt)
  • โœ… Verify API key is valid and has credits
  • โœ… Try with ignore_rendering: true
๐Ÿ“‰ Incomplete Results

Missing expected pages:

  • โœ… Verify URL patterns with pattern tester
  • โœ… Check crawl depth covers target pages
  • โœ… Increase page limit if needed
  • โœ… Review site's internal linking structure
๐ŸŒ Performance Issues

Crawling too slow:

  • โœ… Enable ignore_rendering for 3x speed boost
  • โœ… Use more specific URL patterns
  • โœ… Reduce crawl depth and page limits
  • โœ… Check your internet connection

๐Ÿ“Š Rate Limits & Usage

โ˜๏ธ Watercrawl Cloud

  • Free Plan: 1000 pages/month
  • Startup Plan: 10,000 pages/month
  • Pro Plan: 30,000 pages/month
  • Enterprise: Custom limits
  • Check pricing for more details
  • Monitor usage: Dashboard

๐Ÿ  Self-Hosted

  • No external rate limits
  • Performance depends on your server resources

๐Ÿ”’ Security & Privacy

  • ๐Ÿ” Secure: All API calls use HTTPS encryption
  • ๐Ÿ”‘ API Keys: Store securely, never in client-side code
  • ๐Ÿ  Self-Hosted: Full data control for sensitive content
  • ๐Ÿ“‹ Review: Always check crawled content before use

๐Ÿ“ž Support & Resources

๐Ÿ†˜ Get Help

  • ๐Ÿ’ฌ Plugin Issues: support@watercrawl.dev
  • ๐Ÿ“– Documentation: docs.watercrawl.dev
  • ๐Ÿ› Bug Reports: GitHub Issues

๐Ÿ“š Learn More

  • ๐Ÿ“– API Reference
  • ๐Ÿ  Self-Hosting Guide
  • ๐Ÿš€ Watercrawl Website

Made with โค๏ธ by the Watercrawl Team

Website โ€ข Documentation โ€ข Dashboard