<<<<<<< HEAD
๐ Bright Data Plugin for Dify
A comprehensive web scraping and data extraction plugin powered by Bright Data's enterprise-grade infrastructure. This plugin provides advanced web scraping capabilities with anti-bot detection for the Dify platform.
โจ New! Smart Data Extractor
Our latest addition automatically determines the best extraction method for your needs:
- ๐ง Intelligent Auto-Detection: Just describe what you want or provide any URL
- ๐ฏ 20+ Data Sources: Amazon, LinkedIn, Instagram, YouTube, TikTok, Crunchbase, and more!
- โก One Tool for Everything: No need to choose between multiple extractors
๐ง Features
Smart Data Extractor (Recommended)
- E-commerce: Amazon products/reviews/search, Walmart, eBay, Best Buy
- Social Media: LinkedIn profiles/companies/jobs, Instagram, TikTok, Facebook, X/Twitter
- Business Intelligence: Crunchbase companies, Google Maps reviews, ZoomInfo
- Content & Media: YouTube videos/comments, Reddit posts, GitHub repositories
- Real Estate: Zillow property listings
Classic Tools
- Web Page Scraping: Extract content as Markdown or HTML with anti-bot protection
- Search Engine Scraping: Google, Bing, Yandex search results
๐ Quick Start
Option 1: Download Pre-built Package (Easiest)
-
Download Plugin Package
- Go to Releases
- Download the latest
bright-data.difypkg file
-
Install in Dify
- Open your Dify dashboard
- Go to Plugins โ Install from file
- Upload the
bright-data.difypkg file
- Click Install
-
Configure API Token
- Click on the installed Bright Data plugin
- Add your Bright Data API Token
- Save configuration
-
Start Using!
- Create workflows and use the Smart Data Extractor
- Or use individual tools like Web Scraper and Search Engine
Option 2: Development/Debug Mode
Only use this if you want to modify the plugin:
- Clone Repository
git clone https://github.com/Idanvilenski/BrightData_Dify_Plugin.git
cd BrightData_Dify_Plugin/bright-data
- Set Up Environment
pip install -r requirements.txt
- Configure Debug Connection
cp .env.example .env
# Edit .env with your Dify debug credentials
- Run Development Server
python -m main
๐ Getting Bright Data API Token
- Sign up at Bright Data
- Go to your dashboard โ Zones
- Create a new zone or use existing one
- Copy the API Token
- Use this token when configuring the plugin in Dify
๐ก Usage Examples
Smart Data Extractor Usage
The Smart Data Extractor automatically detects what you want to extract:
Example 1: Amazon Product Analysis
Request: "Get Amazon product details including price and reviews"
URL: "https://amazon.com/dp/B08N5WRWNW"
โ Automatically uses Amazon extractor
Example 2: LinkedIn Profile Research
Request: "Extract LinkedIn professional profile information"
URL: "https://linkedin.com/in/elonmusk"
โ Automatically uses LinkedIn extractor
Example 3: Instagram Analytics
Request: "Get Instagram profile statistics and engagement"
URL: "https://instagram.com/instagram"
โ Automatically uses Instagram extractor
Example 4: Business Intelligence
Request: "Find company funding and investment data"
URL: "https://crunchbase.com/organization/openai"
โ Automatically uses Crunchbase extractor
Workflow Integration
- Create a new Workflow in Dify
- Add Smart Data Extractor node
- Configure parameters:
- Request: Describe what you want
- URL: Paste any supported URL
- Additional Parameters: Leave empty (optional)
- Connect to LLM node to process the extracted data
- Run and enjoy!
๐ฏ Supported Platforms
| Category | Platforms |
|---|
| E-commerce | Amazon, Walmart, eBay, Best Buy, Home Depot, Zara, Etsy |
| Social Media | LinkedIn, Instagram, TikTok, Facebook, X/Twitter, YouTube, Reddit |
| Business Intel | Crunchbase, ZoomInfo, Google Maps, Yahoo Finance |
| Content & Apps | GitHub, Reuters News, Google Play, Apple App Store |
| Real Estate | Zillow, Google Shopping |
๐ง Configuration Parameters
Smart Data Extractor
- Request (required): Describe what data you want to extract
- URL (optional): Specific URL to extract from
- Additional Parameters (optional): JSON with extra settings like
{"pages_to_search": "3", "num_of_reviews": "50"}
Web Scraper
- URL (required): Webpage to scrape
- Output format: Markdown or HTML
Search Engine
- Query (required): Search terms
- Engine: Google (default), Bing, or Yandex
๐จ Troubleshooting
Common Issues
"Could not determine appropriate tool"
- Be more specific in your request
- Use full URLs with https://
- Try: "Get Amazon product info" instead of "get data"
"API Error 400: Invalid input"
- Check that URLs are complete and valid
- Ensure your BrightData API token is correct
- Verify you have sufficient credits
Plugin not connecting
- Restart the plugin:
Ctrl+C then python -m main
- Check your
.env file has correct debug credentials
- Refresh Dify and re-add the plugin
Debug Mode Issues
"Failed to connect to localhost:5003"
- Make sure Dify is running with plugin daemon
- Get fresh debug credentials from Dify
- Update your
.env file with new credentials
๐ What's New in v0.1.0
- โ
Smart Data Extractor: Auto-detects best extraction method
- โ
20+ Data Sources: Massive expansion from 3 to 20+ platforms
- โ
Intelligent URL Detection: Automatically selects tools based on URLs
- โ
Natural Language Requests: Describe what you want in plain English
- โ
Better Error Handling: Helpful suggestions when things go wrong
- โ
Robust API Integration: Handles various response formats gracefully
๐ค Contributing
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature
- Make your changes and test thoroughly
- Commit:
git commit -m 'Add amazing feature'
- Push:
git push origin feature/amazing-feature
- Submit a pull request
๐ License
This project is licensed under the MIT License - see the LICENSE file for details.
๐ Support
๐ Credits
- Bright Data: For providing the enterprise-grade web scraping infrastructure
- Dify Team: For creating an amazing AI workflow platform
- Contributors: Thank you to everyone who helps improve this plugin!
=======
<<<<<<< HEAD
๐ Bright Data Plugin for Dify
A comprehensive web scraping and data extraction plugin powered by Bright Data's enterprise-grade infrastructure. This plugin provides advanced web scraping capabilities with anti-bot detection for the Dify platform.
โจ New! Smart Data Extractor
Our latest addition automatically determines the best extraction method for your needs:
- ๐ง Intelligent Auto-Detection: Just describe what you want or provide any URL
- ๐ฏ 20+ Data Sources: Amazon, LinkedIn, Instagram, YouTube, TikTok, Crunchbase, and more!
- โก One Tool for Everything: No need to choose between multiple extractors
๐ง Features
Smart Data Extractor (Recommended)
- E-commerce: Amazon products/reviews/search, Walmart, eBay, Best Buy
- Social Media: LinkedIn profiles/companies/jobs, Instagram, TikTok, Facebook, X/Twitter
- Business Intelligence: Crunchbase companies, Google Maps reviews, ZoomInfo
- Content & Media: YouTube videos/comments, Reddit posts, GitHub repositories
- Real Estate: Zillow property listings
Classic Tools
- Web Page Scraping: Extract content as Markdown or HTML with anti-bot protection
- Search Engine Scraping: Google, Bing, Yandex search results
๐ Quick Start
Option 1: Download Pre-built Package (Easiest)
-
Download Plugin Package
- Go to Releases
- Download the latest
bright-data.difypkg file
-
Install in Dify
- Open your Dify dashboard
- Go to Plugins โ Install from file
- Upload the
bright-data.difypkg file
- Click Install
-
Configure API Token
- Click on the installed Bright Data plugin
- Add your Bright Data API Token
- Save configuration
-
Start Using!
- Create workflows and use the Smart Data Extractor
- Or use individual tools like Web Scraper and Search Engine
Option 2: Development/Debug Mode
Only use this if you want to modify the plugin:
- Clone Repository
git clone https://github.com/Idanvilenski/BrightData_Dify_Plugin.git
cd BrightData_Dify_Plugin/bright-data
- Set Up Environment
pip install -r requirements.txt
- Configure Debug Connection
cp .env.example .env
# Edit .env with your Dify debug credentials
- Run Development Server
python -m main
๐ Getting Bright Data API Token
- Sign up at Bright Data
- Go to your dashboard โ Zones
- Create a new zone or use existing one
- Copy the API Token
- Use this token when configuring the plugin in Dify
๐ก Usage Examples
Smart Data Extractor Usage
The Smart Data Extractor automatically detects what you want to extract:
Example 1: Amazon Product Analysis
Request: "Get Amazon product details including price and reviews"
URL: "https://amazon.com/dp/B08N5WRWNW"
โ Automatically uses Amazon extractor
Example 2: LinkedIn Profile Research
Request: "Extract LinkedIn professional profile information"
URL: "https://linkedin.com/in/elonmusk"
โ Automatically uses LinkedIn extractor
Example 3: Instagram Analytics
Request: "Get Instagram profile statistics and engagement"
URL: "https://instagram.com/instagram"
โ Automatically uses Instagram extractor
Example 4: Business Intelligence
Request: "Find company funding and investment data"
URL: "https://crunchbase.com/organization/openai"
โ Automatically uses Crunchbase extractor
Workflow Integration
- Create a new Workflow in Dify
- Add Smart Data Extractor node
- Configure parameters:
- Request: Describe what you want
- URL: Paste any supported URL
- Additional Parameters: Leave empty (optional)
- Connect to LLM node to process the extracted data
- Run and enjoy!
๐ฏ Supported Platforms
| Category | Platforms |
|---|
| E-commerce | Amazon, Walmart, eBay, Best Buy, Home Depot, Zara, Etsy |
| Social Media | LinkedIn, Instagram, TikTok, Facebook, X/Twitter, YouTube, Reddit |
| Business Intel | Crunchbase, ZoomInfo, Google Maps, Yahoo Finance |
| Content & Apps | GitHub, Reuters News, Google Play, Apple App Store |
| Real Estate | Zillow, Google Shopping |
๐ง Configuration Parameters
Smart Data Extractor
- Request (required): Describe what data you want to extract
- URL (optional): Specific URL to extract from
- Additional Parameters (optional): JSON with extra settings like
{"pages_to_search": "3", "num_of_reviews": "50"}
Web Scraper
- URL (required): Webpage to scrape
- Output format: Markdown or HTML
Search Engine
- Query (required): Search terms
- Engine: Google (default), Bing, or Yandex
๐จ Troubleshooting
Common Issues
"Could not determine appropriate tool"
- Be more specific in your request
- Use full URLs with https://
- Try: "Get Amazon product info" instead of "get data"
"API Error 400: Invalid input"
- Check that URLs are complete and valid
- Ensure your BrightData API token is correct
- Verify you have sufficient credits
Plugin not connecting
- Restart the plugin:
Ctrl+C then python -m main
- Check your
.env file has correct debug credentials
- Refresh Dify and re-add the plugin
Debug Mode Issues
"Failed to connect to localhost:5003"
- Make sure Dify is running with plugin daemon
- Get fresh debug credentials from Dify
- Update your
.env file with new credentials
๐ What's New in v0.1.0
- โ
Smart Data Extractor: Auto-detects best extraction method
- โ
20+ Data Sources: Massive expansion from 3 to 20+ platforms
- โ
Intelligent URL Detection: Automatically selects tools based on URLs
- โ
Natural Language Requests: Describe what you want in plain English
- โ
Better Error Handling: Helpful suggestions when things go wrong
- โ
Robust API Integration: Handles various response formats gracefully
๐ค Contributing
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature
- Make your changes and test thoroughly
- Commit:
git commit -m 'Add amazing feature'
- Push:
git push origin feature/amazing-feature
- Submit a pull request
๐ License
This project is licensed under the MIT License - see the LICENSE file for details.
๐ Support
๐ Credits
- Bright Data: For providing the enterprise-grade web scraping infrastructure
- Dify Team: For creating an amazing AI workflow platform
- Contributors: Thank you to everyone who helps improve this plugin!
=======
BrightData_Dify_Plugin
A BrightData plugin for the Dify platform, the plugin contains all of Bright's web scraping, unlocking and dataset tools
c79671e0e242770add67da64e39647731b1ebdd5
3fcdbbf934150efbd3a2a0c0432ed5680181d5a2