by scrapeunblocker · v0.0.2
Web scraping and Google SERP API. Scrape any website and bypass anti-bot protections (Cloudflare, DataDome, PerimeterX, Akamai) with residential proxies. Returns rendered HTML or AI-parsed structured data.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Web scraping and Google SERP API as a Dify tool. Scrape any website, bypass anti-bot protections, and pull Google search results directly into your Dify apps, agents and workflows.
ScrapeUnblocker renders every page in a real browser behind anti-bot systems such as Cloudflare, DataDome, PerimeterX and Akamai, using residential proxies, and returns the result as rendered HTML or AI-parsed structured data. Use it whenever a plain HTTP request returns a block page, a captcha, or an empty shell instead of the real content.
get_page_source)Scrapes a single web page and returns its fully rendered content.
| Parameter | Required | Description |
|---|---|---|
url | yes | Full URL of the page to scrape, including the scheme |
parsed_data | no | Return AI-parsed structured JSON instead of raw HTML |
proxy_country | no | Two-letter country code for the exit IP, for geo-restricted content |
Returns the page content as a text message, plus a JSON message with success,
url, parsed_data and content_length.
search_google)Scrapes a Google search engine results page (SERP) and returns the organic results.
| Parameter | Required | Description |
|---|---|---|
keyword | yes | The search term to look up on Google |
pages_to_check | no | How many result pages to scrape (default 1) |
proxy_country | no | Two-letter country code for country-specific results |
Each organic result carries title, url, description and position.
Two optional settings are available: API Base URL (override to point at a different environment) and Timeout Seconds (default 180, since rendering a protected page in a browser takes longer than a plain HTTP request).