by shaba · v0.1.5
Crawl static documentation sites and wikis (VitePress, MkDocs, Docusaurus, Sphinx, MediaWiki, ...) into clean per-page markdown.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
A Dify datasource plugin that crawls a static documentation site or wiki and extracts clean, per-page markdown — ready to feed a Dify Knowledge base.
It works with any site that exposes either a sitemap.xml or navigable in-page links,
which covers the common static-site generators — VitePress, MkDocs, Docusaurus, Sphinx,
Hugo, Jekyll, GitBook, mdBook and similar — as well as MediaWiki wikis (crawled via
the MediaWiki API). Page content is converted to markdown with Trafilatura.
Choose a crawl strategy:
api.php, allpages) and render
each page via action=parse.url (required) — documentation site URL, or a bare domain for auto-discovery.strategy — auto | html | mediawiki (default auto).discover — when a bare domain is given, probe common documentation paths (e.g. /docs/).path_prefix — restrict the crawl to a path, e.g. /docs/.include_paths / exclude_paths — comma-separated glob patterns to include or skip.max_pages — stop after N pages (default 50).max_depth — link-follow depth from the start URL (default 5).python3 -m pytest -q
ruff check .
yamllint .
The crawler logic (discovery, sitemap parsing, BFS link crawl, MediaWiki, and content
extraction) lives in the docs_crawler package, which is independent of the Dify SDK and
covered by unit tests with mocked network calls.
Apache-2.0. Copyright © 2026 Alexey Shabalin.
https://github.com/shaba/dify-datasource-litecrawl — issues and pull requests welcome.