by langgenius · v0.1.0
Dify Extractor
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Dify Extractor converts uploaded documents into Markdown plus structured document records for RAG pipelines and workflows.
.pdf).docx).pptx).xls, .xlsx).md, .markdown, .mdx).htm, .html).csv).json).yaml, .yml).txt, .log, .rst, .ini, .cfg, .conf, and .xmlText MIME types are also accepted when a file has no recognized extension. Unsupported binary formats return a clear error instead of being decoded as text.
documents contains format-appropriate records: rows for CSV/Excel, pages for PDF, sections for
Markdown, and a whole-file record for the other formats.images is emitted when embedded or referenced images are successfully uploaded to Dify.Malformed, empty, undecodable, or resource-limit-breaking files return one error message and do not emit partial document variables.
Remote Markdown and DOCX images are imported on a best-effort basis. Only HTTP(S) URLs are used, with a 30-second request timeout, a maximum of 20 images, a 15 MiB per-image limit, and a 100 MiB cumulative limit. A failed image download or upload does not discard otherwise extractable text.
ZIP-based Office files are rejected before parsing if they contain more than 10,000 entries, expand beyond 200 MiB, or exceed a 1,000:1 compression ratio.
The plugin targets Python 3.12 and uses uv for dependency management.
uv sync --all-groups --frozen --python 3.12
uv run pytest
uv run ruff check .
uv run ruff format --check .