by wonendieee · v0.1.2
Extract text-only physical pages from PDF/Word/Excel or split selected PDF ranges into native file batches. Images are not Base64-encoded into JSON; processing stays local.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Document Cutter is a Dify tool plugin for extracting selected page ranges from PDF, Word, and Excel files. It returns text-only page chunks or native Dify files without placing image binaries in JSON.
Source repository: https://github.com/wonendieee/document_cutter
Install the packaged .difypkg from Dify's plugin management page, then add the Split Document tool to a workflow. The plugin requires no API keys, credentials, external services, or outbound network access.
.pdf), Word (.docx), Excel (.xlsx), and legacy Excel (.xls) inputs.pages_per_chunk pages.1-10, 5-, -3, 7, or empty for all pages.Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Source document file. |
split_mode | select | Yes | page_text returns text-only chunks. page_file returns one or more native file batches. |
page_range | string | No | 1-based inclusive page range. Empty means all pages. |
pages_per_chunk | number | No | Maximum consecutive pages per text chunk or PDF file batch. |
output_filename | string | No | Optional output file base name without extension. |
| Input | Meaning |
|---|---|
1-2 | Pages 1 through 2 |
5- | Page 5 through the end |
-3 | First 3 pages |
7 | Page 7 only |
| empty | All pages |
For Word documents, page boundaries are based on manual page breaks stored in the DOCX file. For Excel documents, pages refer to sheet indexes.
page_text returns JSON containing text-only chunks and physical-page metadata. Embedded images are not Base64-encoded into JSON.
page_file returns JSON batch metadata and one native Dify blob per generated file. A 10-page PDF with pages_per_chunk=4 produces batches for pages 1–4, 5–8, and 9–10.
All processing is performed in the Dify plugin runtime. No document content is sent to third-party services.
MIT