Document Cutter · emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginsDocument Cutter
Plugin
Limited listing

Document Cutter

by wonendieee · v0.1.2

Extract text-only physical pages from PDF/Word/Excel or split selected PDF ranges into native file batches. Images are not Base64-encoded into JSON; processing stays local.

667 installsUpdated Jul 21, 2026
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Tools

Available inside your emploidai workspace after installation.

Data sources

Available inside your emploidai workspace after installation.

Category

tool

Version

0.1.2wonendieee

Requirements

Maximum memory 256MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.0.0+

Permissions

  • Uses tool capability

Dependencies

No additional dependencies

Resources

Privacy policy
emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

Document Cutter

Document Cutter is a Dify tool plugin for extracting selected page ranges from PDF, Word, and Excel files. It returns text-only page chunks or native Dify files without placing image binaries in JSON.

Source repository: https://github.com/wonendieee/document_cutter

Setup

Install the packaged .difypkg from Dify's plugin management page, then add the Split Document tool to a workflow. The plugin requires no API keys, credentials, external services, or outbound network access.

Features

  • Supports PDF (.pdf), Word (.docx), Excel (.xlsx), and legacy Excel (.xls) inputs.
  • Extracts text-only physical-page chunks and preserves empty pages for stable page numbering.
  • Splits PDF ranges into consecutive native-file batches; each batch contains at most pages_per_chunk pages.
  • Returns non-PDF selected ranges as same-format native files.
  • Supports flexible page range syntax: 1-10, 5-, -3, 7, or empty for all pages.
  • Processes files locally inside the Dify plugin runtime.

Tools

Split Document

Parameters:

NameTypeRequiredDescription
filefileYesSource document file.
split_modeselectYespage_text returns text-only chunks. page_file returns one or more native file batches.
page_rangestringNo1-based inclusive page range. Empty means all pages.
pages_per_chunknumberNoMaximum consecutive pages per text chunk or PDF file batch.
output_filenamestringNoOptional output file base name without extension.

Page Range Examples

InputMeaning
1-2Pages 1 through 2
5-Page 5 through the end
-3First 3 pages
7Page 7 only
emptyAll pages

For Word documents, page boundaries are based on manual page breaks stored in the DOCX file. For Excel documents, pages refer to sheet indexes.

Output Modes

page_text returns JSON containing text-only chunks and physical-page metadata. Embedded images are not Base64-encoded into JSON.

page_file returns JSON batch metadata and one native Dify blob per generated file. A 10-page PDF with pages_per_chunk=4 produces batches for pages 1–4, 5–8, and 9–10.

Privacy

All processing is performed in the Dify plugin runtime. No document content is sent to third-party services.

License

MIT