by r3-yamauchi · v0.0.1
Upload various file types (PDF, TXT, DOCX, Markdown, etc.) directly to Dify's knowledge base via API
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Author: r3-yamauchi
Version: 0.0.1
Type: tool
English | Japanese
knowledgebase_update adds two focused tools to your Dify workspace—update (text) and upload (file)—so you can register documents in a knowledge base dataset without writing HTTP requests. Both tools wrap Dify's knowledge APIs, handling validation, UPSERT behaviour, and optional chunking configuration for you.
The source code of this plugin is available in the GitHub repository.
update tool that automatically chooses between create-by-text and update-by-text.upload tool that accepts Dify File objects, base64 strings, or downloadable URLs and invokes the file-based APIs.document_id parameter, ensuring updates land on the intended record.process_rule JSON across both tools for consistent behaviour.name is provided it is stored and used (case-insensitively) to find existing documents even when document_id is omitted.api_uri, api_secret) so workflows stay tidy.document_id for downstream steps.https://api.dify.ai/v1 for SaaSupdate or upload tool node to a workflow and configure the parameters described below.The plugin ships two tools that share the same validation and UPSERT logic:
update: accepts raw text and calls Dify's /document/create-by-text or /documents/{id}/update-by-text.upload: accepts file inputs and calls /document/create-by-file or /documents/{id}/update-by-file.update tool (text upload)| Parameter | Required | Default | Description |
|---|---|---|---|
dataset_id | ✅ | – | Target dataset ID in the knowledge base. |
document_id | ❌ | – | Existing document identifier used to update directly; validated before the request is sent. |
name | ✅ (unless document_id is provided) | – | Document name applied to the record and used to resolve UPSERTs when no document_id is given (case-insensitive). |
text | ✅ | – | Raw text content that will be indexed. |
indexing_technique | ✅ | high_quality | Indexing quality mode (high_quality or economy). |
doc_form | ❌ | – | Chunking strategy (text_model, hierarchical_model, or qa_model). |
doc_language | ❌ | – | Document language hint. Required when doc_form is qa_model; otherwise must be omitted. |
process_rule | ✅ | {"mode": "automatic"} | JSON string describing preprocessing and segmentation. |
timeout_seconds | ❌ | – | Per-tool timeout in seconds; falls back to the provider value when omitted. |
retry_attempts | ❌ | – | Per-tool retry count for network errors and 5xx responses; defaults to the provider setting. |
If name is omitted for the upload tool and no filename can be derived from the file payload, the tool raises an error.
upload tool (file upload)| Parameter | Required | Default | Description |
|---|---|---|---|
dataset_id | ✅ | – | Target dataset ID in the knowledge base. |
document_id | ❌ | – | Existing document identifier used to update directly; validated before the request is sent. |
name | ❌ | – | Document name applied when creating or updating; also used to locate existing records when document_id is omitted (case-insensitive). The uploaded multipart filename mirrors this value—if the name has no extension, the original file extension is appended. Falls back to the uploaded filename when blank. |
file | ✅ | – | Dify file object or compatible file payload. The original file name is reused when name is empty. |
indexing_technique | ✅ | high_quality | Indexing quality mode (high_quality or economy). |
doc_form | ❌ | – | Chunking strategy (text_model, hierarchical_model, or qa_model). |
doc_language | ❌ | – | Document language hint. Required when doc_form is qa_model; otherwise must be omitted. |
process_rule | ✅ | {"mode": "automatic"} | JSON string describing preprocessing and segmentation. |
timeout_seconds | ❌ | – | Per-tool timeout in seconds; falls back to the provider value when omitted. |
retry_attempts | ❌ | – | Per-tool retry count for network errors and 5xx responses; defaults to the provider setting. |
api_uri / api_secret from provider credentials and assert dataset_id is present in tool parameters.text for update, file for upload).process_rule JSON and verify mode-specific requirements (automatic vs custom/hierarchical).document_id is provided, update the document directly; otherwise look up by name (case-insensitive) to decide whether to update or create. When name has no extension, the source file extension is appended before uploading so Dify receives a valid filename.create-by-text / update-by-text or create-by-file / update-by-file) with the prepared payload.operation, fallback document.id when absent).nodes:
- type: file_input
name: pdf_input
config:
file_type: pdf
- type: tool
name: upload_to_kb
tool: upload
parameters:
dataset_id: "550e8400-e29b-41d4-a716-446655440000"
file: "{{pdf_input}}"
indexing_technique: "high_quality"
process_rule: '{"mode": "automatic"}'
nodes:
- type: tool
name: update_summary
tool: update
parameters:
dataset_id: "550e8400-e29b-41d4-a716-446655440000"
name: "Weekly Report"
text: "{{workflow.outputs.summary}}"
indexing_technique: "economy"
process_rule: '{"mode": "automatic"}'
The tool returns the JSON payload produced by Dify, including the detected operation and document.id.
api_uri and api_secret are configured on the provider before running workflows.update needs text, while upload needs file.update tool always needs name; the upload tool can infer it from the file metadata when omitted.document_id is provided but the dataset does not contain it, the tool raises an error before sending the update.name omits it; double-check that the provided name or the source file includes one.This plugin processes the following data:
This plugin does not send data to any third parties other than the Dify API endpoint specified by the user.
This project is released under the MIT License.