File Upload to Knowledge Base · emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginsFile Upload to Knowledge Base
Plugin
Limited listing

File Upload to Knowledge Base

by r3-yamauchi · v0.0.1

Upload various file types (PDF, TXT, DOCX, Markdown, etc.) directly to Dify's knowledge base via API

3.2k installsUpdated Nov 10, 2025
Documentation
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Tools

Available inside your emploidai workspace after installation.

Data sources

Available inside your emploidai workspace after installation.

Category

tool

Version

0.0.1r3-yamauchi

Requirements

Maximum memory 256MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.0.0+

Permissions

  • Uses tool capability
  • Requires encrypted tool credentials

Dependencies

No additional dependencies

Resources

DocumentationPrivacy policy
emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

knowledgebase_update

Author: r3-yamauchi
Version: 0.0.1
Type: tool

English | Japanese

Description

knowledgebase_update adds two focused tools to your Dify workspace—update (text) and upload (file)—so you can register documents in a knowledge base dataset without writing HTTP requests. Both tools wrap Dify's knowledge APIs, handling validation, UPSERT behaviour, and optional chunking configuration for you.

The source code of this plugin is available in the GitHub repository.

Ask DeepWiki

Features

  • Maintain a text-first update tool that automatically chooses between create-by-text and update-by-text.
  • Provide an upload tool that accepts Dify File objects, base64 strings, or downloadable URLs and invokes the file-based APIs.
  • Target documents by ID with the optional document_id parameter, ensuring updates land on the intended record.
  • Perform document UPSERTs by searching for matching names before issuing updates.
  • Share validation of indexing options and process_rule JSON across both tools for consistent behaviour.
  • Enforce compatible chunking and language combinations so that Q&A mode always includes a language hint.
  • Surface clear error hints for common HTTP responses and log each operation outcome for auditability.
  • Honour document names: when name is provided it is stored and used (case-insensitively) to find existing documents even when document_id is omitted.
  • Centralise credential handling via provider settings (api_uri, api_secret) so workflows stay tidy.
  • Emit both a human-readable summary and a JSON payload that surfaces the resulting document_id for downstream steps.

Prerequisites

  • Dify SaaS or self-hosted instance with plugin support enabled
  • API Secret with permission to manage the target knowledge base dataset
  • Dataset ID for the destination knowledge base

Installation & Setup

  1. Install the plugin from the Dify Marketplace or upload this package manually.
  2. In the provider settings, supply:
    • API URI – e.g. https://api.dify.ai/v1 for SaaS
    • API Secret – knowledge-base API key created in the target dataset
    • (Optional) Timeout (seconds) – overrides the default 30-second HTTP timeout
    • (Optional) Retry Attempts – number of retries for network errors or 5xx responses (default 0)
  3. Add either the update or upload tool node to a workflow and configure the parameters described below.

Tool Parameters

The plugin ships two tools that share the same validation and UPSERT logic:

  • update: accepts raw text and calls Dify's /document/create-by-text or /documents/{id}/update-by-text.
  • upload: accepts file inputs and calls /document/create-by-file or /documents/{id}/update-by-file.

update tool (text upload)

ParameterRequiredDefaultDescription
dataset_id✅–Target dataset ID in the knowledge base.
document_id❌–Existing document identifier used to update directly; validated before the request is sent.
name✅ (unless document_id is provided)–Document name applied to the record and used to resolve UPSERTs when no document_id is given (case-insensitive).
text✅–Raw text content that will be indexed.
indexing_technique✅high_qualityIndexing quality mode (high_quality or economy).
doc_form❌–Chunking strategy (text_model, hierarchical_model, or qa_model).
doc_language❌–Document language hint. Required when doc_form is qa_model; otherwise must be omitted.
process_rule✅{"mode": "automatic"}JSON string describing preprocessing and segmentation.
timeout_seconds❌–Per-tool timeout in seconds; falls back to the provider value when omitted.
retry_attempts❌–Per-tool retry count for network errors and 5xx responses; defaults to the provider setting.

If name is omitted for the upload tool and no filename can be derived from the file payload, the tool raises an error.

upload tool (file upload)

ParameterRequiredDefaultDescription
dataset_id✅–Target dataset ID in the knowledge base.
document_id❌–Existing document identifier used to update directly; validated before the request is sent.
name❌–Document name applied when creating or updating; also used to locate existing records when document_id is omitted (case-insensitive). The uploaded multipart filename mirrors this value—if the name has no extension, the original file extension is appended. Falls back to the uploaded filename when blank.
file✅–Dify file object or compatible file payload. The original file name is reused when name is empty.
indexing_technique✅high_qualityIndexing quality mode (high_quality or economy).
doc_form❌–Chunking strategy (text_model, hierarchical_model, or qa_model).
doc_language❌–Document language hint. Required when doc_form is qa_model; otherwise must be omitted.
process_rule✅{"mode": "automatic"}JSON string describing preprocessing and segmentation.
timeout_seconds❌–Per-tool timeout in seconds; falls back to the provider value when omitted.
retry_attempts❌–Per-tool retry count for network errors and 5xx responses; defaults to the provider setting.

Processing Flow

  1. Read api_uri / api_secret from provider credentials and assert dataset_id is present in tool parameters.
  2. Validate required content inputs (text for update, file for upload).
  3. Parse process_rule JSON and verify mode-specific requirements (automatic vs custom/hierarchical).
  4. If document_id is provided, update the document directly; otherwise look up by name (case-insensitive) to decide whether to update or create. When name has no extension, the source file extension is appended before uploading so Dify receives a valid filename.
  5. Call the corresponding Dify endpoint (create-by-text / update-by-text or create-by-file / update-by-file) with the prepared payload.
  6. Return the JSON response with helpful defaults (operation, fallback document.id when absent).

Usage Example

nodes:
  - type: file_input
    name: pdf_input
    config:
      file_type: pdf

  - type: tool
    name: upload_to_kb
    tool: upload
    parameters:
      dataset_id: "550e8400-e29b-41d4-a716-446655440000"
      file: "{{pdf_input}}"
      indexing_technique: "high_quality"
      process_rule: '{"mode": "automatic"}'
nodes:
  - type: tool
    name: update_summary
    tool: update
    parameters:
      dataset_id: "550e8400-e29b-41d4-a716-446655440000"
      name: "Weekly Report"
      text: "{{workflow.outputs.summary}}"
      indexing_technique: "economy"
      process_rule: '{"mode": "automatic"}'

The tool returns the JSON payload produced by Dify, including the detected operation and document.id.

Troubleshooting

  • Missing credentials – Ensure both api_uri and api_secret are configured on the provider before running workflows.
  • Required content missing – update needs text, while upload needs file.
  • Document name is required – The update tool always needs name; the upload tool can infer it from the file metadata when omitted.
  • Document ID not found – When document_id is provided but the dataset does not contain it, the tool raises an error before sending the update.
  • Invalid JSON format for process_rule – Check quoting/escaping; the plugin expects a valid JSON string.
  • Failed to create/update document: ... – Review the Dify response for details (e.g. dataset ID mismatch, file too large, permission error).
  • doc_language can only be provided when doc_form is set to qa_model – Adjust the chunking mode or remove the language hint.
  • HTTP 401/403/404 – The plugin now adds hints to these errors; verify credentials, permissions, or dataset IDs according to the message.
  • Request rejected: Verify payload fields and formats – Ensure the uploaded filename includes a valid extension. The tool appends the original file's extension when name omits it; double-check that the provided name or the source file includes one.

Privacy Policy

Data Handling

This plugin processes the following data:

  1. PDF Files: various file types uploaded by users are sent to the knowledge base via the Dify API
  2. API Credentials: API URI stored in plugin settings and API Secret supplied per tool invocation are used for authenticating requests to the Dify API
  3. Knowledge Base ID: Dataset ID specified as a tool parameter is used to identify the upload destination

Data Storage

  • The plugin does not store data locally
  • All data is sent to the specified knowledge base via the Dify API
  • API URI is managed by Dify's plugin system. API secrets should be injected at runtime through Dify workflows or environment secrets so they are not persisted with the provider configuration.

Security

  • API communications are encrypted via HTTPS
  • API secrets are not stored by the provider; inject them via secure workflow variables or environment secrets at execution time.

Third-Party Disclosure

This plugin does not send data to any third parties other than the Dify API endpoint specified by the user.

License

This project is released under the MIT License.