by hjlarry · v0.0.1
Add the content produced by dify's workflow to its knowledge base
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Author: hjlarry
Version: 0.0.1
Type: tool
Repo: https://github.com/hjlarry/dify-plugin-knowledge
The API_URL and API_KEY are comes from here:
The Dataset ID is from here:
the PROCESS RULE can be default {"mode": "automatic"}, but when you want to custom your process rules, you can config like this:
When your CHUNK METHOD is General or Q&A, the PROCESS RULE should be like this:
{
"mode": "custom",
"rules":{
"pre_processing_rules": [{"id":"remove_extra_spaces", "enabled": true}, {"id":"remove_urls_emails", "enabled": true}],
"segmentation": {
"separator": "\n\n",
"max_tokens": 1024,
"chunk_overlap": 50
}
}
}
When your CHUNK METHOD is Parent-child, the PROCESS RULE should be like this:
{
"mode": "hierarchical",
"rules":{
"pre_processing_rules": [{"id":"remove_extra_spaces", "enabled": true}, {"id":"remove_urls_emails", "enabled": true}],
"segmentation": {
"separator": "\n\n",
"max_tokens": 1024,
"chunk_overlap": 50
}
},
"parent_mode": "full-doc",
"subchunk_segmentation": {
"separator": "\n\n",
"max_tokens": 1024,
"chunk_overlap": 50
}
}
You can run the plugin like this, follow a PDF parsing plugin like MinerU