by zyileven ยท v0.2.1
Enterprise-level multi-GPU document parsing service powered by MinerU
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Enterprise-level multi-GPU document parsing service powered by MinerU
MinerU Tianshu is a powerful Dify plugin that enables high-quality document parsing capabilities through MinerU's enterprise-level infrastructure.
๐ Prerequisites: This plugin requires the MinerU Tianshu API Server to be deployed first. Please deploy the upstream project before installing this plugin.
Convert PDFs, images, and Office documents into structured Markdown format with support for:
This plugin provides 3 tools for flexible document processing workflows:
parse_document - One-click document parsing with automatic wait
parse_document_async - Submit and continue workflow
get_parse_result - Retrieve results later
This plugin requires the MinerU Tianshu API Server to be deployed and running. The API server is provided by the upstream project:
Upstream Project: https://github.com/magicyuan876/mineru-tianshu
You MUST deploy this project before using this plugin, as it provides all the backend API services for document parsing. Please follow the deployment instructions in the upstream project repository.
For self-hosted Dify instances, you MUST configure the FILES_URL environment variable in your Dify server's .env file:
# Add this to your Dify server's .env file
FILES_URL=http://your-dify-server:port
# Example: FILES_URL=http://localhost:3000
# Example: FILES_URL=https://your-dify-domain.com
This allows the plugin to download files from your Dify instance. Without this configuration, you'll see errors like:
Error: Invalid file URL '/files/...': Request URL is missing an 'http://' or 'https://' protocol
Note: Dify Cloud users don't need this configuration as it's already set up.
Install the plugin in your Dify instance:
Configure API Server:
http://localhost:8100Start using in your workflows or agents!
Download the plugin package from GitHub Releases
Upload to Dify:
Configure and use as described in Option 1
Tool: parse_document
Inputs:
- file: {{uploaded_document}}
- backend: pipeline
- lang: ch
- formula_enable: true
- table_enable: true
- max_wait_time: 300
Output:
{{markdown_content}}
Step 1: Submit Document
Tool: parse_document_async
Inputs:
- file: {{document}}
- priority: 5
Output:
{{parse_document_async.text}} # Returns task_id directly as a string
Step 2: Retrieve Result Later
Tool: get_parse_result
Inputs:
- task_id: {{parse_document_async.text}} # Use the task_id from previous step
Output:
{{markdown_content}}
Create an agent that:
parse_document to convert PDF to Markdownhttp://your-server:porthttp://localhost:8100pipeline (Recommended): Balanced performance and accuracyvlm-transformers: Vision-language model with Transformersvlm-vllm-engine: Optimized VLM engine for large-scale processingch: Chinese (Simplified)en: Englishkorean: Koreanjapan: JapaneseIMPORTANT: This plugin requires the MinerU Tianshu API Server from the upstream project:
๐ Upstream Project: https://github.com/magicyuan876/mineru-tianshu
Clone the upstream project:
git clone https://github.com/magicyuan876/mineru-tianshu.git
cd mineru-tianshu
Follow the deployment instructions in the upstream project's README:
Verify the API server is running:
8100 (or as configured)http://your-server:8100/health (if available)Use the API server URL when configuring this plugin in Dify:
http://localhost:8100http://your-server-ip:8100For detailed deployment instructions, configuration options, and troubleshooting, please refer to the upstream project documentation.
Error: "API Server URL is not configured"
Error: "Network error: Connection refused"
Error: "Timeout: Processing exceeded 300 seconds"
max_wait_time parameterparse_document_async + get_parse_result for large documentsWarning: "Task completed but no content found"
Choose the right backend:
pipeline: Best for general documentsvlm-vllm-engine: Best for large-scale batch processingUse async mode for large documents:
parse_document_asyncget_parse_resultOptimize parameters:
Clone the repository
git clone https://github.com/zyileven/mineru-tianshu.git
cd mineru-tianshu
Install dependencies
pip install -r requirements.txt
Configure environment
.env.example to .envRun the plugin
python -m main
mineru-tianshu/
โโโ manifest.yaml # Plugin metadata
โโโ provider/
โ โโโ mineru-tianshu.yaml # Provider configuration
โ โโโ mineru-tianshu.py # Provider implementation
โโโ tools/
โ โโโ parse_document.yaml # Sync tool definition
โ โโโ parse_document.py # Sync tool implementation
โ โโโ parse_document_async.yaml
โ โโโ parse_document_async.py
โ โโโ get_parse_result.yaml
โ โโโ get_parse_result.py
โโโ requirements.txt
โโโ LICENSE
โโโ README.md
# Run unit tests
pytest tests/
Contributions are welcome! Please:
git checkout -b feature/AmazingFeature)git commit -m 'Add some AmazingFeature')git push origin feature/AmazingFeature)See CONTRIBUTING.md for more details.
Apache License 2.0 - see LICENSE for details.
If you encounter any issues or have questions:
We strive to respond to all support requests within 48 hours.
Made with โค๏ธ by zyileven
โญ If you find this plugin helpful, please consider giving it a star on GitHub!