by asukhodko Β· v2.1.7
Advanced Markdown chunking powered by chunkana library with structural awareness for better RAG performance
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Advanced Markdown Chunker is a Dify plugin that intelligently splits Markdown documents into semantically meaningful chunks optimized for RAG (Retrieval-Augmented Generation) systems. Powered by the chunkana engine, it provides advanced structural awareness that goes beyond simple text splitting, preserving document structure, keeping code blocks intact, and automatically selecting the best chunking strategy based on content analysis.
This plugin is designed primarily for RAG (Retrieval-Augmented Generation) workflows where chunks are embedded and stored in vector databases for semantic search. Built on the robust chunkana library, it provides enterprise-grade chunking capabilities through a user-friendly Dify interface. By default, each chunk includes embedded metadata (header paths, content type, line numbers) directly in the chunk text, which improves retrieval quality by providing additional context for embeddings.
Note for Model Training: If you need clean text without metadata (e.g., for fine-tuning language models), set
include_metadata: falseor post-process chunks to remove the<metadata>block.
| Simple Chunking Problems | Advanced Markdown Chunker Solution |
|---|---|
| Breaks code blocks mid-function | Preserves code blocks as atomic units |
| Loses header context | Maintains hierarchical section structure |
| Splits tables and lists | Keeps tables and lists intact |
| One-size-fits-all approach | 4 adaptive strategies based on content |
| No overlap support | Smart overlap for better retrieval |
| Destroys list hierarchies | Smart list grouping with context binding |
| Breaks nested code examples | Handles nested fencing (````, ``````, ~~~~) |
| Code examples lose explanatory context | Enhanced code-context binding with pattern recognition |
| Before/After comparisons split apart | Intelligent Before/After pairing |
| Code and output separated | Automatic Code+Output binding |
| Mathematical formulas split | LaTeX formula preservation ($$...$$, environments) |
$$...$$) never split across chunks\begin{equation}, \begin{align}) preserved completemax_distance_lines)require_same_section)max_group_size, max_grouped_tables)Local Processing Only
The Plugin processes all Markdown content locally within your Dify instance. No data is transmitted to external services.
What the Plugin does:
What the Plugin does NOT do:
For complete details, see PRIVACY.md.
.difypkg file from Releases.difypkg fileRequirements:
# Clone the repository
git clone https://github.com/asukhodko/dify-markdown-chunker.git
cd dify-markdown-chunker
# Create virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Verify installation
make test
Requirements:
Add the chunker to your Dify workflow:
- node: chunk_markdown
type: tool
tool: advanced_markdown_chunker
config:
max_chunk_size: 2048
strategy: auto
chunk_overlap: 100
include_metadata: true
| Parameter | Type | Default | Description |
|---|---|---|---|
input_text | string | required | Markdown text to chunk |
max_chunk_size | number | 4096 | Maximum chunk size in characters |
chunk_overlap | number | 200 | Base overlap size (adaptive: actual max = min(overlap_size, chunk_size * 0.35)) |
strategy | select | auto | Chunking strategy (auto/code_aware/list_aware/structural/fallback) |
include_metadata | boolean | true | Embed metadata in chunk text (see below) |
enable_hierarchy | boolean | false | Create parent-child relationships between chunks |
debug | boolean | false | Include all chunks (root, intermediate, leaf) in hierarchical mode |
leaf_only | boolean | false | Return only leaf chunks in hierarchical mode (recommended for vector DB) |
The plugin parameters map to chunkana configuration as follows:
| Plugin Parameter | Chunkana Config | Notes |
|---|---|---|
max_chunk_size | max_chunk_size | Direct mapping |
chunk_overlap | overlap_size | Capped at 35% of chunk size |
strategy | strategy_override | auto β automatic selection based on content analysis |
include_metadata | include_metadata | Controls metadata embedding in chunk text |
enable_hierarchy | enable_hierarchy | Enables parent-child chunk relationships |
debug | debug_mode | Controls visibility of all chunk types in hierarchical mode |
leaf_only | leaf_only | Filters to content chunks only (excludes structural headers) |
Advanced chunkana features not exposed in plugin UI:
For direct chunkana usage with full feature access, see the chunkana documentation.
When enable_hierarchy=true, the plugin returns chunks organized in a tree structure with parent-child relationships.
Chunk Types:
| Type | is_root | is_leaf | indexable | Description |
|---|---|---|---|---|
| Root | true | false | false | Document root, covers entire document |
| Internal | false | false | true | Section headers with children |
| Leaf | false | true | true | Content chunks for indexing |
Filtering Behavior:
debug=false (default): Root chunk excluded from resultsdebug=true: All chunks included for debuggingleaf_only=true: Only leaf chunks returned (recommended for vector DB)Recommended Usage for Vector DB:
- node: chunk_for_indexing
type: tool
tool: advanced_markdown_chunker
config:
enable_hierarchy: true
leaf_only: true # Only indexable content chunks
For Debugging Hierarchy:
- node: debug_hierarchy
type: tool
tool: advanced_markdown_chunker
config:
enable_hierarchy: true
debug: true # Include root and internal nodes
chunk_overlapChunk Overlap controls how many characters of context are shared between consecutive chunks to preserve semantic continuity.
Behavior depends on include_metadata:
include_metadata | Overlap Behavior |
|---|---|
true (default) | Overlap stored in metadata fields previous_content / next_content. Chunk content stays clean. |
false | Overlap embedded directly into chunk text: previous_content + "\n" + main + "\n" + next_content |
Example with include_metadata: true:
<metadata>
{
"previous_content": "...end of previous chunk...",
"next_content": "...start of next chunk..."
}
</metadata>
# Current Section
Main content of this chunk...
Example with include_metadata: false:
...end of previous chunk...
# Current Section
Main content of this chunk...
...start of next chunk...
This allows chunk_overlap to work predictably in both modes:
include_metadata: true): Overlap available as structured metadata for embeddingsinclude_metadata: false): Overlap physically present in text for sliding window processinginclude_metadataWhen include_metadata: true (default), each chunk includes a <metadata> block prepended to the content:
<metadata>
{
"content_type": "text",
"header_path": "/Installation/Requirements",
"start_line": 45,
"end_line": 52
}
</metadata>
# Requirements
Python 3.12 or higher is required...
Typical metadata fields:
content_type β type of content (text, code, table, list, mixed)header_path β hierarchical path of section headersstart_line / end_line β source line numberscode_language β programming language (for code blocks)previous_content / next_content β overlap context from adjacent chunksadaptive_size β calculated optimal chunk size (when adaptive sizing enabled) newcontent_complexity β complexity score 0.0-1.0 (when adaptive sizing enabled) newsize_scale_factor β applied scaling factor (when adaptive sizing enabled) newcode_role β code block role (example, setup, output, before, after, error)has_related_code β whether chunk contains related code blockscode_relationship β relationship type (before_after, code_output, sequential)explanation_bound β whether explanation is bound to codeWhen to disable metadata:
With include_metadata: false, chunks contain only the raw Markdown content:
# Requirements
Python 3.12 or higher is required...
workflow:
- node: load_document
type: document_loader
- node: chunk_markdown
type: tool
tool: advanced_markdown_chunker
input: ${load_document.content}
config:
max_chunk_size: 2048
strategy: auto
chunk_overlap: 100
- node: embed_chunks
type: embedding
input: ${chunk_markdown.chunks}
- node: store_vectors
type: vector_store
input: ${embed_chunks.vectors}
- node: chunk_api_docs
type: tool
tool: advanced_markdown_chunker
config:
max_chunk_size: 1500
strategy: code
include_metadata: true
from chunkana import MarkdownChunker
# Simple chunking
chunker = MarkdownChunker()
chunks = chunker.chunk("# Hello\n\nWorld")
# With analysis
result = chunker.chunk("# Hello\n\nWorld", include_analysis=True)
print(f"Strategy: {result.strategy_used}")
print(f"Chunks: {len(result.chunks)}")
from chunkana import MarkdownChunker
# Create hierarchical structure with parent-child relationships
chunker = MarkdownChunker()
result = chunker.chunk_hierarchical(markdown_text)
# Access document root
root = result.get_chunk(result.root_id)
print(f"Document: {root.content[:100]}...")
# Navigate hierarchy
sections = result.get_children(result.root_id)
for section in sections:
print(f"Section: {section.metadata['header_path']}")
# Get subsections
subsections = result.get_children(section.metadata['chunk_id'])
for subsection in subsections:
print(f" - {subsection.metadata['header_path']}")
# Multi-level retrieval: Find chunk and get context
matched_chunk = sections[0] # Example: search result
parent_context = result.get_parent(matched_chunk.metadata['chunk_id'])
breadcrumb = [a.metadata['header_path'] for a in result.get_ancestors(matched_chunk.metadata['chunk_id'])]
print(f"Breadcrumb: {' > '.join(reversed(breadcrumb))}")
# Backward-compatible flat access
leaf_chunks = result.get_flat_chunks()
print(f"Total leaf chunks: {len(leaf_chunks)}")
from chunkana import MarkdownChunker
chunker = MarkdownChunker()
# Automatic selection (recommended)
chunks = chunker.chunk(text)
# Force specific strategy
chunks = chunker.chunk(text, strategy="code_aware")
chunks = chunker.chunk(text, strategy="list_aware") # For list-heavy docs
chunks = chunker.chunk(text, strategy="structural")
from chunkana import MarkdownChunker, ChunkConfig
# Changelog processing with list-aware strategy
changelog = """
# Changelog
## Version 2.0
New features:
- **Authentication**
- OAuth 2.0 support
- SAML integration
- MFA with SMS and authenticator apps
- **Performance**
- 50% faster processing
- Reduced memory usage
"""
config = ChunkConfig(
max_chunk_size=2000,
list_ratio_threshold=0.35, # Lower threshold for changelogs
list_count_threshold=3 # Activate with fewer lists
)
chunker = MarkdownChunker(config)
chunks = chunker.chunk(changelog)
# Result: Nested items stay together, context preserved
for chunk in chunks:
print(f"Chunk type: {chunk.metadata['content_type']}")
print(f"List depth: {chunk.metadata.get('max_list_depth', 0)}")
from chunkana import MarkdownChunker
# Meta-documentation with nested code blocks
meta_doc = '''
# How to Write Documentation
## Code Examples
When documenting code, use triple backticks. For showing markdown examples,
use quadruple backticks:
````markdown
Here's how to show Python code:
```python
def example():
return "Hello, World!"
```
````
This preserves the nested structure correctly.
'''
chunker = MarkdownChunker()
chunks = chunker.chunk(meta_doc)
# Result: Nested fences preserved as single code block
# Inner ```python``` stays inside outer ````markdown````
for chunk in chunks:
if '````' in chunk.content:
print("Nested fencing preserved!")
print(f"Has code blocks: {chunk.metadata.get('has_code', False)}")
from chunkana import MarkdownChunker, ChunkConfig, TableGroupingConfig
# API documentation with related tables
api_docs = """
## GET /users/{id}
### Parameters
| Parameter | Type | Required |
|-----------|------|----------|
| id | string | yes |
### Response Fields
| Field | Type | Description |
|-------|------|-------------|
| name | string | User name |
| email | string | Email |
### Error Codes
| Code | Message |
|------|---------|
| 404 | Not Found |
"""
# Enable table grouping
config = ChunkConfig(
group_related_tables=True,
table_grouping_config=TableGroupingConfig(
max_distance_lines=10,
require_same_section=True,
)
)
chunker = MarkdownChunker(config)
chunks = chunker.chunk(api_docs)
# Related tables grouped together for better retrieval
for chunk in chunks:
if chunk.metadata.get("is_table_group"):
print(f"Grouped {chunk.metadata['table_group_count']} tables")
from chunkana import MarkdownChunker, StreamingConfig
import os
# Process large files with minimal memory usage
chunker = MarkdownChunker()
# Configure streaming for memory-constrained environments
streaming_config = StreamingConfig(
buffer_size=100_000, # 100KB buffer windows
max_memory_mb=50 # Strict 50MB memory limit
)
# Stream process large file (e.g., 50MB documentation)
file_path = "large_documentation.md"
chunk_count = 0
for chunk in chunker.chunk_file_streaming(file_path, streaming_config):
# Process each chunk immediately (e.g., insert to vector DB)
chunk_count += 1
print(f"Processed chunk {chunk_count}: {len(chunk.content)} chars")
# Access streaming-specific metadata
window_idx = chunk.metadata.get('stream_window_index', 0)
print(f" From window: {window_idx}")
print(f"Total chunks processed: {chunk_count}")
print(f"Memory usage stayed below {streaming_config.max_memory_mb}MB")
# Progress tracking example
file_size = os.path.getsize(file_path)
processed_bytes = 0
for chunk in chunker.chunk_file_streaming(file_path):
processed_bytes += len(chunk.content)
progress = (processed_bytes / file_size) * 100
print(f"\rProgress: {progress:.1f}%", end="")
print("\nDone!")
from chunkana import MarkdownChunker, ChunkConfig
# For code-heavy documents (handles nested fencing)
config = ChunkConfig.for_code_heavy()
chunker = MarkdownChunker(config)
# For Dify RAG systems
config = ChunkConfig.for_dify_rag()
chunker = MarkdownChunker(config)
# For search indexing
config = ChunkConfig.for_search_indexing()
chunker = MarkdownChunker(config)
from chunkana import MarkdownChunker
chunker = MarkdownChunker()
result = chunker.chunk(markdown_text, include_analysis=True)
for chunk in result.chunks:
print(f"Content: {chunk.content[:50]}...")
print(f"Lines: {chunk.start_line}-{chunk.end_line}")
print(f"Size: {chunk.size} chars")
print(f"Type: {chunk.content_type}")
print(f"Strategy: {chunk.strategy}")
from chunkana import chunk_text, chunk_file
# Chunk text directly
chunks = chunk_text("# My Document\n\nContent here...")
# Chunk from file
chunks = chunk_file("README.md")
The system automatically selects the optimal strategy based on content analysis:
| Strategy | Priority | Activation Conditions | Best For |
|---|---|---|---|
| Code-Aware | 1 (highest) | code_ratio β₯ 30% OR has code blocks/tables | Technical docs, API docs |
| List-Aware | 2 | list_ratio > 40% OR list_count β₯ 5 (AND logic for structured docs) | Changelogs, feature lists, task lists, outlines |
| Structural | 3 | β₯3 headers with hierarchy | Documentation, guides |
| Fallback | 4 (default) | Always applicable | Simple text, mixed content |
Unique capability not found in competing solutions (LangChain, LlamaIndex, Unstructured, Chonkie):
Intelligent List Processing:
Activation Logic:
# For documents with strong hierarchical structure (many headers):
activate_if: list_ratio > 0.40 AND list_count >= 5
# For documents without strong structure:
activate_if: list_ratio > 0.40 OR list_count >= 5
Perfect for:
Example Input:
Our product includes:
- **Authentication**
- OAuth 2.0 support
- SAML integration
- MFA options
- SMS
- Authenticator app
- Hardware keys
- **Authorization**
- Role-based access
- Permission groups
Result: Two coherent chunks preserving full hierarchies:
Competitor Behavior: Would split nested items, losing context and relationships.
Unique capability for code-heavy documentation that intelligently binds code blocks to their explanations:
Pattern Recognition:
Enhanced Metadata: Each code chunk includes:
code_role β Classification (example, setup, output, before, after, error)has_related_code β Boolean flag for grouped blockscode_relationship β Relationship type (before_after, code_output, sequential)explanation_bound β Whether explanation context is availableExample: Before/After Refactoring
# Code Improvement
## Refactoring
Before:
```python
def old_way():
x = 1
y = 2
return x + y
After:
def new_way():
return 1 + 2
**Result:** Single chunk containing both code blocks with metadata:
```json
{
"code_relationship": "before_after",
"code_roles": ["before", "after"],
"has_related_code": true,
"related_code_count": 2
}
Example: Code + Output
Run this command:
```bash
echo "Hello, World!"
Output:
Hello, World!
**Result:** Grouped chunk with code-output relationship preserved.
**Configuration:**
```python
config = ChunkConfig(
enable_code_context_binding=True, # Enable feature
bind_output_blocks=True, # Auto-detect output
preserve_before_after_pairs=True, # Keep comparisons together
max_context_chars_before=500, # Explanation search limit
)
Perfect for:
from chunkana import ChunkConfig
config = ChunkConfig(
# Size limits
max_chunk_size=4096, # Maximum chunk size (chars)
min_chunk_size=512, # Minimum chunk size
# Overlap (adaptive sizing)
overlap_size=200, # Base overlap size (0 = disabled)
# Actual max = min(overlap_size, chunk_size * 0.35)
# Behavior
preserve_atomic_blocks=True, # Keep code blocks and tables intact
extract_preamble=True, # Extract content before first header
# Strategy selection thresholds
code_threshold=0.3, # Code ratio for CodeAwareStrategy
structure_threshold=3, # Min headers for StructuralStrategy
list_ratio_threshold=0.40, # List ratio for ListAwareStrategy
list_count_threshold=5, # Min list blocks for ListAwareStrategy
# Code-Context Binding (NEW)
enable_code_context_binding=True, # Enable enhanced code-context binding
max_context_chars_before=500, # Max chars for backward explanation search
max_context_chars_after=300, # Max chars for forward explanation search
related_block_max_gap=5, # Max line gap for related block detection
bind_output_blocks=True, # Auto-bind output blocks to code
preserve_before_after_pairs=True, # Keep Before/After pairs together
# Adaptive Chunk Sizing (NEW)
use_adaptive_sizing=False, # Enable adaptive chunk sizing
adaptive_config=None, # AdaptiveSizeConfig instance (see below)
# Override
strategy_override=None, # Force specific strategy (code_aware/list_aware/structural/fallback)
)
Group related tables in the same chunk for better retrieval quality:
from chunkana import ChunkConfig, TableGroupingConfig
# Enable table grouping
config = ChunkConfig(
group_related_tables=True,
table_grouping_config=TableGroupingConfig(
max_distance_lines=10, # Max lines between tables to group
max_grouped_tables=5, # Max tables per group
max_group_size=5000, # Max chars for grouped content
require_same_section=True # Only group within same header section
)
)
chunker = MarkdownChunker(config)
chunks = chunker.chunk(api_docs)
# Grouped table chunks have metadata:
# - is_table_group: True
# - table_group_count: number of tables in group
When to Use:
Enable automatic size optimization based on content complexity:
from chunkana import ChunkConfig, AdaptiveSizeConfig
# Enable with default settings
config = ChunkConfig(
use_adaptive_sizing=True,
adaptive_config=AdaptiveSizeConfig(
base_size=1500, # Base chunk size for medium complexity
min_scale=0.5, # Minimum scaling factor (0.5x = 750 chars)
max_scale=1.5, # Maximum scaling factor (1.5x = 2250 chars)
# Complexity weights (must sum to 1.0)
code_weight=0.4, # Weight for code ratio
table_weight=0.3, # Weight for table ratio
list_weight=0.2, # Weight for list ratio
sentence_length_weight=0.1, # Weight for average sentence length
)
)
chunker = MarkdownChunker(config)
chunks = chunker.chunk(text)
# Chunks now have adaptive sizing metadata:
# - adaptive_size: calculated optimal size
# - content_complexity: complexity score (0.0-1.0)
# - size_scale_factor: applied scale factor
Quick Enable with Profile:
# Use pre-configured adaptive sizing profile
config = ChunkConfig.with_adaptive_sizing()
chunker = MarkdownChunker(config)
How It Works:
optimal_size = base_size * (min_scale + complexity * scale_range)Behavior:
Example Results:
| Document Type | Code Ratio | Complexity | Scale Factor | Size (base=1500) |
|---|---|---|---|---|
| API docs with code | 60% | 0.68 | 1.4x | 2100 chars |
| Technical blog | 20% | 0.30 | 0.8x | 1200 chars |
| Plain text guide | 0% | 0.10 | 0.6x | 900 chars |
When to Use:
### Configuration Profiles
| Profile | Use Case | Max Size | Overlap |
|---------|----------|----------|---------|
| `default()` | General use | 4096 | 200 |
| `for_code_heavy()` | Code documentation | 8192 | 100 |
| `for_structured()` | Structured docs | 4096 | 200 |
| `minimal()` | Fine-grained | 1024 | 50 |
### Overlap Handling
Two modes for overlap handling:
- **Metadata mode** (`include_metadata=True`): Overlap stored in `previous_content`/`next_content` fields
- **Content mode** (`include_metadata=False`): Overlap merged into chunk content
---
## π API Reference
### MarkdownChunker
```python
class MarkdownChunker:
def __init__(
self,
config: Optional[ChunkConfig] = None,
enable_performance_monitoring: bool = False
)
def chunk(
self,
md_text: str,
strategy: Optional[str] = None,
include_analysis: bool = False,
return_format: Literal["objects", "dict"] = "objects",
include_metadata: bool = True
) -> Union[List[Chunk], ChunkingResult, dict]
def chunk_hierarchical(
self,
md_text: str
) -> HierarchicalChunkingResult
def get_available_strategies(self) -> List[str]
def add_strategy(self, strategy: BaseStrategy) -> None
def remove_strategy(self, strategy_name: str) -> None
@dataclass
class Chunk:
content: str # Chunk content
start_line: int # Start line (1-based)
end_line: int # End line
metadata: Dict[str, Any]
# Properties
size: int # Size in characters
line_count: int # Number of lines
content_type: str # Content type (code/text/list/table/mixed)
strategy: str # Strategy used
language: Optional[str] # Programming language (for code)
@dataclass
class ChunkingResult:
chunks: List[Chunk]
strategy_used: str
processing_time: float
fallback_used: bool
fallback_level: int
errors: List[str]
warnings: List[str]
# Statistics
total_chars: int
total_lines: int
content_type: str
complexity_score: float
@dataclass
class HierarchicalChunkingResult:
chunks: List[Chunk] # All chunks including root document chunk
root_id: str # ID of document-level chunk
strategy_used: str # Name of chunking strategy applied
# Navigation methods (O(1) performance)
def get_chunk(chunk_id: str) -> Optional[Chunk]
def get_children(chunk_id: str) -> List[Chunk]
def get_parent(chunk_id: str) -> Optional[Chunk]
def get_ancestors(chunk_id: str) -> List[Chunk] # Parent to root
def get_siblings(chunk_id: str) -> List[Chunk] # Includes self
def get_flat_chunks() -> List[Chunk] # Leaf chunks only
def get_by_level(level: int) -> List[Chunk] # 0=doc, 1=section, 2=subsection, 3=paragraph
def to_tree_dict() -> Dict # Serializable tree structure
Hierarchy Metadata Fields:
Each chunk in hierarchical mode includes these additional metadata fields:
| Field | Type | Description |
|---|---|---|
chunk_id | str | Unique 8-char hash identifier |
parent_id | str | Parent chunk ID (None for root) |
children_ids | List[str] | Child chunk IDs |
prev_sibling_id | str | Previous sibling ID |
next_sibling_id | str | Next sibling ID |
hierarchy_level | int | 0=document, 1=section, 2=subsection, 3=paragraph |
is_leaf | bool | Has no children |
is_root | bool | Document-level chunk |
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Dify Plugin Layer β
β (dify-markdown-chunker) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
ββββββββββββββββββββ ββββββββββββββββββββ βββββββββββββββββββ
β Plugin Tools β β Migration β β Input/Output β
β β β Adapter β β Filtering β
β β’ Tool Schema β β β β β
β β’ Parameter β β β’ API Mapping β β β’ Metadata β
β Validation β β β’ Compatibility β β Embedding β
β β’ Dify Interface β β β’ Error Handling β β β’ Debug Control β
ββββββββββββββββββββ ββββββββββββββββββββ βββββββββββββββββββ
β β β
βββββββββββββββββΌββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Chunkana Engine β
β (Core Chunking Library) β
β β
β βββββββββββββββ βββββββββββββββ βββββββββββββββ β
β β Parser β β Strategy β β Hierarchy β β
β β β β Selector β β Builder β β
β β β’ AST Build β β β β β β
β β β’ Content β β β’ Auto β β β’ Parent- β β
β β Analysis β β β’ Code β β Child β β
β β β’ Element β β β’ List β β β’ Navigationβ β
β β Detection β β β’ Struct β β β’ Metadata β β
β βββββββββββββββ β β’ Fallback β βββββββββββββββ β
β βββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The plugin now uses a layered architecture with clear separation of concerns:
Plugin Layer (dify-markdown-chunker):
Core Engine (chunkana):
| Module | Description |
|---|---|
tools/markdown_chunk_tool.py | Dify tool implementation with parameter mapping |
adapter.py | Migration adapter providing API compatibility |
input_validator.py | Input validation and preprocessing |
output_filter.py | Output filtering and metadata control |
| chunkana library | Core chunking engine (external dependency) |
dify-markdown-chunker/
βββ tools/ # Dify plugin tools
β βββ markdown_chunk_tool.py # Main tool implementation
β βββ markdown_chunk_tool.yaml # Tool schema and metadata
βββ provider/ # Dify plugin provider
β βββ markdown_chunker.yaml # Provider configuration
βββ adapter.py # Migration adapter (API compatibility)
βββ input_validator.py # Input validation and preprocessing
βββ output_filter.py # Output filtering and metadata control
βββ main.py # Plugin entry point
βββ tests/ # Test suite (migration-compatible)
β βββ test_migration_*.py # Migration compatibility tests
β βββ test_integration_*.py # Integration tests
β βββ test_*_adapted.py # Adapted legacy tests
βββ docs/ # Documentation
βββ manifest.yaml # Dify plugin manifest
βββ requirements.txt # Dependencies (includes chunkana)
The migration to chunkana provides several advantages:
The v2 architecture delivers excellent performance with linear scaling:
| Document Size | Processing Time | Throughput | Memory |
|---|---|---|---|
| Tiny (1KB) | 2.3ms | 435 KB/s | 12.3 MB |
| Small (10KB) | 8.5ms | 1,177 KB/s | 14.5 MB |
| Medium (100KB) | 45.2ms | 2,212 KB/s | 28.4 MB |
| Large (1MB) | 412.5ms | 2,424 KB/s | 156.2 MB |
Performance Characteristics:
Note: Performance data based on benchmarks from
docs/research/07_benchmark_results.mdconducted on Windows 11, Intel Core i7, 16GB RAM, SSD. Actual performance may vary depending on system configuration, document complexity, and content type.
chunker = MarkdownChunker(enable_performance_monitoring=True)
for doc in documents:
chunker.chunk(doc)
stats = chunker.get_performance_stats()
print(f"Average time: {stats['chunk']['avg_time']:.3f}s")
For detailed benchmarks and methodology, see Performance Guide.
This project uses pytest for testing. The test suite has been cleaned up and optimized:
# Run all tests
make test-all
# Run only migration-compatible tests
make test
# Run specific test categories
pytest tests/test_migration_*.py # Migration tests
pytest tests/test_integration_*.py # Integration tests
The test suite is organized as follows:
tests/test_migration_*.py - Tests using the migration adaptertests/test_integration_*.py - Integration teststests/test_*_adapted.py - Adapted legacy testsmake test-all
make test
pytest tests/test_migration_.py # Migration tests pytest tests/test_integration_.py # Integration tests
### Test Structure
The test suite is organized as follows:
- `tests/test_migration_*.py` - Tests using the migration adapter
- `tests/test_integration_*.py` - Integration tests
- `tests/test_*_adapted.py` - Adapted legacy tests
# Run all tests
make test-all
# Run only migration-compatible tests
make test
# Run specific test categories
pytest tests/test_migration_*.py # Migration tests
pytest tests/test_integration_*.py # Integration tests
The test suite is organized as follows:
tests/test_migration_*.py - Tests using the migration adaptertests/test_integration_*.py - Integration teststests/test_*_adapted.py - Adapted legacy testsmake test-all
make test
pytest tests/test_migration_.py # Migration tests pytest tests/test_integration_.py # Integration tests
### Test Structure
The test suite is organized as follows:
- `tests/test_migration_*.py` - Tests using the migration adapter
- `tests/test_integration_*.py` - Integration tests
- `tests/test_*_adapted.py` - Adapted legacy tests
# Run all tests (473 total: 97 plugin + 376 chunkana)
make test
# Verbose output
make test-verbose
# With coverage report
make test-coverage
# Quick tests
make test-quick
# Performance benchmarks
python tests/performance/run_benchmarks_standalone.py
# Format code
make format
# Run linter
make lint
# Type checking
make quality-check
# Validate structure
make validate
# Build package
make package
# Full release
make release
markdown-it-py>=3.0.0 β Markdown parsingmistune>=3.0.0 β Alternative parserpydantic>=2.0.0 β Data validationdify_plugin==0.5.0b15 β Dify integrationpytest>=8.0.0 β Testinghypothesis>=6.0.0 β Property-based testingblack>=23.0.0 β Code formattingmypy>=1.5.0 β Type checkingContributions are welcome! See CONTRIBUTING.md for guidelines.
# 1. Fork the repository
# 2. Create feature branch
git checkout -b feature/amazing-feature
# 3. Make changes with tests
# 4. Check quality
make test && make quality-check
# 5. Submit Pull Request
Author: Aleksandr Sukhodko (@asukhodko)
Repository: https://github.com/asukhodko/dify-markdown-chunker
MIT License β see LICENSE
Current Version: 2.1.6 (January 2026)
Released: January 4, 2026
Changes:
markdown_chunker and markdown_chunker_v2 directoriesadapter.py) for full API compatibilitydify-plugin CLI installation in MakefileReleased: December 23, 2025
Changes:
dify_plugin dependency from 0.5.0b15 to 0.7.0.difyignore to properly include README.md and PRIVACY.md in packageReleased: December 14, 2025
New Features:
max_distance_lines)require_same_section)max_group_size, max_grouped_tables)For full release history, see CHANGELOG.md.