Advanced Markdown Chunker Β· emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginsAdvanced Markdown Chunker
Plugin
Limited listing

Advanced Markdown Chunker

by asukhodko Β· v2.1.7

Advanced Markdown chunking powered by chunkana library with structural awareness for better RAG performance

4.0k installsUpdated Jan 12, 2026
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Tools

Available inside your emploidai workspace after installation.

Data sources

Available inside your emploidai workspace after installation.

Category

tool

Version

2.1.7asukhodko

Requirements

Maximum memory 512MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.0.0+

Permissions

No permissions were declared by the publisher.

Dependencies

No additional dependencies

Resources

Privacy policy
emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

πŸ”– Advanced Markdown Chunker for Dify

Intelligent Markdown document chunking for RAG systems with structural awareness

GitHub Repository Version License: MIT Python 3.12+ Dify Plugin Tests


πŸ“‹ Table of Contents

  • Overview
  • Features
  • Data & Privacy
  • Installation
  • Dify Integration
  • Quick Start
  • Chunking Strategies
  • Configuration
  • API Reference
  • Architecture
  • Performance
  • Development
  • Contributing
  • Author & Support
  • License

Overview

Advanced Markdown Chunker is a Dify plugin that intelligently splits Markdown documents into semantically meaningful chunks optimized for RAG (Retrieval-Augmented Generation) systems. Powered by the chunkana engine, it provides advanced structural awareness that goes beyond simple text splitting, preserving document structure, keeping code blocks intact, and automatically selecting the best chunking strategy based on content analysis.

Primary Use Case: RAG Systems

This plugin is designed primarily for RAG (Retrieval-Augmented Generation) workflows where chunks are embedded and stored in vector databases for semantic search. Built on the robust chunkana library, it provides enterprise-grade chunking capabilities through a user-friendly Dify interface. By default, each chunk includes embedded metadata (header paths, content type, line numbers) directly in the chunk text, which improves retrieval quality by providing additional context for embeddings.

Note for Model Training: If you need clean text without metadata (e.g., for fine-tuning language models), set include_metadata: false or post-process chunks to remove the <metadata> block.

Why Use This Plugin?

Simple Chunking ProblemsAdvanced Markdown Chunker Solution
Breaks code blocks mid-functionPreserves code blocks as atomic units
Loses header contextMaintains hierarchical section structure
Splits tables and listsKeeps tables and lists intact
One-size-fits-all approach4 adaptive strategies based on content
No overlap supportSmart overlap for better retrieval
Destroys list hierarchiesSmart list grouping with context binding
Breaks nested code examplesHandles nested fencing (````, ``````, ~~~~)
Code examples lose explanatory contextEnhanced code-context binding with pattern recognition
Before/After comparisons split apartIntelligent Before/After pairing
Code and output separatedAutomatic Code+Output binding
Mathematical formulas splitLaTeX formula preservation ($$...$$, environments)

✨ Features

🎯 Adaptive Chunking

  • 4 intelligent strategies β€” automatic selection based on content analysis
  • Adaptive Chunk Sizing β€” automatic size optimization based on content complexity (new)
    • Code-heavy content β†’ larger chunks (up to 1.5x base size)
    • Simple text β†’ smaller chunks (down to 0.5x base size)
    • Configurable complexity weights and scaling bounds
    • Optional feature (disabled by default for backward compatibility)
  • Hierarchical Chunking β€” parent-child relationships between chunks (new)
    • Multi-level retrieval support (overview vs. detail)
    • Programmatic navigation (siblings, ancestors, children)
    • O(1) chunk lookup performance
    • Backward compatible with flat chunking
  • Streaming Processing β€” memory-efficient processing for large files (new)
    • Process files >10MB with <50MB RAM usage
    • Configurable buffer management (100KB default window)
    • Progress tracking support for long-running operations
    • Maintains quality through smart window boundary detection
  • List-Aware Strategy β€” preserves nested list hierarchies and context (unique competitive advantage)
  • Nested Fencing Support β€” correctly handles quadruple/quintuple backticks and tilde fencing for meta-documentation (unique capability)
  • Enhanced Code-Context Binding β€” intelligently binds code blocks to explanations, recognizes Before/After patterns, Code+Output pairs, and sequential examples (unique competitive advantage)
  • LaTeX Formula Handling β€” preserves mathematical formulas as atomic blocks (new)
    • Display math ($$...$$) never split across chunks
    • Environment blocks (\begin{equation}, \begin{align}) preserved complete
    • Supported in all 4 chunking strategies
    • Critical for scientific papers and technical documentation
  • Table Grouping Option β€” groups related tables in same chunk for better retrieval (new)
    • Configurable proximity threshold (max_distance_lines)
    • Section boundary awareness (require_same_section)
    • Size and count limits (max_group_size, max_grouped_tables)
    • Perfect for API documentation with Parameters/Response/Error tables
  • Structure preservation β€” headers, lists, tables, and code stay intact
  • Adaptive overlap β€” context window scales with chunk size (up to 35%)

πŸ” Deep Content Analysis

  • AST parsing β€” full Markdown syntax analysis
  • Content type detection β€” code-heavy, text-heavy, mixed
  • Complexity scoring β€” optimizes strategy selection

πŸ›‘οΈ Reliability

  • 473 tests β€” comprehensive test coverage with property-based testing (97 plugin tests + 376 chunkana library tests)
  • Property-Based Testing β€” formal correctness guarantees with Hypothesis
  • Automatic fallback β€” graceful degradation on errors
  • Performance benchmarks β€” automated performance regression detection

πŸ”Œ Integration

  • Dify Plugin β€” ready-to-use in Dify workflows
  • Python Library β€” standalone usage
  • REST API Ready β€” adapters for API integration

πŸ”’ Data & Privacy

Local Processing Only
The Plugin processes all Markdown content locally within your Dify instance. No data is transmitted to external services.

What the Plugin does:

  • βœ… Parses Markdown structure using local AST analysis
  • βœ… Generates chunks based on document structure
  • βœ… Adds metadata for improved retrieval quality

What the Plugin does NOT do:

  • ❌ Send data to external APIs
  • ❌ Store data outside of Dify's standard mechanisms
  • ❌ Log or track user content
  • ❌ Collect analytics or telemetry

For complete details, see PRIVACY.md.


πŸ“¦ Installation

Dify Plugin Installation

  1. Download the .difypkg file from Releases
  2. In Dify: Settings β†’ Plugins β†’ Install Plugin
  3. Upload the .difypkg file
  4. The plugin is now available in your workflows

Requirements:

  • Dify version 1.9.0 or higher
  • No additional configuration needed

Development Installation

# Clone the repository
git clone https://github.com/asukhodko/dify-markdown-chunker.git
cd dify-markdown-chunker

# Create virtual environment
python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

# Verify installation
make test

Requirements:

  • Python 3.12 or higher

πŸ”Œ Dify Integration

Workflow Configuration

Add the chunker to your Dify workflow:

- node: chunk_markdown
  type: tool
  tool: advanced_markdown_chunker
  config:
    max_chunk_size: 2048
    strategy: auto
    chunk_overlap: 100
    include_metadata: true

Plugin Parameters

ParameterTypeDefaultDescription
input_textstringrequiredMarkdown text to chunk
max_chunk_sizenumber4096Maximum chunk size in characters
chunk_overlapnumber200Base overlap size (adaptive: actual max = min(overlap_size, chunk_size * 0.35))
strategyselectautoChunking strategy (auto/code_aware/list_aware/structural/fallback)
include_metadatabooleantrueEmbed metadata in chunk text (see below)
enable_hierarchybooleanfalseCreate parent-child relationships between chunks
debugbooleanfalseInclude all chunks (root, intermediate, leaf) in hierarchical mode
leaf_onlybooleanfalseReturn only leaf chunks in hierarchical mode (recommended for vector DB)

Parameter Mapping: Plugin β†’ Chunkana

The plugin parameters map to chunkana configuration as follows:

Plugin ParameterChunkana ConfigNotes
max_chunk_sizemax_chunk_sizeDirect mapping
chunk_overlapoverlap_sizeCapped at 35% of chunk size
strategystrategy_overrideauto β†’ automatic selection based on content analysis
include_metadatainclude_metadataControls metadata embedding in chunk text
enable_hierarchyenable_hierarchyEnables parent-child chunk relationships
debugdebug_modeControls visibility of all chunk types in hierarchical mode
leaf_onlyleaf_onlyFilters to content chunks only (excludes structural headers)

Advanced chunkana features not exposed in plugin UI:

  • Adaptive chunk sizing based on content complexity
  • Custom strategy thresholds and weights
  • Streaming processing for large files
  • Fine-grained code-context binding controls
  • Table grouping configuration

For direct chunkana usage with full feature access, see the chunkana documentation.

Hierarchical Chunking Mode

When enable_hierarchy=true, the plugin returns chunks organized in a tree structure with parent-child relationships.

Chunk Types:

Typeis_rootis_leafindexableDescription
RoottruefalsefalseDocument root, covers entire document
InternalfalsefalsetrueSection headers with children
LeaffalsetruetrueContent chunks for indexing

Filtering Behavior:

  • debug=false (default): Root chunk excluded from results
  • debug=true: All chunks included for debugging
  • leaf_only=true: Only leaf chunks returned (recommended for vector DB)

Recommended Usage for Vector DB:

- node: chunk_for_indexing
  type: tool
  tool: advanced_markdown_chunker
  config:
    enable_hierarchy: true
    leaf_only: true  # Only indexable content chunks

For Debugging Hierarchy:

- node: debug_hierarchy
  type: tool
  tool: advanced_markdown_chunker
  config:
    enable_hierarchy: true
    debug: true  # Include root and internal nodes

Understanding chunk_overlap

Chunk Overlap controls how many characters of context are shared between consecutive chunks to preserve semantic continuity.

Behavior depends on include_metadata:

include_metadataOverlap Behavior
true (default)Overlap stored in metadata fields previous_content / next_content. Chunk content stays clean.
falseOverlap embedded directly into chunk text: previous_content + "\n" + main + "\n" + next_content

Example with include_metadata: true:

<metadata>
{
  "previous_content": "...end of previous chunk...",
  "next_content": "...start of next chunk..."
}
</metadata>
# Current Section

Main content of this chunk...

Example with include_metadata: false:

...end of previous chunk...
# Current Section

Main content of this chunk...
...start of next chunk...

This allows chunk_overlap to work predictably in both modes:

  • RAG mode (include_metadata: true): Overlap available as structured metadata for embeddings
  • Clean text mode (include_metadata: false): Overlap physically present in text for sliding window processing

Understanding include_metadata

When include_metadata: true (default), each chunk includes a <metadata> block prepended to the content:

<metadata>
{
  "content_type": "text",
  "header_path": "/Installation/Requirements",
  "start_line": 45,
  "end_line": 52
}
</metadata>
# Requirements

Python 3.12 or higher is required...

Typical metadata fields:

  • content_type β€” type of content (text, code, table, list, mixed)
  • header_path β€” hierarchical path of section headers
  • start_line / end_line β€” source line numbers
  • code_language β€” programming language (for code blocks)
  • previous_content / next_content β€” overlap context from adjacent chunks
  • adaptive_size β€” calculated optimal chunk size (when adaptive sizing enabled) new
  • content_complexity β€” complexity score 0.0-1.0 (when adaptive sizing enabled) new
  • size_scale_factor β€” applied scaling factor (when adaptive sizing enabled) new
  • code_role β€” code block role (example, setup, output, before, after, error)
  • has_related_code β€” whether chunk contains related code blocks
  • code_relationship β€” relationship type (before_after, code_output, sequential)
  • explanation_bound β€” whether explanation is bound to code

When to disable metadata:

  • Fine-tuning language models (need clean training data)
  • Exporting chunks for external processing
  • When metadata would interfere with downstream tasks

With include_metadata: false, chunks contain only the raw Markdown content:

# Requirements

Python 3.12 or higher is required...

Example: Knowledge Base Ingestion

workflow:
  - node: load_document
    type: document_loader
  
  - node: chunk_markdown
    type: tool
    tool: advanced_markdown_chunker
    input: ${load_document.content}
    config:
      max_chunk_size: 2048
      strategy: auto
      chunk_overlap: 100
  
  - node: embed_chunks
    type: embedding
    input: ${chunk_markdown.chunks}
  
  - node: store_vectors
    type: vector_store
    input: ${embed_chunks.vectors}

Example: API Documentation Processing

- node: chunk_api_docs
  type: tool
  tool: advanced_markdown_chunker
  config:
    max_chunk_size: 1500
    strategy: code
    include_metadata: true

πŸš€ Quick Start

Basic Usage

from chunkana import MarkdownChunker

# Simple chunking
chunker = MarkdownChunker()
chunks = chunker.chunk("# Hello\n\nWorld")

# With analysis
result = chunker.chunk("# Hello\n\nWorld", include_analysis=True)
print(f"Strategy: {result.strategy_used}")
print(f"Chunks: {len(result.chunks)}")

Hierarchical Chunking

from chunkana import MarkdownChunker

# Create hierarchical structure with parent-child relationships
chunker = MarkdownChunker()
result = chunker.chunk_hierarchical(markdown_text)

# Access document root
root = result.get_chunk(result.root_id)
print(f"Document: {root.content[:100]}...")

# Navigate hierarchy
sections = result.get_children(result.root_id)
for section in sections:
    print(f"Section: {section.metadata['header_path']}")
    
    # Get subsections
    subsections = result.get_children(section.metadata['chunk_id'])
    for subsection in subsections:
        print(f"  - {subsection.metadata['header_path']}")

# Multi-level retrieval: Find chunk and get context
matched_chunk = sections[0]  # Example: search result
parent_context = result.get_parent(matched_chunk.metadata['chunk_id'])
breadcrumb = [a.metadata['header_path'] for a in result.get_ancestors(matched_chunk.metadata['chunk_id'])]
print(f"Breadcrumb: {' > '.join(reversed(breadcrumb))}")

# Backward-compatible flat access
leaf_chunks = result.get_flat_chunks()
print(f"Total leaf chunks: {len(leaf_chunks)}")

Strategy Selection

from chunkana import MarkdownChunker

chunker = MarkdownChunker()

# Automatic selection (recommended)
chunks = chunker.chunk(text)

# Force specific strategy
chunks = chunker.chunk(text, strategy="code_aware")
chunks = chunker.chunk(text, strategy="list_aware")  # For list-heavy docs
chunks = chunker.chunk(text, strategy="structural")

List-Aware Strategy Example

from chunkana import MarkdownChunker, ChunkConfig

# Changelog processing with list-aware strategy
changelog = """
# Changelog

## Version 2.0

New features:
- **Authentication**
  - OAuth 2.0 support
  - SAML integration
  - MFA with SMS and authenticator apps
- **Performance**
  - 50% faster processing
  - Reduced memory usage
"""

config = ChunkConfig(
    max_chunk_size=2000,
    list_ratio_threshold=0.35,  # Lower threshold for changelogs
    list_count_threshold=3       # Activate with fewer lists
)

chunker = MarkdownChunker(config)
chunks = chunker.chunk(changelog)

# Result: Nested items stay together, context preserved
for chunk in chunks:
    print(f"Chunk type: {chunk.metadata['content_type']}")
    print(f"List depth: {chunk.metadata.get('max_list_depth', 0)}")

Nested Fencing Support Example

from chunkana import MarkdownChunker

# Meta-documentation with nested code blocks
meta_doc = '''
# How to Write Documentation

## Code Examples

When documenting code, use triple backticks. For showing markdown examples,
use quadruple backticks:

````markdown
Here's how to show Python code:

```python
def example():
    return "Hello, World!"
```
````

This preserves the nested structure correctly.
'''

chunker = MarkdownChunker()
chunks = chunker.chunk(meta_doc)

# Result: Nested fences preserved as single code block
# Inner ```python``` stays inside outer ````markdown````
for chunk in chunks:
    if '````' in chunk.content:
        print("Nested fencing preserved!")
        print(f"Has code blocks: {chunk.metadata.get('has_code', False)}")

Table Grouping Example

from chunkana import MarkdownChunker, ChunkConfig, TableGroupingConfig

# API documentation with related tables
api_docs = """
## GET /users/{id}

### Parameters

| Parameter | Type | Required |
|-----------|------|----------|
| id | string | yes |

### Response Fields

| Field | Type | Description |
|-------|------|-------------|
| name | string | User name |
| email | string | Email |

### Error Codes

| Code | Message |
|------|---------|
| 404 | Not Found |
"""

# Enable table grouping
config = ChunkConfig(
    group_related_tables=True,
    table_grouping_config=TableGroupingConfig(
        max_distance_lines=10,
        require_same_section=True,
    )
)

chunker = MarkdownChunker(config)
chunks = chunker.chunk(api_docs)

# Related tables grouped together for better retrieval
for chunk in chunks:
    if chunk.metadata.get("is_table_group"):
        print(f"Grouped {chunk.metadata['table_group_count']} tables")

Streaming Processing Example

from chunkana import MarkdownChunker, StreamingConfig
import os

# Process large files with minimal memory usage
chunker = MarkdownChunker()

# Configure streaming for memory-constrained environments
streaming_config = StreamingConfig(
    buffer_size=100_000,  # 100KB buffer windows
    max_memory_mb=50      # Strict 50MB memory limit
)

# Stream process large file (e.g., 50MB documentation)
file_path = "large_documentation.md"
chunk_count = 0

for chunk in chunker.chunk_file_streaming(file_path, streaming_config):
    # Process each chunk immediately (e.g., insert to vector DB)
    chunk_count += 1
    print(f"Processed chunk {chunk_count}: {len(chunk.content)} chars")
    
    # Access streaming-specific metadata
    window_idx = chunk.metadata.get('stream_window_index', 0)
    print(f"  From window: {window_idx}")

print(f"Total chunks processed: {chunk_count}")
print(f"Memory usage stayed below {streaming_config.max_memory_mb}MB")

# Progress tracking example
file_size = os.path.getsize(file_path)
processed_bytes = 0

for chunk in chunker.chunk_file_streaming(file_path):
    processed_bytes += len(chunk.content)
    progress = (processed_bytes / file_size) * 100
    print(f"\rProgress: {progress:.1f}%", end="")

print("\nDone!")

Configuration Profiles

from chunkana import MarkdownChunker, ChunkConfig

# For code-heavy documents (handles nested fencing)
config = ChunkConfig.for_code_heavy()
chunker = MarkdownChunker(config)

# For Dify RAG systems
config = ChunkConfig.for_dify_rag()
chunker = MarkdownChunker(config)

# For search indexing
config = ChunkConfig.for_search_indexing()
chunker = MarkdownChunker(config)

Accessing Chunk Metadata

from chunkana import MarkdownChunker

chunker = MarkdownChunker()
result = chunker.chunk(markdown_text, include_analysis=True)

for chunk in result.chunks:
    print(f"Content: {chunk.content[:50]}...")
    print(f"Lines: {chunk.start_line}-{chunk.end_line}")
    print(f"Size: {chunk.size} chars")
    print(f"Type: {chunk.content_type}")
    print(f"Strategy: {chunk.strategy}")

Convenience Functions

from chunkana import chunk_text, chunk_file

# Chunk text directly
chunks = chunk_text("# My Document\n\nContent here...")

# Chunk from file
chunks = chunk_file("README.md")

🎨 Chunking Strategies

The system automatically selects the optimal strategy based on content analysis:

StrategyPriorityActivation ConditionsBest For
Code-Aware1 (highest)code_ratio β‰₯ 30% OR has code blocks/tablesTechnical docs, API docs
List-Aware2list_ratio > 40% OR list_count β‰₯ 5 (AND logic for structured docs)Changelogs, feature lists, task lists, outlines
Structural3β‰₯3 headers with hierarchyDocumentation, guides
Fallback4 (default)Always applicableSimple text, mixed content

List-Aware Strategy: Competitive Advantage

Unique capability not found in competing solutions (LangChain, LlamaIndex, Unstructured, Chonkie):

Intelligent List Processing:

  • Hierarchy Preservation β€” nested lists never split across depth levels
  • Context Binding β€” introduction paragraphs automatically attached to their lists
  • Smart Grouping β€” related list items kept together based on structure
  • Type Detection β€” handles bullet lists, numbered lists, and checkboxes

Activation Logic:

# For documents with strong hierarchical structure (many headers):
activate_if: list_ratio > 0.40 AND list_count >= 5

# For documents without strong structure:
activate_if: list_ratio > 0.40 OR list_count >= 5

Perfect for:

  • Changelogs β€” version releases with nested changes
  • Feature lists β€” product capabilities with descriptions
  • Task lists β€” todos with sub-tasks and checkboxes
  • Outlines β€” structured notes and cheatsheets

Example Input:

Our product includes:

- **Authentication**
  - OAuth 2.0 support
  - SAML integration
  - MFA options
    - SMS
    - Authenticator app
    - Hardware keys

- **Authorization**
  - Role-based access
  - Permission groups

Result: Two coherent chunks preserving full hierarchies:

  • Chunk 1: Introduction + Authentication with all nested items
  • Chunk 2: Authorization with all nested items

Competitor Behavior: Would split nested items, losing context and relationships.

Code-Context Binding: Competitive Advantage

Unique capability for code-heavy documentation that intelligently binds code blocks to their explanations:

Pattern Recognition:

  • Before/After Comparisons β€” keeps refactoring examples together
  • Code + Output Pairs β€” binds execution results to code
  • Setup + Example β€” groups installation with usage
  • Sequential Steps β€” maintains tutorial order

Enhanced Metadata: Each code chunk includes:

  • code_role β€” Classification (example, setup, output, before, after, error)
  • has_related_code β€” Boolean flag for grouped blocks
  • code_relationship β€” Relationship type (before_after, code_output, sequential)
  • explanation_bound β€” Whether explanation context is available

Example: Before/After Refactoring

# Code Improvement

## Refactoring

Before:

```python
def old_way():
    x = 1
    y = 2
    return x + y

After:

def new_way():
    return 1 + 2

**Result:** Single chunk containing both code blocks with metadata:
```json
{
  "code_relationship": "before_after",
  "code_roles": ["before", "after"],
  "has_related_code": true,
  "related_code_count": 2
}

Example: Code + Output

Run this command:

```bash
echo "Hello, World!"

Output:

Hello, World!

**Result:** Grouped chunk with code-output relationship preserved.

**Configuration:**
```python
config = ChunkConfig(
    enable_code_context_binding=True,    # Enable feature
    bind_output_blocks=True,              # Auto-detect output
    preserve_before_after_pairs=True,     # Keep comparisons together
    max_context_chars_before=500,         # Explanation search limit
)

Perfect for:

  • API documentation with examples
  • Tutorial-style technical writing
  • Code migration guides
  • Troubleshooting documentation

βš™οΈ Configuration

Basic Parameters

from chunkana import ChunkConfig

config = ChunkConfig(
    # Size limits
    max_chunk_size=4096,      # Maximum chunk size (chars)
    min_chunk_size=512,       # Minimum chunk size
    
    # Overlap (adaptive sizing)
    overlap_size=200,         # Base overlap size (0 = disabled)
                              # Actual max = min(overlap_size, chunk_size * 0.35)
    
    # Behavior
    preserve_atomic_blocks=True,  # Keep code blocks and tables intact
    extract_preamble=True,        # Extract content before first header
    
    # Strategy selection thresholds
    code_threshold=0.3,           # Code ratio for CodeAwareStrategy
    structure_threshold=3,        # Min headers for StructuralStrategy
    list_ratio_threshold=0.40,    # List ratio for ListAwareStrategy
    list_count_threshold=5,       # Min list blocks for ListAwareStrategy
    
    # Code-Context Binding (NEW)
    enable_code_context_binding=True,   # Enable enhanced code-context binding
    max_context_chars_before=500,       # Max chars for backward explanation search
    max_context_chars_after=300,        # Max chars for forward explanation search
    related_block_max_gap=5,            # Max line gap for related block detection
    bind_output_blocks=True,            # Auto-bind output blocks to code
    preserve_before_after_pairs=True,   # Keep Before/After pairs together
    
    # Adaptive Chunk Sizing (NEW)
    use_adaptive_sizing=False,          # Enable adaptive chunk sizing
    adaptive_config=None,               # AdaptiveSizeConfig instance (see below)
    
    # Override
    strategy_override=None,   # Force specific strategy (code_aware/list_aware/structural/fallback)
)

Table Grouping Configuration

Group related tables in the same chunk for better retrieval quality:

from chunkana import ChunkConfig, TableGroupingConfig

# Enable table grouping
config = ChunkConfig(
    group_related_tables=True,
    table_grouping_config=TableGroupingConfig(
        max_distance_lines=10,    # Max lines between tables to group
        max_grouped_tables=5,     # Max tables per group
        max_group_size=5000,      # Max chars for grouped content
        require_same_section=True # Only group within same header section
    )
)

chunker = MarkdownChunker(config)
chunks = chunker.chunk(api_docs)

# Grouped table chunks have metadata:
# - is_table_group: True
# - table_group_count: number of tables in group

When to Use:

  • βœ… API documentation with Parameters/Response/Error tables
  • βœ… Data reports with related comparison tables
  • βœ… Technical specs with multiple related tables
  • ❌ Documents where tables are independent

Adaptive Chunk Sizing Configuration

Enable automatic size optimization based on content complexity:

from chunkana import ChunkConfig, AdaptiveSizeConfig

# Enable with default settings
config = ChunkConfig(
    use_adaptive_sizing=True,
    adaptive_config=AdaptiveSizeConfig(
        base_size=1500,           # Base chunk size for medium complexity
        min_scale=0.5,            # Minimum scaling factor (0.5x = 750 chars)
        max_scale=1.5,            # Maximum scaling factor (1.5x = 2250 chars)
        
        # Complexity weights (must sum to 1.0)
        code_weight=0.4,          # Weight for code ratio
        table_weight=0.3,         # Weight for table ratio
        list_weight=0.2,          # Weight for list ratio
        sentence_length_weight=0.1,  # Weight for average sentence length
    )
)

chunker = MarkdownChunker(config)
chunks = chunker.chunk(text)

# Chunks now have adaptive sizing metadata:
# - adaptive_size: calculated optimal size
# - content_complexity: complexity score (0.0-1.0)
# - size_scale_factor: applied scale factor

Quick Enable with Profile:

# Use pre-configured adaptive sizing profile
config = ChunkConfig.with_adaptive_sizing()
chunker = MarkdownChunker(config)

How It Works:

  1. Content Analysis - Calculates code ratio, table ratio, list ratio, avg sentence length
  2. Complexity Scoring - Weighted sum of factors produces score 0.0-1.0
  3. Size Calculation - optimal_size = base_size * (min_scale + complexity * scale_range)
  4. Chunk Application - Chunks respect calculated size while preserving atomic blocks

Behavior:

  • Code-heavy documents (high complexity) β†’ larger chunks (up to 1.5x base size)
  • Simple text (low complexity) β†’ smaller chunks (down to 0.5x base size)
  • Mixed content β†’ balanced sizing

Example Results:

Document TypeCode RatioComplexityScale FactorSize (base=1500)
API docs with code60%0.681.4x2100 chars
Technical blog20%0.300.8x1200 chars
Plain text guide0%0.100.6x900 chars

When to Use:

  • βœ… Mixed corpus with varying complexity
  • βœ… Want optimal retrieval precision across content types
  • βœ… Need larger chunks for code preservation
  • ❌ Require predictable chunk sizes (disable for consistency)

### Configuration Profiles

| Profile | Use Case | Max Size | Overlap |
|---------|----------|----------|---------|
| `default()` | General use | 4096 | 200 |
| `for_code_heavy()` | Code documentation | 8192 | 100 |
| `for_structured()` | Structured docs | 4096 | 200 |
| `minimal()` | Fine-grained | 1024 | 50 |

### Overlap Handling

Two modes for overlap handling:

- **Metadata mode** (`include_metadata=True`): Overlap stored in `previous_content`/`next_content` fields
- **Content mode** (`include_metadata=False`): Overlap merged into chunk content

---

## πŸ“š API Reference

### MarkdownChunker

```python
class MarkdownChunker:
    def __init__(
        self,
        config: Optional[ChunkConfig] = None,
        enable_performance_monitoring: bool = False
    )
    
    def chunk(
        self,
        md_text: str,
        strategy: Optional[str] = None,
        include_analysis: bool = False,
        return_format: Literal["objects", "dict"] = "objects",
        include_metadata: bool = True
    ) -> Union[List[Chunk], ChunkingResult, dict]
    
    def chunk_hierarchical(
        self,
        md_text: str
    ) -> HierarchicalChunkingResult
    
    def get_available_strategies(self) -> List[str]
    def add_strategy(self, strategy: BaseStrategy) -> None
    def remove_strategy(self, strategy_name: str) -> None

Chunk

@dataclass
class Chunk:
    content: str           # Chunk content
    start_line: int        # Start line (1-based)
    end_line: int          # End line
    metadata: Dict[str, Any]
    
    # Properties
    size: int              # Size in characters
    line_count: int        # Number of lines
    content_type: str      # Content type (code/text/list/table/mixed)
    strategy: str          # Strategy used
    language: Optional[str] # Programming language (for code)

ChunkingResult

@dataclass
class ChunkingResult:
    chunks: List[Chunk]
    strategy_used: str
    processing_time: float
    fallback_used: bool
    fallback_level: int
    errors: List[str]
    warnings: List[str]
    
    # Statistics
    total_chars: int
    total_lines: int
    content_type: str
    complexity_score: float

HierarchicalChunkingResult

@dataclass
class HierarchicalChunkingResult:
    chunks: List[Chunk]         # All chunks including root document chunk
    root_id: str                # ID of document-level chunk
    strategy_used: str          # Name of chunking strategy applied
    
    # Navigation methods (O(1) performance)
    def get_chunk(chunk_id: str) -> Optional[Chunk]
    def get_children(chunk_id: str) -> List[Chunk]
    def get_parent(chunk_id: str) -> Optional[Chunk]
    def get_ancestors(chunk_id: str) -> List[Chunk]  # Parent to root
    def get_siblings(chunk_id: str) -> List[Chunk]   # Includes self
    def get_flat_chunks() -> List[Chunk]             # Leaf chunks only
    def get_by_level(level: int) -> List[Chunk]      # 0=doc, 1=section, 2=subsection, 3=paragraph
    def to_tree_dict() -> Dict                       # Serializable tree structure

Hierarchy Metadata Fields:

Each chunk in hierarchical mode includes these additional metadata fields:

FieldTypeDescription
chunk_idstrUnique 8-char hash identifier
parent_idstrParent chunk ID (None for root)
children_idsList[str]Child chunk IDs
prev_sibling_idstrPrevious sibling ID
next_sibling_idstrNext sibling ID
hierarchy_levelint0=document, 1=section, 2=subsection, 3=paragraph
is_leafboolHas no children
is_rootboolDocument-level chunk

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Dify Plugin Layer                        β”‚
β”‚                 (dify-markdown-chunker)                     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό               β–Ό               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Plugin Tools   β”‚ β”‚  Migration       β”‚ β”‚ Input/Output    β”‚
β”‚                  β”‚ β”‚  Adapter         β”‚ β”‚ Filtering       β”‚
β”‚ β€’ Tool Schema    β”‚ β”‚                  β”‚ β”‚                 β”‚
β”‚ β€’ Parameter      β”‚ β”‚ β€’ API Mapping    β”‚ β”‚ β€’ Metadata      β”‚
β”‚   Validation     β”‚ β”‚ β€’ Compatibility  β”‚ β”‚   Embedding     β”‚
β”‚ β€’ Dify Interface β”‚ β”‚ β€’ Error Handling β”‚ β”‚ β€’ Debug Control β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚               β”‚               β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Chunkana Engine                          β”‚
β”‚                 (Core Chunking Library)                     β”‚
β”‚                                                             β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”            β”‚
β”‚ β”‚   Parser    β”‚ β”‚  Strategy   β”‚ β”‚  Hierarchy  β”‚            β”‚
β”‚ β”‚             β”‚ β”‚  Selector   β”‚ β”‚  Builder    β”‚            β”‚
β”‚ β”‚ β€’ AST Build β”‚ β”‚             β”‚ β”‚             β”‚            β”‚
β”‚ β”‚ β€’ Content   β”‚ β”‚ β€’ Auto      β”‚ β”‚ β€’ Parent-   β”‚            β”‚
β”‚ β”‚   Analysis  β”‚ β”‚ β€’ Code      β”‚ β”‚   Child     β”‚            β”‚
β”‚ β”‚ β€’ Element   β”‚ β”‚ β€’ List      β”‚ β”‚ β€’ Navigationβ”‚            β”‚
β”‚ β”‚   Detection β”‚ β”‚ β€’ Struct    β”‚ β”‚ β€’ Metadata  β”‚            β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β€’ Fallback  β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜            β”‚
β”‚                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Current Architecture (Post-Migration)

The plugin now uses a layered architecture with clear separation of concerns:

Plugin Layer (dify-markdown-chunker):

  • Dify tool interface and parameter validation
  • Migration adapter for API compatibility
  • Input/output filtering and metadata control
  • Debug mode and hierarchical filtering

Core Engine (chunkana):

  • Advanced Markdown parsing and content analysis
  • Intelligent strategy selection and chunking algorithms
  • Hierarchical structure building and navigation
  • Performance optimization and memory management

Modules

ModuleDescription
tools/markdown_chunk_tool.pyDify tool implementation with parameter mapping
adapter.pyMigration adapter providing API compatibility
input_validator.pyInput validation and preprocessing
output_filter.pyOutput filtering and metadata control
chunkana libraryCore chunking engine (external dependency)

Project Structure

dify-markdown-chunker/
β”œβ”€β”€ tools/                     # Dify plugin tools
β”‚   β”œβ”€β”€ markdown_chunk_tool.py # Main tool implementation
β”‚   └── markdown_chunk_tool.yaml # Tool schema and metadata
β”œβ”€β”€ provider/                  # Dify plugin provider
β”‚   └── markdown_chunker.yaml  # Provider configuration
β”œβ”€β”€ adapter.py                 # Migration adapter (API compatibility)
β”œβ”€β”€ input_validator.py         # Input validation and preprocessing
β”œβ”€β”€ output_filter.py           # Output filtering and metadata control
β”œβ”€β”€ main.py                    # Plugin entry point
β”œβ”€β”€ tests/                     # Test suite (migration-compatible)
β”‚   β”œβ”€β”€ test_migration_*.py    # Migration compatibility tests
β”‚   β”œβ”€β”€ test_integration_*.py  # Integration tests
β”‚   └── test_*_adapted.py      # Adapted legacy tests
β”œβ”€β”€ docs/                      # Documentation
β”œβ”€β”€ manifest.yaml              # Dify plugin manifest
└── requirements.txt           # Dependencies (includes chunkana)

Migration Benefits

The migration to chunkana provides several advantages:

  • Maintainability: Core logic maintained in dedicated library
  • Performance: Optimized algorithms and memory management
  • Features: Access to latest chunking innovations
  • Compatibility: Full backward compatibility through adapter layer
  • Testing: Comprehensive test coverage for migration scenarios

⚑ Performance

Benchmark Results

The v2 architecture delivers excellent performance with linear scaling:

Document SizeProcessing TimeThroughputMemory
Tiny (1KB)2.3ms435 KB/s12.3 MB
Small (10KB)8.5ms1,177 KB/s14.5 MB
Medium (100KB)45.2ms2,212 KB/s28.4 MB
Large (1MB)412.5ms2,424 KB/s156.2 MB

Performance Characteristics:

  • Processing Speed: 0.42 ms/KB (based on regression analysis)
  • Throughput: ~2.4 MB/s peak (1MB files)
  • Scaling: Linear (RΒ² = 0.9987)
  • Memory: Base 12.3MB + 0.14MB per KB input

Performance Monitoring

Note: Performance data based on benchmarks from docs/research/07_benchmark_results.md conducted on Windows 11, Intel Core i7, 16GB RAM, SSD. Actual performance may vary depending on system configuration, document complexity, and content type.

chunker = MarkdownChunker(enable_performance_monitoring=True)

for doc in documents:
    chunker.chunk(doc)

stats = chunker.get_performance_stats()
print(f"Average time: {stats['chunk']['avg_time']:.3f}s")

For detailed benchmarks and methodology, see Performance Guide.


πŸ§ͺ Development

Testing

This project uses pytest for testing. The test suite has been cleaned up and optimized:

  • Migration-compatible tests: Tests that work with the current adapter-based architecture
  • Adapted tests: Legacy tests that have been adapted to use the migration adapter
  • Removed tests: 0 redundant tests were removed during cleanup

Running Tests

# Run all tests
make test-all

# Run only migration-compatible tests
make test

# Run specific test categories
pytest tests/test_migration_*.py  # Migration tests
pytest tests/test_integration_*.py  # Integration tests

Test Structure

The test suite is organized as follows:

  • tests/test_migration_*.py - Tests using the migration adapter
  • tests/test_integration_*.py - Integration tests
  • tests/test_*_adapted.py - Adapted legacy tests

Run all tests

make test-all

Run only migration-compatible tests

make test

Run specific test categories

pytest tests/test_migration_.py # Migration tests pytest tests/test_integration_.py # Integration tests


### Test Structure

The test suite is organized as follows:
- `tests/test_migration_*.py` - Tests using the migration adapter
- `tests/test_integration_*.py` - Integration tests
- `tests/test_*_adapted.py` - Adapted legacy tests
# Run all tests
make test-all

# Run only migration-compatible tests
make test

# Run specific test categories
pytest tests/test_migration_*.py  # Migration tests
pytest tests/test_integration_*.py  # Integration tests

Test Structure

The test suite is organized as follows:

  • tests/test_migration_*.py - Tests using the migration adapter
  • tests/test_integration_*.py - Integration tests
  • tests/test_*_adapted.py - Adapted legacy tests

Run all tests

make test-all

Run only migration-compatible tests

make test

Run specific test categories

pytest tests/test_migration_.py # Migration tests pytest tests/test_integration_.py # Integration tests


### Test Structure

The test suite is organized as follows:
- `tests/test_migration_*.py` - Tests using the migration adapter
- `tests/test_integration_*.py` - Integration tests
- `tests/test_*_adapted.py` - Adapted legacy tests
# Run all tests (473 total: 97 plugin + 376 chunkana)
make test

# Verbose output
make test-verbose

# With coverage report
make test-coverage

# Quick tests
make test-quick

# Performance benchmarks
python tests/performance/run_benchmarks_standalone.py

Code Quality

# Format code
make format

# Run linter
make lint

# Type checking
make quality-check

Building Plugin

# Validate structure
make validate

# Build package
make package

# Full release
make release

πŸ“¦ Dependencies

Core

  • markdown-it-py>=3.0.0 β€” Markdown parsing
  • mistune>=3.0.0 β€” Alternative parser
  • pydantic>=2.0.0 β€” Data validation
  • dify_plugin==0.5.0b15 β€” Dify integration

Development

  • pytest>=8.0.0 β€” Testing
  • hypothesis>=6.0.0 β€” Property-based testing
  • black>=23.0.0 β€” Code formatting
  • mypy>=1.5.0 β€” Type checking

🀝 Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.

# 1. Fork the repository
# 2. Create feature branch
git checkout -b feature/amazing-feature

# 3. Make changes with tests
# 4. Check quality
make test && make quality-check

# 5. Submit Pull Request

πŸ‘€ Author & Support

Author: Aleksandr Sukhodko (@asukhodko)
Repository: https://github.com/asukhodko/dify-markdown-chunker

Getting Help

  • Bug Reports & Feature Requests: GitHub Issues
  • Questions & Discussions: GitHub Discussions

πŸ“„ License

MIT License β€” see LICENSE


πŸ“ Changelog

Current Version: 2.1.6 (January 2026)

Latest: v2.1.6

Released: January 4, 2026

Changes:

  • βœ… Migration to chunkana 0.1.0 β€” Complete migration from embedded code to external chunkana library
    • Removed embedded markdown_chunker and markdown_chunker_v2 directories
    • Added migration adapter (adapter.py) for full API compatibility
    • All functionality preserved with improved maintainability and performance
    • Enhanced chunking algorithms and memory optimization through chunkana engine
  • βœ… Build System Improvements β€” Enhanced packaging and development workflow
    • Added automatic dify-plugin CLI installation in Makefile
    • Fixed package creation and validation commands
    • Improved code quality checks and linting
  • βœ… Testing Infrastructure β€” Comprehensive test coverage for migration
    • 99 migration-compatible tests passing
    • Property-based testing for correctness validation
    • Regression testing against pre-migration snapshots

Previous: v2.1.4

Released: December 23, 2025

Changes:

  • Bumped dify_plugin dependency from 0.5.0b15 to 0.7.0
  • Fixed .difyignore to properly include README.md and PRIVACY.md in package

v2.1.3

Released: December 14, 2025

New Features:

  • βœ… Table Grouping Option β€” Groups related tables in same chunk for better retrieval
    • Proximity-based grouping (max_distance_lines)
    • Section boundary awareness (require_same_section)
    • Size and count limits (max_group_size, max_grouped_tables)
    • Perfect for API documentation with Parameters/Response/Error tables

Previous: v2.1.2 (December 11, 2025)

  • Enhanced Code-Context Binding β€” Intelligent binding of code blocks to explanations
  • Adaptive Chunk Sizing β€” Automatic size optimization based on content complexity
  • Hierarchical Chunking β€” Parent-child relationships with navigation API

For full release history, see CHANGELOG.md.


πŸ“š Documentation & Resources

Core Documentation

  • Usage Guide - Comprehensive usage examples for plugin UI and direct chunkana
  • API Reference - Complete API documentation and parameter mapping
  • Configuration Guide - Detailed configuration options
  • Chunking Strategies - Strategy details and selection guide

Migration & Compatibility

  • Migration to Chunkana Guide - Complete migration information
  • Troubleshooting Guide - Common issues and solutions

Advanced Usage

  • chunkana Documentation - Full chunkana library documentation
  • Performance Guide - Performance optimization tips
  • Developer Guide - Development and contribution guide

Quick Links

  • GitHub Repository - Source code and issues
  • Latest Release - Download plugin
  • GitHub Discussions - Questions and community

⬆ Back to Top