by langgenius · v0.1.4
SoMark is a DocAI that can convert diverse documents—such as PDFs, images, and more—into structured Markdown or JSON. It is designed to work seamlessly across all scenarios.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
SoMark is a DocAI that can convert diverse documents—such as PDFs, images, and more—into structured Markdown or JSON. It is designed to work seamlessly across all scenarios.
It breaks the traditional trade-off between accuracy, speed, and cost, delivering precise document parsing in milliseconds with minimal hardware resources.
The resulting structured data is AI-native, ready to power LLM training, enhance RAG systems, and enable intelligent agents.
SoMark pioneers the proprietary "OXR" algorithm, extending traditional OCR (Optical Character Recognition) into Optical Everything Recognition.
From basic layout segmentation and reading-order recovery to complex elements such as tables, formulas, images, and even chemical notations, every component can be accurately extracted and reconstructed. The output is a complete, highly structured representation of the document.
Built on this powerful OXR algorithm, SoMark achieves the perfect balance of accuracy, speed, and cost:
SoMark delivers strong general-purpose recognition capability. A single API call handles document parsing across all formats and scenarios.
Log into your Dify platform.
Go to "Tools" -> "Plugin Market", search for the "SoMark" plugin and add it.
Configure the SoMark plugin parameters:
Base URL (required): The address of the SoMark service.
https://somark.cn/api/v1 (Mainland China) or https://somark.ai/api/v1 (Taiwan, China; Hong Kong, China; Macau, China; and other overseas regions).https://somark.your-domain.com/api/v1). The URL must start with http:// or https://.API Key:
Save your configuration.
In your Dify workflow, click "+" to add a new node, select "Tools", then find and add the SoMark > SoMark Document Parser node.
In the SoMark Document Parser node panel, configure the File input:
{x} in the File input field.For more parameters (such as output format and feature toggles), see Input Parameters below.
Note:
After the node executes, its output variables become available for all downstream nodes (e.g., LLM, Text Splitter, Code node). Click {x} in any downstream node's input field and select from the SoMark node's output variables.
| Parameter | Type | Required | Description |
|---|---|---|---|
File | File | Yes | Supported files: PDF, PNG, JPG, JPEG, BMP, TIFF, JP2, DIB, PPM, PGM, PBM, GIF, HEIC, HEIF, WEBP, XPM, TGA, DDS, XBM, DOC, DOCX, PPT, PPTX, XLS, XLSX. Max 200 MB / 300 pages. |
Output Formats | Single-select | No | Select the output format. Supported options: Markdown, JSON, Both "Markdown" and "JSON". Default: Both "Markdown" and "JSON". |
Image Format | Single-select | No | Image output format. Supported options: URL, Base64, None. Default: URL. |
Formula Format | Single-select | No | Formula output format. Supported options: LaTeX, MathML, ASCII. Default: LaTeX. |
Table Format | Single-select | No | Table output format. Supported options: HTML, Markdown, Image. Default: HTML. In Markdown mode, merged cells are expanded into individual cells with duplicated content. |
Chemical Structure Formula Format | Single-select | No | Chemical structure output format. Supported options: Image. Default: Image. |
Enable Text Cross Page | True / False | No | Merge text that spans across pages into a continuous paragraph. Default: False. |
Enable Table Cross Page | True / False | No | Merge tables that span across pages into a continuous table. Default: False. |
Enable Title Level Recognition | True / False | No | Recognize heading hierarchy such as H1/H2/H3. Default: False. |
Enable Inline Image | True / False | No | Return images embedded in text paragraphs. Default: False. |
Enable Table Image | True / False | No | Return images embedded in table cells. Default: True. |
Enable Image Understanding | True / False | No | Perform semantic understanding and structured description for images in the document. Default: True. |
Keep Header Footer | True / False | No | Keep page headers and footers instead of filtering them out. Default: False. |
The node exposes the following output variables:
markdown string — The parsed document content in Markdown format, preserving the original layout structure including headings, tables, lists, formulas, and images.
json_str string — The parsed document content in JSON string format, containing structured data for document elements such as text blocks, tables, formulas, images, coordinates, and page information. Suitable for advanced downstream processing in a Code node after JSON parsing.
text / files — Dify built-in variables, not populated by this plugin.
Add the SoMark tools to an agent or workflow, fill in the required inputs, and run the node to call the upstream service.
This plugin sends the inputs required by the selected operation to the upstream service. Review the upstream service's privacy policy before use.