pdfminer_six · emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginspdfminer_six
Plugin
Limited listing

pdfminer_six

by tavan · v0.0.1

pdfminer_six extracts the text from a page directly from the sourcecode of the PDF.

1.2k installsUpdated May 29, 2025
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Tools

Available inside your emploidai workspace after installation.

Data sources

Available inside your emploidai workspace after installation.

Category

tool

Version

0.0.1tavan

Requirements

Maximum memory 256MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.0.0+

Permissions

  • Uses app capability
  • Uses endpoint capability
  • Uses storage capability
  • Uses tool capability

Dependencies

No additional dependencies

Resources

Privacy policy
emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

pdfminer_six

Author: tavan Version: 0.0.1 Type: Tool plugin

Introduction

pdfminer.six is ​​a powerful PDF document parsing tool that focuses on text extraction and analysis. It can directly extract text content from PDF source code and supports obtaining the precise location, font and color information of the text.

Main features

  • Support PDF-1.7 specification
  • Support CJK languages ​​and vertical writing scripts
  • Support multiple font types (Type1, TrueType, Type3, CID)
  • Support RC4 and AES encryption
  • Support form extraction
  • Support directory extraction
  • Support automatic layout analysis

Supported output formats

The plugin supports the following output formats:

  • markdown - Markdown format (MIME: text/markdown)

  • html - HTML format (MIME: text/html)

  • text - Plain text format (MIME: text/plain)

  • tag - Tagged text format (MIME: text/plain)

  • xml - XML ​​format (MIME: application/xml)

Usage Guide

  1. Upload file
  • Support single PDF file upload
  • The file must be in valid PDF format
  1. Select output format
  • Specify output_type in the parameter
  • Optional values: markdown, html, text, tag, xml
  • Text format is used by default

  1. Processing results The plugin will return responses in three formats:
  • Text message: Processing status description
  • Blob message: Converted content
  • JSON message: Processing result metadata

Future Enhancements

  • Support hocr

License

This project is licensed under the MIT License.