by textin · v1.1.1
Designed specifically for downstream tasks of large language models (LLMs), this service can recognize text from documents or images and intelligently restore the reading order of the content, outputting a standardized Markdown format. It supports OCR recognition across more than 10 common document layouts and over 52 languages, enabling LLMs to more efficiently utilize document content in scenarios such as understanding, generation, and question answering.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Author: textin
Version: 1.1.1
Type: tool
TextIn OCR is a general-purpose document parsing service designed specifically for downstream tasks of large language models (LLMs). It can recognize text information from documents or images, including financial reports, national standards, academic papers, corporate announcements, user manuals, financial invoices, and more. It extracts key information and supports restoring the document content into standard Markdown format. TextIn offers OCR text recognition covering more than 10 common document layouts and supports over 52 languages, helping various large models efficiently utilize document data in scenarios such as comprehension, generation, and question answering.
import json
def main(response):
res = response[0]
return {
"markdown": res["result"]["markdown"]
}
Q: What types of documents does TextIn support? A: It supports a wide range of formats including images (JPG, PNG, etc.), PDF files, Word documents (DOC/DOCX), HTML, Excel(xlsx, csv), PPT(pptx), TXT and more.
Q: What languages does TextIn support for text recognition? A: Currently, it supports over 52 languages, including major ones such as Chinese, English, Japanese, Korean, German, and French.
Q: What is the output format after document parsing? A: The default output is structured Markdown format, and JSON format is also supported for content extraction.
Q: Are there any usage limits? A: Usage limits vary based on the subscription plan and account type, including limits on the number of calls and concurrency. Please refer to the official TextIn OCR documentation for details.
Q: How do I obtain and manage my API key? A: After registering via the TextIn OCR, you can generate and manage your APP ID and APP SECRET in the User Center for calling the TextIn OCR service.
Q: Does TextIn support batch document processing? A: Yes, it supports batch uploading of documents via API for recognition and parsing, which is suitable for large-scale document processing.
Q: Will the parsed document retain the original layout? A: TextIn restores the document structure based on common reading order, but extremely complex layouts may undergo some adjustments.
Q: Are SDKs or code samples provided? A: Yes, TextIn offers multi-language SDKs (e.g., Python, Java) and comprehensive API usage examples to facilitate quick integration.