Introduction
What is Image to Text?
Image to Text is an AI-powered OCR tool that extracts text from images and documents with industry-leading accuracy. Built on PaddleOCR-VL, it supports 111+ languages with structured data output.
Key Features
- Multi-Language OCR: Support for 111+ languages including Chinese, English, Japanese, Korean, Arabic, and more
- Document AI Parsing: Convert complex PDFs and documents into structured Markdown and JSON
- Vision-Language Model: PaddleOCR-VL 1.5 uses a compact 0.9B parameter model for fast, accurate text recognition
- Real-World Scenarios: Handles skewed, warped, scanned, and low-light images
Documentation
- Getting Started - Create an API key and make your first OCR request
- OCR Integration Guide - Practical patterns for integrating OCR into your application
- OCR API Reference - Endpoint parameters, response schema, and examples
- Image Processing API - Developer guide for image conversion and compression