Image OCR

Upload an image or screenshot; AI extracts editable plain text via OCR.

OCR Text Recognition extracts editable plain text from images and screenshots so the content can be searched, quoted, or pipelined into other modules for summarization, translation, or analysis. Upload a photo, scanned document, screenshot, or screen capture and the AI scans the entire frame and returns the recognized text in reading order. Common use cases include digitizing paper receipts and invoices for bookkeeping, converting handwritten notes into searchable text, lifting specific paragraphs out of books or papers before sending them to the paper or summary module, turning whiteboard photos into editable meeting-minute drafts, and pasting code or error-message screenshots into the programming module for debugging. Mixed Chinese/English, Traditional/Simplified Chinese, and partial Japanese are supported; very messy handwriting or low-resolution images may degrade accuracy, so a clearer mobile photo or better lighting helps. Output can be copied directly, paired with sc2tc / tc2sc for script conversion, or chained into translate_to_english / translate_to_tc for one-stop digitization-plus-translation. Maximum context window (input and output combined): about 1,000,000 tokens. Maximum output per reply: about 64,000 tokens (model reasoning, if any, counts toward this limit). Maximum total upload size per request: 256MB (all files and message content combined). Accepted document formats: pdf, docx, xlsx, pptx, txt, csv (file content is provided to the model as extracted text). Accepted image formats: gif, jpeg, jpg, png, webp. Animated GIF images are analyzed using only the first frame. Maximum number of images: 600. When uploading more than 20 images, each side must be at most 2,000 pixels. Maximum size per image: 10MB (base64-encoded). Maximum total image size per message: 20MB (text and history share the same request size limit; uploads close to the limit may still be rejected). Image dimensions: 200 to 8,000 pixels per side. Output content formats: Text This module does not support continuous conversations. AI can make mistakes. Please verify important information.