<< All versions
Skill v1.0.0
currentAutomated scan100/100opencoven/coven/ocr
──Details
PublishedOctober 3, 2026 at 10:46 AM
Content Hashsha256:e94193782e06154f...
Git SHA24e4c1f3b904
──Files
Files (1 file, 3.1 KB)
SKILL.md3.1 KBactive
SKILL.md · 72 lines · 3.1 KB
version: "1.0.0" name: ocr description: Extract text from screenshots, scanned documents, image files, and PDFs. Use when a user asks to OCR, read text from an image/screenshot/scan, transcribe visible text, extract text from a scanned PDF, compare OCR text, or convert image/PDF text into Markdown/JSON/plain text.
OCR
Use this skill to extract text from local screenshots, images, scans, and PDFs with a repeatable workflow.
Quick start
Prefer the bundled helper when the user needs exact-ish text extraction from a local file:
bash
python3 ~/.openclaw/workspace/skills/ocr/scripts/ocr.py <path> --format text
Useful variants:
bash
# Structured output with per-line confidence/bounding boxespython3 ~/.openclaw/workspace/skills/ocr/scripts/ocr.py screenshot.png --format json# Markdown output from a scanned PDF, OCRing up to the first 5 pagespython3 ~/.openclaw/workspace/skills/ocr/scripts/ocr.py scan.pdf --force-ocr --max-pages 5 --format markdown# Multiple languages for macOS Vision OCRpython3 ~/.openclaw/workspace/skills/ocr/scripts/ocr.py receipt.jpg --languages en-US,es-ES --format text
Workflow
- Locate the file. If the image/PDF is attached in chat, use the local attachment path the runtime provides.
- For a single image or scanned PDF, run
scripts/ocr.py. - For text-native PDFs, let
ocr.pyusepdftotextfirst; only use--force-ocrwhen the PDF is a scan or the extracted text is wrong. - Inspect output before relying on it. OCR can confuse punctuation, columns, totals, handwriting, low contrast, and small text.
- If the user needs high confidence, run JSON output and mention low-confidence lines or ambiguous characters.
- If the user asks for a clean transcript, lightly normalize whitespace but do not silently rewrite wording.
Tool choices
- Use
scripts/ocr.pyfor local images/PDFs and repeatable extraction. - Use the model
imagetool when the user asks for visual understanding, layout interpretation, or when OCR output alone is not enough. - Use
pdftotextdirectly only for text-native PDFs when no OCR is needed.
Script behavior
scripts/ocr.py supports:
- image inputs: PNG, JPEG, TIFF, BMP, GIF, WebP, HEIC/HEIF
- PDF inputs: embedded text via
pdftotext, scanned pages viapdftoppm+ macOS Vision OCR - directory inputs: processes supported files in sorted order
- output formats:
text,json,markdown
Dependencies:
- macOS
swift+ Vision framework for image OCR pdftotextfor text-native PDFspdftoppmfor scanned PDF rasterization
If a dependency is missing, install it or fall back to the image tool for one-off extraction.
Reporting results
Be explicit about uncertainty:
- Say “OCR reads…” or “I extracted…” rather than presenting uncertain OCR as ground truth.
- Preserve line breaks when they matter, especially for receipts, forms, addresses, code, and tables.
- For sensitive documents, summarize only what the user requested and avoid exposing unnecessary personal data.
- For financial/legal/medical text, flag that OCR may need human verification before decisions.