Skip to main content
ZICQ

Skills ZICQ category:Documents ocr-document-processor

Ocr Document Processor

Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.

4725 installs

Official URL:skills.sh

What this skill does

Intro in this page language first. The official description stays in its original wording; we do not rewrite SKILL.md.

What it does

Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.

When to use it

The official description does not include a separate “Use when”. Per the spec, agents activate this skill when the task matches keywords in that description.

How agents load it

Per Agent Skills progressive disclosure: name and description load at startup (~100 tokens); the full SKILL.md body loads when the skill activates; scripts/, references/, and assets/ load only as needed. This file's sections: OCR Document Processor; Use This For; Workflow; Guardrails. It includes spec-recommended sections: step-by-step instructions.

File analysis

File analysis: besides SKILL.md, the body references scripts/ocr_processor.py, scripts/business_card_scanner.py, scripts/receipt_scanner.py. Those resources load on demand.

OCR Document ProcessorUse This ForWorkflowGuardrails

Source category:skills.sh agent-skill

SKILL.md & Agent activation

Official spec ↗
name
ocr-document-processor
description
Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.
  1. DiscoverThe client exposes names and descriptions to the agent.
  2. ActivateYour request or the task context selects the skill and loads its instructions.
  3. Load resourcesReferenced scripts, documentation and assets are used when needed.
Files referenced by the instructions · 3
  • scripts/ocr_processor.py
  • scripts/business_card_scanner.py
  • scripts/receipt_scanner.py

These paths are extracted from the text. Check the upstream package to verify the files exist.

Invocation syntax and available tools depend on your Agent client. Client integration guide ↗

Install this skill

Skills CLI ↗

Choose the target agent and installation scope, keep referenced package files, then verify the skill appears in the client's catalog.

This skill references supporting files. Retrieve the complete directory from the source; copying SKILL.md alone may leave missing dependencies.

Ask your Agent to install

Copy these instructions to a compatible agent and confirm the target directory matches your client.

Install the agent skill "ocr-document-processor" into my project. The full SKILL.md and official description are at https://zicq.com/en/skills/skl-9ea686d493b59fac-Ocr-Document-Processor.html
Save it as .cursor/skills/ocr-document-processor/SKILL.md or .claude/skills/ocr-document-processor/SKILL.md and keep the frontmatter name and description exactly as-is.
This skill also ships scripts/, references/, or assets/ — fetch the whole folder from https://github.com/dkyazzentwatwa/chatgpt-skills instead of creating only a SKILL.md.

Full package on GitHub ↗

Install from the terminal · Skills CLI

Requires Node.js and npx. First inspect the repository's skill list to confirm the name.

npx skills add 'https://github.com/dkyazzentwatwa/chatgpt-skills' --list

npx skills add 'https://github.com/dkyazzentwatwa/chatgpt-skills' --skill 'ocr-document-processor'

The CLI lets you choose the agent interactively. The default scope is the project; use -g for user scope. Confirm package availability with the discovery command, then use npx skills list to inspect installed skills.

Readable layout

name: ocr-document-processor description: Extract text and structure from scans, images, and scanned PDFs. Use for OCR, searchable PDFs, table extraction, receipt parsing, and business card parsing.

OCR Document Processor

Handle OCR-heavy inputs where text must be recovered from images or scanned pages.

Use This For

  • OCR on images and scanned PDFs
  • Searchable PDF export
  • Structured extraction to text, markdown, JSON, or HTML
  • Table extraction from scanned material
  • Receipt parsing and business card parsing

Workflow

  1. Decide whether plain OCR, structured extraction, or document-specific parsing is needed.
  2. Preprocess noisy inputs before extraction when skew, blur, or shadows are present.
  3. Use scripts/ocr_processor.py for core OCR tasks.
  4. Use the focused helpers when the input is specialized:
    • scripts/business_card_scanner.py
    • scripts/receipt_scanner.py
  5. Return confidence caveats when the source is low quality, rotated, handwritten, or multilingual.

Guardrails

  • Prefer explicit language selection when accuracy matters.
  • Do not claim fields are exact when OCR confidence is weak.
  • Route non-scanned digital PDFs to document-converter-suite instead of OCR by default.

Related skills

Documents

Ontology

Typed knowledge graph for structured agent memory and composable skills. Use when creating/querying entities (Person, Project, Task, Event, …

Documents

Nano Pdf

Edit PDFs with natural-language instructions using the nano-pdf CLI.

Documents

Word / DOCX

Create, inspect, and edit Microsoft Word documents and DOCX files with reliable styles, numbering, tracked changes, tables, sections, and co…

Documents

Baidu Search

Search the web using Baidu AI Search Engine (BDSE). Use for live information, documentation, or research topics.