Skip to main content
Guide9 min read·Updated August 20, 2026
🧩

Best AI Agent Skills for Document Processing & OCR 2026

B

A. Frans

Published August 20, 2026

Document ProcessingOCRPDFAgent SkillsClaude Code

Anthropic's official PDF skill handles about 80% of document work and does it well. The other 20% is why a community stack exists: scanned invoices with no text layer, Korean HWP filings, Chinese-layout academic PDFs, tables that live inside images, and the 400-page discovery dump nobody wants to page through.

Eight skills cover that territory in 2026. Two are official and audited. The rest range from carefully maintained to one-person side projects, and the difference matters more here than in most skill categories, because document processing means handing files to code you didn't write.

The eight

SkillMaintainerTrust tierSecurityLicenseJob
PDF ToolsAnthropicOfficialAuditedMITRead, create, merge, split, extract tables, watermark
DOCX (Word)AnthropicOfficialAuditedMITRead, edit, generate Word documents
XLSX (Spreadsheet)AnthropicOfficialAuditedMITRead, write, analyze Excel and CSV
KreuzbergGoldziherVerifiedCommunity-reviewedOtherPolyglot document intelligence, Rust core
Skill SeekersyusufkaraaslanVerifiedCommunity-reviewedMITTurn docs sites, repos, and PDFs into skills
KordocchrisryugjCommunityUnreviewedMITHWP, HWPX, PDF, XLSX, DOCX to Markdown
PDF Reader MCPSylphxAICommunityUnreviewedMITParallel PDF processing, 5-10x throughput
MinerU Tianshumagicyuan876CommunityUnreviewedApache-2.0PDF and Office to Markdown at platform scale

Start with the official three

The Anthropic skills repo covers PDF, DOCX, and XLSX, all audited, all MIT. There is no claude skill add command, whatever a directory listing may tell you: the Claude CLI exposes claude mcp and claude plugin, and skills arrive either through a plugin marketplace or by living in your skills directory.

The marketplace route, from inside Claude Code:

/plugin marketplace add anthropics/skills

Or clone them straight into your personal skills directory, where each skill is a folder holding a SKILL.md:

git clone https://github.com/anthropics/skills.git /tmp/anthropic-skills
cp -r /tmp/anthropic-skills/skills/{pdf,docx,xlsx} ~/.claude/skills/

PDF Tools reads, creates, merges, splits, watermarks, and extracts text and tables. DOCX handles Word documents including tracked changes and comments. XLSX covers .xlsx, .xlsm, and .csv with formulas, charts, and cell formatting.

Install all three before you evaluate anything else, because they set the baseline. Half the people who go hunting for a document skill are solving a problem the official PDF skill already handles. Table extraction from a normal text-layer PDF is one of those problems.

Where the official set stops: no OCR. If the PDF is a scan, there's no text to extract and you get an empty result rather than an error, which is the worst kind of failure because it looks like the document was empty. That's the line where you need something else.

Kreuzberg — the serious extraction engine

Repo: github.com/Goldziher/kreuzberg. Verified tier, community-reviewed, and the license reads as Other rather than MIT, which is worth reading before commercial use. It ships Python and Node bindings rather than a one-line MCP install, so follow the repo's setup instructions rather than guessing an npx invocation.

Kreuzberg calls itself a polyglot document intelligence framework with a Rust core, extracting text, metadata, and images.

The Rust core is the reason to care. Document extraction at volume is CPU-bound, and Python-based pipelines fall over on batches that Kreuzberg chews through. If your work is one document at a time, you won't notice. If it's 5,000 contracts, you will.

Format coverage is the other draw. Where the official PDF skill is a PDF skill, Kreuzberg treats format as an implementation detail and gives you text plus structure regardless of what came in.

This is the pick if document processing is a core part of your workflow rather than an occasional task.

Kordoc — the one that handles HWP

claude mcp add kordoc -- npx -y kordoc

Kordoc converts HWP3-5, HWPX, HWPML, PDF, XLS/XLSX, and DOCX to Markdown, with both a CLI and an MCP server. Community tier, unreviewed, MIT, maintained by chrisryugj.

HWP is the reason this exists. Hangul Word Processor is the standard format for Korean government and corporate documents, and almost nothing outside Korea reads it. If your work touches Korean filings, regulatory submissions, or academic material, Kordoc goes from niche to essential immediately.

The unreviewed security status is real, and the repo documentation is largely in Korean. Read it anyway before you point it at anything sensitive.

MinerU Tianshu — layout-aware conversion at scale

Repo: github.com/magicyuan876/mineru-tianshu. Community tier, unreviewed, Apache-2.0. It's an enterprise-oriented data preprocessing platform, PDF and Office to Markdown, Vue3 plus FastAPI, with MCP protocol support layered on. There's no npm package: you deploy the stack and point your MCP config at it, per the repo.

The strength is layout handling on documents that break naive extractors: multi-column academic papers, mixed CJK and Latin text, formulas, and tables that span pages. If you've watched a PDF extractor interleave two columns into unreadable soup, this is the class of tool that fixes it.

It's the heaviest install here by a distance. You're standing up a platform, not adding a skill. Justified for a research group processing thousands of papers; absurd for someone who needs to read one report.

PDF Reader MCP — throughput and nothing else

Repo: github.com/SylphxAI/pdf-reader-mcp. Community tier, unreviewed, MIT. It advertises 5-10x faster processing through parallelism and claims 94%+ test coverage. Build and run it from the repo; it isn't published to npm under that name.

Narrow by design and honest about it. If your bottleneck is wall-clock time on a large PDF batch and the official skill is too slow, this replaces it. If your bottleneck is anything else, it doesn't help. The test coverage claim is a good sign for a single-maintainer project, and it's still unreviewed.

Skill Seekers — the sideways one

Repo: github.com/yusufkaraaslan/Skill_Seekers. Verified tier, community-reviewed, MIT. It's a Python project, so clone and follow its README rather than reaching for npx. It converts documentation sites, GitHub repos, and PDFs into Claude skills, with automatic conflict detection.

This belongs on the list because of what it does with a PDF rather than to it. Instead of extracting text for one-time reading, it turns a manual into a reusable skill your agent loads on demand. Point it at a 300-page API reference and you get something the agent consults, instead of 300 pages of context you pay for on every prompt.

For teams with internal documentation locked in PDFs, this pays off faster than anything else on the list.

Where SaaS still beats skills

Agent skills lose to purpose-built document AI on one axis: trained extraction accuracy on high-volume repetitive forms.

Nanonets trains custom models per document type including handwriting, priced from $0.30/page with enterprise plans from $499/mo. Rossum processes transactional documents in 276 languages with a 14-day trial and custom pricing. Docsumo claims 95%+ accuracy on invoices and bank statements, free for 100 pages a month then $500/mo Growth. Mindee covers invoices, receipts, and passports, free for 500 pages a month then from $30/mo. Parseur mixes AI extraction with template-based parsing for consistent layouts, free to start.

The dividing line is repetition. If you process 3,000 invoices a month in twelve known formats, a trained model beats a general-purpose skill on accuracy and it isn't close. Pay for it. We compared three of these directly in Nanonets vs Rossum vs Vic.ai.

If you process varied documents where each one is a bit different, skills win, because there's nothing to train on and a general extractor plus a capable model beats a model trained on the wrong distribution.

Notice that Mindee's 500 free pages a month covers a lot of small-business volume. Check the free tiers before assuming this category requires a budget.

The security part

Four of the eight skills here are unreviewed community projects, and document processing has a specific risk profile that makes that worth pausing on.

You're handing files to code you didn't audit. Those files are frequently the sensitive ones: contracts, filings, financial statements, medical records. An MCP server with filesystem access and a network connection can do more than parse.

Reasonable precautions, in order of how much they buy you:

Run unreviewed servers against a copy in a scratch directory rather than pointing them at your document store. Pin package versions instead of tracking latest through npx -y. Read what network calls the server makes; a local parser has no reason to phone home. And for anything under privilege or regulatory constraint, stay on the audited official skills or a vendor with a signed DPA.

The trust tiers on this site exist for exactly this. Official-and-audited versus community-and-unreviewed is a real distinction, not a badge.

Adjacent reading: contract review skills and legal research skills cover the downstream work once documents are parsed, and if your documents are financial, the bookkeeping skills roundup is the closer fit. Professionals who live in documents all day may also want our lists for lawyers and accountants.

Picking one

Install the three official skills. That's not a recommendation, it's the floor.

Add Kreuzberg if documents are central to your work and you need format-agnostic extraction that holds up at volume. It's the best non-official option on this list and the verified tier plus community review makes it defensible for client work.

Add Kordoc only for Korean formats, MinerU Tianshu only for heavy academic or CJK layout work, and PDF Reader MCP only when throughput is measurably your problem.

Add Skill Seekers regardless of the above, because turning reference PDFs into loadable skills is a different kind of win and most people haven't thought to do it.

Skip the whole category and buy Mindee or Docsumo if your documents are repetitive forms at volume. Trained models exist for a reason.

FAQ

Do any of these do OCR on scanned documents? Kreuzberg and MinerU Tianshu handle scanned input; the official PDF skill does not and will return nothing rather than warning you. If your source is a scan, test with a known document first so you can tell empty output from failed output.

What's the difference between a skill and an MCP server here? A skill is a folder with a SKILL.md that the agent loads as instructions plus scripts, installed through a plugin marketplace or by dropping it in ~/.claude/skills/. An MCP server installs with claude mcp add, runs as a separate process, and exposes tools. Kordoc, MinerU, Kreuzberg, PDF Reader, and Skill Seekers are MCP servers; the official three and Markdown Exporter are skills. Functionally the agent uses both the same way.

Can I extract tables reliably? From a text-layer PDF, yes, the official PDF skill does it well. From a scan, expect to check every result. From a multi-column academic paper, MinerU Tianshu is the one built for it. No tool in this category is accurate enough on tables to skip verification when the numbers matter.

Is Kreuzberg's license a problem for commercial use? It's listed as Other rather than a standard MIT or Apache tag, which means read the repo before you ship it in a product. For internal use this rarely matters; for anything redistributed, it does.

How do I convert extracted content back into documents? Markdown Exporter covers the return trip, turning Markdown into DOCX, PPTX, XLSX, PNG, PDF, HTML, CSV, JSON, and XML. Community tier, unreviewed, Apache-2.0, maintained by bowenliang123 at github.com/bowenliang123/markdown-exporter.

Share this article

📬

Get More AI Tool Guides

New comparisons and guides every week. Join thousands of professionals staying ahead of the AI curve.