
book-to-skill fixes a familiar problem: you read a great technical book once, and 3 months later you can't remember chapter 7 existed. Point it at the PDF and it builds an agent skill from it: the book's frameworks, decision rules and anti-patterns, with one file per chapter. Then you ask your coding agent a question, and it reads the right chapter and answers from the real content.
It's structure, not a summary. And it works on anything you re-read often: internal docs, brand guidelines, specs, or a pile of research papers.

| Author | virgiliojr94 on GitHub |
| Released | May 2026, regular releases since |
| Popularity | About 33,000 GitHub stars |
| Language | Python for the extractor, a SKILL.md spec for the generator |
| Licence | MIT |
| Works with | Claude Code, GitHub Copilot CLI, Amp, OpenCode, OpenClaw, Hermes Agent: any host that reads the open Agent Skills format |
| Inputs | PDF, EPUB, DOCX, HTML, RTF, MOBI, Markdown, plain text, folders and globs |
What it generates
A run creates a skill folder with 5 kinds of file:
| File | What it holds | Size |
|---|---|---|
SKILL.md |
The core mental models and a chapter index | about 4,000 tokens |
chapters/ch01-*.md ... |
1 file per chapter, loaded only when asked | about 1,000 tokens each |
glossary.md |
Every key term, with chapter references | about 1,500 tokens |
patterns.md |
Techniques, algorithms and design patterns | about 2,000 tokens |
cheatsheet.md |
Decision tables and quick rules | about 1,000 tokens |
The chapter files cost nothing until a question needs them. The author measured 24 to 51 times fewer tokens than putting the whole book in the agent's context to answer 1 question.
For business people
The problem: books, playbooks and internal documentation hold knowledge that nobody re-reads. Searching a PDF gives you pages, not answers. Asking an AI about a book it hasn't read gives you guesses.
What it changes: the book becomes a reference your AI assistant actually uses. Type /my-book pricing and the agent answers from that book's chapter on pricing.
Where it fits beyond books:
- Internal documentation: runbooks, onboarding guides and decision records as 1 skill.
- Brand and design guidelines: a 60-page brand book the team can query.
- Research: papers plus your own notes, merged and updated as new material arrives.
- Specs and standards: API contracts, RFCs and compliance documents.
Costs: the tool is free. The generation step runs on your agent's model, once per book.
Copyright, from the README: the tool ships no book content. You convert files you own, and the result counts as your own study notes. Don't publish a skill made from someone else's book: keep those private.
For technical people
Install it as a skill:
npx skills add virgiliojr94/book-to-skill
# or
git clone https://github.com/virgiliojr94/book-to-skill.git ~/.claude/skills/book-to-skill
Use it from the agent:
/book-to-skill ./designing-data-intensive-applications.pdf
/book-to-skill ./docs/ internal-handbook
/ddia replication
- 2 halves: a deterministic Python extractor (document to clean text and metadata) and a generator, which is your agent following the
SKILL.mdspec. - Where skills land:
~/.agents/skills/<slug>/, the shared folder most agents read, with a checked symlink into~/.claude/skills/<slug>/for Claude Code. - PDF extraction: it asks whether the book is technical or prose.
doclingkeeps tables and code blocks (about 1.5 seconds a page),pdftotextis instant for prose. Scanned PDFs need OCR first, for example withocrmypdf. - Check your setup:
python3 scripts/extract.py --checklists the installed extractors and what's missing. - Modes: analyse only, generate from an earlier analysis, and update to fold new material into an existing skill. It can also publish the skill to a private GitHub repo for
npx skills addon other machines. - Validation:
tools/validate_skill.pychecks a generated skill against each host's rules.
Value
I keep reading notes on sales and leadership books in this notebook, but notes are for me. A skill is for my agents. The obvious first candidates are the books my sales work is built on, and the Kaltura brand and product guidelines, so a drafting agent answers from the source instead of from memory.
The update mode matters most: a skill that grows as I add material, instead of a one-off conversion. And the copyright rule is simple: skills from bought books stay private, on my machine.
Further Reading

