book-to-skill

An agent skill that turns a technical book, a docs folder or a stack of papers into a structured skill: core models, one file per chapter, glossary and cheatsheet, loaded only when needed.

book-to-skill fixes a familiar problem: you read a great technical book once, and 3 months later you can't remember chapter 7 existed. Point it at the PDF and it builds an agent skill from it: the book's frameworks, decision rules and anti-patterns, with one file per chapter. Then you ask your coding agent a question, and it reads the right chapter and answers from the real content.

It's structure, not a summary. And it works on anything you re-read often: internal docs, brand guidelines, specs, or a pile of research papers.

book-to-skill banner: Booklin, a cartoon wizard in a purple robe and hat, holds an open book whose pages turn into sparkles that settle into a neat grid of squares, on a purple background

Author virgiliojr94 on GitHub
Released May 2026, regular releases since
Popularity About 33,000 GitHub stars
Language Python for the extractor, a SKILL.md spec for the generator
Licence MIT
Works with Claude Code, GitHub Copilot CLI, Amp, OpenCode, OpenClaw, Hermes Agent: any host that reads the open Agent Skills format
Inputs PDF, EPUB, DOCX, HTML, RTF, MOBI, Markdown, plain text, folders and globs

What it generates

A run creates a skill folder with 5 kinds of file:

File What it holds Size
SKILL.md The core mental models and a chapter index about 4,000 tokens
chapters/ch01-*.md ... 1 file per chapter, loaded only when asked about 1,000 tokens each
glossary.md Every key term, with chapter references about 1,500 tokens
patterns.md Techniques, algorithms and design patterns about 2,000 tokens
cheatsheet.md Decision tables and quick rules about 1,000 tokens

The chapter files cost nothing until a question needs them. The author measured 24 to 51 times fewer tokens than putting the whole book in the agent's context to answer 1 question.

For business people

The problem: books, playbooks and internal documentation hold knowledge that nobody re-reads. Searching a PDF gives you pages, not answers. Asking an AI about a book it hasn't read gives you guesses.

What it changes: the book becomes a reference your AI assistant actually uses. Type /my-book pricing and the agent answers from that book's chapter on pricing.

Where it fits beyond books:

  • Internal documentation: runbooks, onboarding guides and decision records as 1 skill.
  • Brand and design guidelines: a 60-page brand book the team can query.
  • Research: papers plus your own notes, merged and updated as new material arrives.
  • Specs and standards: API contracts, RFCs and compliance documents.

Costs: the tool is free. The generation step runs on your agent's model, once per book.

Copyright, from the README: the tool ships no book content. You convert files you own, and the result counts as your own study notes. Don't publish a skill made from someone else's book: keep those private.

For technical people

Install it as a skill:

npx skills add virgiliojr94/book-to-skill
# or
git clone https://github.com/virgiliojr94/book-to-skill.git ~/.claude/skills/book-to-skill

Use it from the agent:

/book-to-skill ./designing-data-intensive-applications.pdf
/book-to-skill ./docs/ internal-handbook
/ddia replication
  • 2 halves: a deterministic Python extractor (document to clean text and metadata) and a generator, which is your agent following the SKILL.md spec.
  • Where skills land: ~/.agents/skills/<slug>/, the shared folder most agents read, with a checked symlink into ~/.claude/skills/<slug>/ for Claude Code.
  • PDF extraction: it asks whether the book is technical or prose. docling keeps tables and code blocks (about 1.5 seconds a page), pdftotext is instant for prose. Scanned PDFs need OCR first, for example with ocrmypdf.
  • Check your setup: python3 scripts/extract.py --check lists the installed extractors and what's missing.
  • Modes: analyse only, generate from an earlier analysis, and update to fold new material into an existing skill. It can also publish the skill to a private GitHub repo for npx skills add on other machines.
  • Validation: tools/validate_skill.py checks a generated skill against each host's rules.

Value

I keep reading notes on sales and leadership books in this notebook, but notes are for me. A skill is for my agents. The obvious first candidates are the books my sales work is built on, and the Kaltura brand and product guidelines, so a drafting agent answers from the source instead of from memory.

The update mode matters most: a skill that grows as I add material, instead of a one-off conversion. And the copyright rule is simple: skills from bought books stay private, on my machine.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.