⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
translate-book
Translate an arXiv paper, or any PDF/DOCX/EPUB book, into any language as a printable book. For arXiv papers it reads the LaTeX source, so equations, tables, figures and numbering survive. Parallel sub-agents translate the chunks; output is HTML, DOCX, EPUB and PDF.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
Translate Book: arXiv
An agent skill for Claude Code, Codex and OpenClaw that turns an arXiv paper into a translated, printable book. It reads the paper's LaTeX source, so equations, figures, tables and the paper's own numbering carry over. PDF, DOCX and EPUB books work too, through Calibre.
Left: page 7 of arXiv:2609.11801, Thinking with Looped Flows (CC BY 4.0). Right: page 12 of the Korean book this skill built from its LaTeX source. Every number in Table 1, and every ±, is the paper's. The whole book, 29 pages: looped-flows_ko.pdf. More pages, with source and licence: assets/demo.
npx skills add kcy4334-lgtm/translate-book-arxiv -a claude-code -g
Then ask: "translate /path/to/paper.pdf to Korean". The skill recognises an arXiv preprint from its first page and asks before it downloads the source. Other agents and a manual install are under Quick Start.
A PDF translator only gets what a PDF reader can recover, and equations do not survive that: pdftohtml breaks every formula into positioned text spans, and no option puts it back together. This skill works from the LaTeX the authors wrote.
The target language is a flag: zh, en, ja, ko, fr, de, es, and others. The print layout is measured against Korean, and ships with Korean typography tuned.
Forked from deusyu/translate-book, which grew out of claude_translater and built the agent-skill workflow: sub-agents translating chunks in parallel, manifest checks, resumable runs and several output formats from one pipeline. This fork develops the arXiv LaTeX path and the logs and advisors that ship with it.
How It Works
arXiv paper (PDF) │ any other book (PDF/DOCX/EPUB)
│ detected from the page-1 stamp │
▼ ▼
Fetch /e-print → flatten the LaTeX Calibre ebook-convert → HTMLZ → HTML
│ the paper's own macros resolved from the .sty files it ships
│ equation, theorem, section and float numbers read from the source
│ figures rasterised from the original vector PDFs, captions attached
▼ ▼
Markdown, with $...$ math intact ←───┘
│
▼
Split into chunks (chunk0001.md, chunk0002.md, ...)
│ manifest.json tracks chunk hashes
│ the reference list becomes its own chunk and is copied, not translated
▼
Parallel sub-agents (work queue, 8 in flight by default)
│ each sub-agent: read 1 chunk → translate → write output_chunk*.md
│ each chunk is verified, recorded and merged as it lands
▼
Validate (manifest hash check, 1:1 source↔output match)
│
▼
Merge → Pandoc → HTML (with TOC) → Pandoc DOCX / Calibre EPUB / Chromium PDF
Each chunk goes to a sub-agent with a fresh context, so a long book never fills one session and nothing gets cut off at the end.
What it prints
The same page in each language: numbered sections, a display equation, citations, a results table and a figure. Click any page for the full-size render.
The float label follows the language (Figure 1, 그림 1 (Fig. 1), 図 1). Table headers and method names are translated; numbers, units and citations are not. The body font is chosen per script, and line breaking follows each language's rules.
This page was written for the repository, with made-up results, so it can be shown in seven languages without depending on any paper's licence. Its source is tests/fixtures/sample_page.md with one translation per language beside it, and python tests/sample_pages.py renders every image again with the shipped a4-book profile.
Features
- Numbers come from the paper: equation, theorem, section, float and subfigure numbers are read from the LaTeX source and checked against the original PDF by
tests/source_probe.py - The paper's own macros are resolved: a paper's
.styis never\input, so pandoc would print\ieor\parheadas they are.scripts/paper_macros.pyexpands the definitions the paper ships: 4,099 calls across 21 papers, with nothing lost but the macro names on the six papers diffed word by word. When it cannot expand one safely it refuses and names the macro and the reason - The reference list is copied as is: it becomes its own chunk, usually the largest, at 27–34% of a paper's characters
- Parallel sub-agents: a work queue keeps 8 translators running, each with its own context; the next chunk starts as soon as a slot frees
- Every chunk is checked before it counts:
scripts/verify_chunk.pycompares each output with its source, the glossary it was given and the chunk it quoted - Consistent terms: a glossary built before translation, a per-chunk term table, and short read-only excerpts from the neighbouring chunks for names and pronouns
- Resumable: SHA-256 hashes in a manifest keep stale outputs out of the merge, and a changed glossary re-translates only the chunks that used the changed terms
- Print-ready PDF: headless Chromium against a real
@pagebox (A4, 18/18/22/18 mm, 11.5 pt), page numbers stamped afterwards because Chrome has no margin boxes.scripts/layout.pyholds the page geometry and fonts - Output: HTML with a floating TOC, DOCX, EPUB and PDF, with an optional EPUB cover, working folder and export name
- Tests: 2,275, standard library only, run in CI
Growing the skill
Each paper brings LaTeX constructs the last one did not have. Four stores ship with the skill so a construct met once is handled the next time:
| what it holds | |
|---|---|
KNOWLEDGE.md | What a tool actually did, with the measurement that proved it. 217 entries |
KNOWHOW.md | What a way of working cost, so it is not paid twice. 44 entries |
REFEREE.md | Whether a repeated failure belongs to a tool, a briefing or a role. 6 entries |
corpus/shapes.json | Every LaTeX construct each paper carried, written by the build itself. 28 papers |
The census answers "has this ever been seen?" with a count, and lists what has never been seen, so a pattern that has never met a real example is not trusted. tests/test_source_lint.py fails when the corpus has met a construct nobody has classified, so a new construct is dealt with before release.
Four advisor sub-agents read the stores:
- old-man: before concluding a paper does not contain something, or writing a pattern whose match decides it; names the spellings and layouts the pattern would miss
- question-monster: after concluding something is impossible; hands back candidates to test
- fast-finder: instead of reading the logs; returns the few entries that bear on the question
- referee: once a whole run is gated; tells a tool fault from a briefing fault from a role's
An example
From one working session:
\iewas printing mid-sentence in a finished Korean book, five
// HOW IT'S BUILT
KEY FILES