⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

translate-book

@deusyu⭐ 2.0k stars

Translate books (PDF/DOCX/EPUB) into any language using parallel sub-agents. Converts input -> Markdown chunks -> translated chunks -> HTML/DOCX/EPUB/PDF.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐⭐ 2.0k on GitHubGitHub ↗

Very popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

Rainman Translate Book

English | 中文

An agent skill for Codex, Claude Code, and OpenClaw that translates entire books (PDF/DOCX/EPUB) into any language using parallel subagents.

Inspired by claude_translater. The original project uses shell scripts as its entry point, coordinating the Claude CLI with multiple step scripts to perform chunked translation. This project restructures the workflow as an agent skill for Codex, Claude Code, and OpenClaw, using subagents to translate chunks in parallel, with manifest-driven integrity checks, resumable runs, and multi-format output unified into a single pipeline. As the project structure and implementation differ significantly from the original, this is an independent project rather than a fork.


How It Works

Input (PDF/DOCX/EPUB)
  │
  ▼
Calibre ebook-convert → HTMLZ → HTML → Markdown
  (or Markdown input, e.g. from MinerU / Marker, skipping Calibre)
  │
  ▼
Split into chunks (chunk0001.md, chunk0002.md, ...)
  │  manifest.json tracks chunk hashes
  ▼
Parallel subagents (8 concurrent by default)
  │  each subagent: read 1 chunk → translate → write output_chunk*.md
  │  batched to respect API rate limits
  ▼
Validate (manifest hash check, 1:1 source↔output match)
  │
  ▼
Merge → Pandoc → HTML (with TOC) → Calibre → DOCX / EPUB / PDF

Each chunk gets its own independent subagent with a fresh context window. This prevents context accumulation and output truncation that happen when translating a full book in a single session.

Features

  • Parallel subagents — 8 concurrent translators per batch, each with isolated context
  • Resumable + selective re-translation — chunk-level resume, with run_state.json tracking glossary-sensitive re-translation
  • Neighbor context — each chunk can see short read-only excerpts from adjacent chunks for pronoun and entity resolution
  • Manifest validation — SHA-256 hash tracking prevents stale or corrupt outputs from being merged
  • Multi-format output — HTML (with floating TOC), DOCX, EPUB, PDF
  • Optional output controls — explicit EPUB cover, custom temp root, and user-facing export aliases
  • Multi-language — zh, en, ja, ko, fr, de, es (extensible)
  • PDF/DOCX/EPUB/Markdown input — Calibre handles the conversion heavy lifting; formula- and table-heavy PDFs can be pre-extracted to Markdown with MinerU or Marker

Prerequisites

  • Agent runtime — Codex, Claude Code, or OpenClaw, installed and ready to run skills
  • Calibre — ebook-convert command must be available (download)
  • Pandoc — for HTML↔Markdown conversion (download)
  • Python 3 with:
    • pypandoc — required (pip install pypandoc)
    • beautifulsoup4 — optional, for better TOC generation (pip install beautifulsoup4)
  • MinerU or Marker — optional, only to pre-extract formula- and table-heavy PDFs to Markdown (see Step 1)

Quick Start

1. Install the skill

Codex

npx skills add deusyu/translate-book -a codex -g

Or install it manually:

mkdir -p ~/.agents/skills
git clone https://github.com/deusyu/translate-book.git ~/.agents/skills/translate-book

Restart Codex if the newly installed skill does not appear.

Claude Code

npx skills add deusyu/translate-book -a claude-code -g

Or install it manually:

mkdir -p ~/.claude/skills
git clone https://github.com/deusyu/translate-book.git ~/.claude/skills/translate-book

OpenClaw

openclaw skills install @deusyu/translate-book

2. Translate a book

Codex

In the Codex CLI or IDE extension, enter:

$translate-book Translate /path/to/book.pdf into Chinese.

Codex can also select the skill automatically when your request matches its description.

Claude Code and OpenClaw

Ask the agent:

translate /path/to/book.pdf to Chinese

In Claude Code, you can also use the slash command:

/translate-book translate /path/to/book.pdf to Japanese

The skill handles the full pipeline automatically — convert, chunk, translate in parallel, validate, merge, and build all output formats.

3. Find your outputs

All files are in {book_name}_temp/:

FileDescription
output.mdMerged translated Markdown
book.htmlWeb version with floating TOC
book.docxWord document
book.epubE-book
book.pdfPrint-ready PDF

Repository Test Assets

  • Checked-in baseline inputs live under tests/baselines/<book-id>/.
  • Generated full-pipeline outputs live under tests/.artifacts/ and should not be committed.
  • Because scripts/convert.py writes {book_name}_temp/ under the current working directory, run repository baseline tests from inside tests/.artifacts/ to keep generated files out of the repo root.

Full-Pipeline Baseline Example

mkdir -p tests/.artifacts
cd tests/.artifacts
python3 ../../scripts/convert.py ../baselines/standard-alice/standard-alice.epub --olang zh
# then run translation via the skill
python3 ../../scripts/merge_and_build.py --temp-dir standard-alice_temp --title "test"

Feedback and Contributions

Please open a detailed GitHub issue instead of starting with a pull request. This project is maintained as an AI-assisted skill pipeline, and changes need to be evaluated against the current orchestration rules, chunk/manifest contracts, baseline assets, and release flow in one maintainer-owned context.

Pull requests are not the preferred contribution path and may be closed in favor of an issue. If you already have a patch, include the idea, key diff, failing case, or verification notes in the issue; the maintainer may rework or split the implementation before merging.

A useful issue should include:

  • Current behavior and expected behavior
  • Input format and environment, such as PDF/DOCX/EPUB, OS, Python, Calibre, and Pandoc versions
  • Minimal reproduction steps or a small public-domain sample when possible
  • Logs, screenshots, or generated file names that show the failure

Contact

For bugs and feature requests, please open a GitHub issue as described above. For questions, ideas, or anything else, you can reach me here:

ChannelLink
X (Twitter)@0xdeusyu
Telegram@DeusThink
Telegram Group (Chinese)@talkdeusyu
Telegram Channel (Chinese)@lovedesuyu
Email[email protected]

The channel is my running log of thinking alongside AI; the group is its discussion space, and translate-book questions are welcome there too.

Pipeline Details

Step 1: Convert

python3 scripts/convert.py /path/to/book.pdf --olang zh

Calibre converts the input to HTMLZ, which is extracted and converted to Markdown, then split into chunks (~6000 chars each). A manifest.json records the SHA-256 hash of each source chunk for later validation, and a source_fingerprint.json ties the temp dir to the exact source bytes it was built from — re-running against a replaced source file aborts instead of silently reusing stale chunks. Temp dirs created before fingerprinting are adopted with a warning on first re-run.

By default the working directory is {book_name}_temp/ under the current directory. Use --temp-root /path/to/work to keep the same leaf directory name under a different parent.

Formula- and table-heavy PDFs: convert Markdown instead (optional)

Calibre refl

// HOW IT'S BUILT

KEY FILES

SKILL.mdREADME.md

// REPO STATS

2.0k stars