⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

context-compression

@maxforai⭐ 46 stars

Use Tokenless for large files and noisy tool outputs before they enter context. Must be used when commands, Read output, diffs, logs, or search results are large.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐ 46 on GitHubGitHub ↗

Growing

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README


Quick start

npm install -g github:MaxForAI/Tokenless
tokenless repair-hooks --user
tokenless install-commands --user
tokenless launch

Then use Claude Code normally. Switch profiles anytime:

tokenless style chat     # readable, shorter replies
tokenless style coding   # dense coding output
tokenless style off      # full hard-off

Claude Code gets expensive when every log, file read, diff, and long reply keeps getting carried into the next request.

Tokenless fixes that.

It keeps the raw evidence on your machine, sends Claude a compact version, and lets you expand the original output only when you need it.

Before / After

Normal Claude CodeClaude Code with Tokenless
Reads a large file or log into future context repeatedly.Stores the raw output locally and sends a compact packet.
Verbose final replies become part of the next request history.chat and coding profiles keep replies short.
Agent trajectory can grow through repeated exploration and task-plan history.Launcher trims Task/Plan tools by default; packets reduce large read context.

Example large-read replacement:

Raw contextTokenless context
Full file/log output is carried through API requests.TOKENLESS-READ-PACKET/0.1 with artifact id, imports, symbols, snippets, nearby files, and exact expansion commands.

Why Tokenless

Claude Code sessions can become expensive because tool outputs, file reads, task-plan history, and verbose assistant replies are repeatedly carried through future API requests. Tokenless targets three sources of growth:

  • Large tool output: test logs, build logs, search results, tree output, diffs, large reads, and large successful edit/write results.
  • Agent trajectory overhead: repeated request context, high-overhead Task/Plan tools, and large raw file payloads.
  • Response verbosity: optional chat and coding profiles reduce assistant output tokens.

Benchmarks & Evidence

Tokenless has two evidence layers: real Claude Code API-body measurements, and external research showing why shorter, denser context can reduce cost without automatically reducing quality.

Real Claude Code benchmark runs

These are API-body measurements from actual Claude Code sessions. The main metric is estimated request-body or response-body tokens from raw API logs, not local hook-side savings estimates.

ScenarioBaselineTokenlessReduction
5-turn CRM vibe coding, off vs coding4,697,867 request tokens2,476,39147.3%
6-turn natural conversation, off vs chat7,223 response tokens1,44280.0%
Large CSS visual edit1,017,642 request tokens403,995-473,354~54-60%
10k-line React/TSX edit917,137 request tokens545,45640.5%
Multifile React dashboard628,261 request tokens512,52118.4%
Task/Plan tools enabled vs default launcher1,524,894 request tokens1,087,75328.7%

The strongest current product benchmark is the 5-turn CRM vibe-coding run: a non-specialist user gave vague iterative product-polish prompts. The public coding profile reduced request tokens by 47.3%, response tokens by 44.4%, and request count by 39.3% versus clean off.

The clean natural-conversation run isolates chat: no file tools or packet reducers were involved, and response tokens dropped by 80.0%.

Detailed methodology and raw run notes are in docs/benchmarking.md and docs/style-benchmark.md.

Research backing

The research does not prove Tokenless automatically helps every session. It supports the benchmark premise: context and response length are controllable engineering variables, and less text can sometimes be cheaper, faster, and more accurate.

PaperWhy it matters for Tokenless
Brevity Constraints Reverse Performance Hierarchies in Language ModelsBrevity constraints improved large-model accuracy by 26.3 percentage points on inverse-scaling problems. Verbose is not always better.
Prompt Compression in the WildPrompt compression can deliver real end-to-end speedups when workload, compression ratio, and hardware match; quality can remain statistically unchanged.
LLMLinguaPrompt compression can reduce inference cost while preserving semantic integrity under high compression ratios.
LongLLMLinguaLong-context compression can improve key-information perception while reducing cost and latency.
Selective ContextPruning redundant context reported 50% context-cost reduction, 36% memory reduction, and 32% inference-time reduction with minor quality loss.
Gist TokensLearned prompt compression reached up to 26x prompt compression and up to 40% FLOPs reduction.

Roadmap

Tokenless currently focuses on Claude Code context growth from tool output, file reads, and response verbosity. Next areas:

  • User prompt compression: identify repeated prompt patterns, compress user intent without losing constraints, and keep the original prompt recoverable.
  • Router-side optimization: reduce duplicated context and style overhead before requests hit the model backend.
  • Broader workflow support: keep Claude Code as the primary target, then evaluate adapters for other agentic coding tools where the same context-growth problem appears.

Installation

Install from GitHub:

npm install -g github:MaxForAI/Tokenless
tokenless repair-hooks --user
tokenless launch

Tokenless is currently distributed through GitHub. It has not been published to the public npm registry yet.

For local development from a checkout:

git clone https://github.com/MaxForAI/Tokenless.git
cd Tokenless
npm install
npm link
tokenless repair-hooks --user
tokenless launch

If Claude Code is not available as claude on your PATH, set CLAUDE_BIN:

CLAUDE_BIN=/path/to/claude tokenless launch

Check installation status:

tokenless status --user

Output profiles

Tokenless has three public profiles:

ProfileBehavior
chatDefault. Short, readable natural-language responses. Only changes output style.
codingDense structured responses for coding workflows. Only changes output style.
offFull Tokenless hard-off. Disables style injection and compression hooks.

Set a profile:

tokenless style chat
tokenless style coding
tokenless style 

// HOW IT'S BUILT

KEY FILES

plugins/claude-code/skills/context-compression/SKILL.mdREADME.md

// REPO STATS

46 stars