⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
context-compression
Use Tokenless for large files and noisy tool outputs before they enter context. Must be used when commands, Read output, diffs, logs, or search results are large.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
// README
Quick start
npm install -g github:MaxForAI/Tokenless
tokenless repair-hooks --user
tokenless install-commands --user
tokenless launch
Then use Claude Code normally. Switch profiles anytime:
tokenless style chat # readable, shorter replies
tokenless style coding # dense coding output
tokenless style off # full hard-off
Claude Code gets expensive when every log, file read, diff, and long reply keeps getting carried into the next request.
Tokenless fixes that.
It keeps the raw evidence on your machine, sends Claude a compact version, and lets you expand the original output only when you need it.
Before / After
| Normal Claude Code | Claude Code with Tokenless |
|---|---|
| Reads a large file or log into future context repeatedly. | Stores the raw output locally and sends a compact packet. |
| Verbose final replies become part of the next request history. | chat and coding profiles keep replies short. |
| Agent trajectory can grow through repeated exploration and task-plan history. | Launcher trims Task/Plan tools by default; packets reduce large read context. |
Example large-read replacement:
| Raw context | Tokenless context |
|---|---|
| Full file/log output is carried through API requests. | TOKENLESS-READ-PACKET/0.1 with artifact id, imports, symbols, snippets, nearby files, and exact expansion commands. |
Why Tokenless
Claude Code sessions can become expensive because tool outputs, file reads, task-plan history, and verbose assistant replies are repeatedly carried through future API requests. Tokenless targets three sources of growth:
- Large tool output: test logs, build logs, search results, tree output, diffs, large reads, and large successful edit/write results.
- Agent trajectory overhead: repeated request context, high-overhead Task/Plan tools, and large raw file payloads.
- Response verbosity: optional
chatandcodingprofiles reduce assistant output tokens.
Benchmarks & Evidence
Tokenless has two evidence layers: real Claude Code API-body measurements, and external research showing why shorter, denser context can reduce cost without automatically reducing quality.
Real Claude Code benchmark runs
These are API-body measurements from actual Claude Code sessions. The main metric is estimated request-body or response-body tokens from raw API logs, not local hook-side savings estimates.
| Scenario | Baseline | Tokenless | Reduction |
|---|---|---|---|
5-turn CRM vibe coding, off vs coding | 4,697,867 request tokens | 2,476,391 | 47.3% |
6-turn natural conversation, off vs chat | 7,223 response tokens | 1,442 | 80.0% |
| Large CSS visual edit | 1,017,642 request tokens | 403,995-473,354 | ~54-60% |
| 10k-line React/TSX edit | 917,137 request tokens | 545,456 | 40.5% |
| Multifile React dashboard | 628,261 request tokens | 512,521 | 18.4% |
| Task/Plan tools enabled vs default launcher | 1,524,894 request tokens | 1,087,753 | 28.7% |
The strongest current product benchmark is the 5-turn CRM vibe-coding run: a non-specialist user gave vague iterative product-polish prompts. The public coding profile reduced request tokens by 47.3%, response tokens by 44.4%, and request count by 39.3% versus clean off.
The clean natural-conversation run isolates chat: no file tools or packet reducers were involved, and response tokens dropped by 80.0%.
Detailed methodology and raw run notes are in docs/benchmarking.md and docs/style-benchmark.md.
Research backing
The research does not prove Tokenless automatically helps every session. It supports the benchmark premise: context and response length are controllable engineering variables, and less text can sometimes be cheaper, faster, and more accurate.
| Paper | Why it matters for Tokenless |
|---|---|
| Brevity Constraints Reverse Performance Hierarchies in Language Models | Brevity constraints improved large-model accuracy by 26.3 percentage points on inverse-scaling problems. Verbose is not always better. |
| Prompt Compression in the Wild | Prompt compression can deliver real end-to-end speedups when workload, compression ratio, and hardware match; quality can remain statistically unchanged. |
| LLMLingua | Prompt compression can reduce inference cost while preserving semantic integrity under high compression ratios. |
| LongLLMLingua | Long-context compression can improve key-information perception while reducing cost and latency. |
| Selective Context | Pruning redundant context reported 50% context-cost reduction, 36% memory reduction, and 32% inference-time reduction with minor quality loss. |
| Gist Tokens | Learned prompt compression reached up to 26x prompt compression and up to 40% FLOPs reduction. |
Roadmap
Tokenless currently focuses on Claude Code context growth from tool output, file reads, and response verbosity. Next areas:
- User prompt compression: identify repeated prompt patterns, compress user intent without losing constraints, and keep the original prompt recoverable.
- Router-side optimization: reduce duplicated context and style overhead before requests hit the model backend.
- Broader workflow support: keep Claude Code as the primary target, then evaluate adapters for other agentic coding tools where the same context-growth problem appears.
Installation
Install from GitHub:
npm install -g github:MaxForAI/Tokenless
tokenless repair-hooks --user
tokenless launch
Tokenless is currently distributed through GitHub. It has not been published to the public npm registry yet.
For local development from a checkout:
git clone https://github.com/MaxForAI/Tokenless.git
cd Tokenless
npm install
npm link
tokenless repair-hooks --user
tokenless launch
If Claude Code is not available as claude on your PATH, set CLAUDE_BIN:
CLAUDE_BIN=/path/to/claude tokenless launch
Check installation status:
tokenless status --user
Output profiles
Tokenless has three public profiles:
| Profile | Behavior |
|---|---|
chat | Default. Short, readable natural-language responses. Only changes output style. |
coding | Dense structured responses for coding workflows. Only changes output style. |
off | Full Tokenless hard-off. Disables style injection and compression hooks. |
Set a profile:
tokenless style chat
tokenless style coding
tokenless style
// HOW IT'S BUILT
KEY FILES