⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
adversarial-claims-reviewer
Use when adversarially reviewing a document that makes formal or technical claims — math derivations, physics papers, statistical analyses, benchmark reports, whitepapers. Inventories every equation and quantitative claim, verifies each AS NAMED in the text (never a paraphrase or a neighboring statement), and classifies VERIFIED / REFUTED / UNVERIFIABLE / VACUOUS. Triggers on "check this paper", "verify these claims", "is this derivation right", "review this proof", "audit this benchmark", "does the math hold up". For source-code review see code-review-and-quality; for skill/agent library audits see skill-library-review; for content quality scoring see content-ops.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
// README
Engineering Heresy — Agentic Framework
A collection of skills and agents for Claude Code that encode engineering workflows, content pipelines, game development, marketing ops, and more into reusable AI playbooks.
Install once, use in any project.
Part of Engineering Heresy by Glenn Eggleton — challenging conventional wisdom in AI and software engineering. Subscribe on Substack →
Features
AgenticOS (this repository) is a curated library of AI playbooks for serious, long-running work in software engineering, game development, marketing, and operations. You install it once into Claude Code; from then on, your agent can follow proven workflows instead of improvising every task from scratch.
The product is not "more agents for the sake of agents." It is a harness — structure, guardrails, and reusable expertise wrapped around frontier models so you get predictable quality at lower token cost, especially across multi-step sessions where context compression would otherwise cause drift and rework. See NORTH_STAR.md for the design thesis.
What you get
| Piece | What it is | How you use it |
|---|---|---|
| Skills (38) | Step-by-step playbooks for a kind of work — code review, Rust engineering, SEO ops, game balancing, etc. | Ask the agent to use a skill by name, or configure your IDE so skills are checked before every task (see Usage). |
| Agents (16) | Role definitions with a mandate and tool allowlist — engineer, security-reviewer, marketer, rust-engineer, etc. | Spawn explicitly ("use the code-reviewer agent") or let the orchestrator dispatch subagents for multi-step work. |
| Commands (3 ship to consumers) | Slash shortcuts — scaffold new skills/agents, record session facts. | /skill-new, /agent-new, /state. |
| Hooks | Small shell scripts that run on IDE events (session start, before shell, before compaction). | Installed and registered automatically; power the awareness harness. Disable by editing your global hook config. |
| Operating rules | Always-on doctrine for orchestration, memory, grounding, and review tiers. | Clone this repo into a project to use the flat CLAUDE.md (generated from .claude/rules/ by scripts/build-claude-md.sh). Not copied by the global installer. |
Supported platform
- Claude Code — full install: skills, agents, hooks, and three consumer commands. Remote one-liner installs a pinned, SHA-256–verified release (
install.sh/install.ps1).
Workflow domains
Skills and agents are grouped by the work they cover. Invoke the one that matches your task; shaper skills (prompt-shaper, marketing-shaper, game-design-shaper) turn vague requests into scoped briefs first.
Software engineering
- Full-stack implementation (
engineer,devops-engineer) - Language specialists: Rust (
rust-engineer), TypeScript testing (frontend/backend), data pipelines and analytics - Smart contracts and EVM development (
web3-engineer,web3-smart-contract-engineering) - CI/CD and deployment pipelines
- Planning and task breakdown for parallel subagent dispatch
- Release coordination across a monorepo
- Codebase cost estimation (LOC/complexity → build cost)
Quality, security, and review
- Multi-axis code review (
code-reviewer,code-review-and-quality) - Cross-stack security audit (
security-reviewer,security-engineering) - PII scan and redaction (
security) - Adversarial review of formal claims — math, stats, benchmarks (
adversarial-claims-reviewer) - API/persistence cataloging into
DATA_MODEL.md(data-model-documenter,data-model-documentation) with adversarial verification (data-model-verifier)
Game development
- Godot 4 + C# (
godot-engineer) - Phaser 3 + TypeScript (
phaser-engineer) - Systems design, economy balancing, IAP catalog design (
game-systems-designer,game-balancer,iap-manager) - End-to-end game design pipeline (
game-design-shaper)
Marketing, growth, and revenue
- Marketing intake and full-spectrum execution (
marketing-shaper,marketer) - Content scoring with an expert panel (
content-ops) - Content production pipeline — quotes, clips, repurposing (
content-pipeline) - SEO, CRO, growth experiments, cold outbound, revenue attribution (
seo-ops,conversion-ops,growth-engine,outbound-engine,revenue-intelligence) - Karpathy-style autoresearch on conversion content (
autoresearch)
Browser and tooling
- Real-browser testing via Chrome DevTools MCP (
browser-testing-with-devtools)
Library maintenance (for contributors and fork maintainers)
- Scaffold conforming skills and agents (
/skill-new,/agent-new) - Structural audit of the library (
library-reviewer,library-investigator,skill-library-review) - Sharded full-library audit workflow (
/audit-library— repo-local) - Stochastic finding triage and ratchet promotion (
findings-ledger,/triage-findings)
Harness capabilities
These are features of the framework itself, not individual skills:
Awareness harness (experimental) — Fights the dominant failure mode of long agent sessions: losing track of settled decisions and existing infrastructure. Externalizes live session state in SESSION-STATE.md, re-injects it via hooks at session start and each turn, checkpoints before compaction, and nudges "survey before you provision." Includes deterministic metrics to compare hook-ON vs hook-OFF sessions.
Orchestrator + ship gates — Operating doctrine treats the main agent as an orchestrator that dispatches specialists (the Agent tool) instead of doing multi-step work inline. After implementation, a fixed review DAG runs: code review, security review, optional library review, and conditional data-model documentation/verification. PR checkboxes and CI (check-pr-ship-gates) enforce the gate. Canonical graph: gate-dag.md.
Review tiers — Findings are sorted by reproducibility: Tier 0 deterministic checks hard-block; Tier 1 LLM findings need an evidence artifact; Tier 2 is advisory and logged to the findings ledger for recurrence-based promotion into deterministic checks.
Persistent memory — Two layers on Claude Code: within-session facts in SESSION-STATE.md and cross-session memory in .claude/memory/. The Codex install keeps the portable session-state layer but leaves cross-session recall to Codex's native memory system. Never hand-edit session state — use /state on Claude or invoke session-state on Codex.
Deterministic validation — scripts/validate.sh is an LLM-free gate on every PR: frontmatter, naming, dangling links, ship manifest, hook safety, tombstones for removed skills, and more. Install scripts refuse to copy a library that fails validation. Enable .githooks pre-commit for local enforcement.
Pinned, verified releases — Remote installs download a tagged tarball and abort on SHA-256 mismatch. No "track main" remote path — unreleased work installs from a local clone. Maintainer runbook: RELEASING.md.
Telemetry (opt-in) — Local-first usage telemetry skill; privacy-respecting, not shipped as silent analytics.
Security model for hooks — Shipped hook scripts are auditable shell, no runtime network, human-reviewed before merge, and scanned by validator Invariant 8. Injected session files are treated as untrusted data. Details: SECURITY.md.
How to get started
- Install — choose Claude Code or Codex.
- Configure skill discipline — use
CLAUDE.mdon Claude orAGENTS.mdon Codex for durable repository guidance. - Initialize session state (optional for long sessions) —
/state initon Claude, or ask Codex to usesession-stateand initialize it. - **Pick a workf
// HOW IT'S BUILT
KEY FILES