⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

benchmark-lightflows

@google⭐ 12 stars

Run and interpret the unified A/B Lightflow benchmark (benchmarks/run_benchmark.py) comparing Arm A (Lightflow DAG) against Arm B (Traditional Skill) across 7 matched domain workflow pairs (3 to 22 stages).

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐ 12 on GitHubGitHub ↗

Growing

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

Lightflow Unified A/B Benchmark: Lightflow DAG (Arm A) vs. Traditional Skill (Arm B)

Executive Summary: The Goal Is Execution Consistency — And Token Efficiency Comes Mainly for Free

When evaluating Lightflow (Arm A) against its natural baseline — a Steelmanned Traditional Skill (Arm B: SKILL.md + actions.py) where the Python business logic is already factored out into actions.py — the conversation often starts with token count. However, evaluating both arms across 7 matched domain workflow pairs (3 to 22 stages, N=5 trials per workflow in run_benchmark.py) and 15 independent live zero-context LLM subagent sessions (N=5 per cohort = 30 live workflow runs) shows a deeper engineering conclusion:

The primary reason to use a declarative DAG is execution consistency and side-effect safety — and token efficiency (2.0x lower session tokens) comes along for free as a byproduct of moving orchestration out of the LLM's improvisation loop.

  1. Why Traditional Skills Drift Even When Python Work Is Factored Out: Even with domain logic cleanly factored into actions.py, a traditional SKILL.md still relies on the LLM to act as the runtime state machine across turns: writing ad-hoc python3 -c glue scripts, persisting intermediate outputs between turns, evaluating branch conditions, catching exceptions to trigger compensating rollbacks, and resuming from the exact point of failure without re-executing completed upstream stages.
    • Generative Glue Variance (even with 100% synchronized docs, Cohort 2): Across N=5 live zero-context LLM sessions given a meticulously synchronized SKILL.md, Arm B agents generated 3.1x more code (~1,565 ± 97 tok vs. ~502 ± 4 tok) with 23x higher code-generation SD (±389 chars vs. ±17 chars) and 38x higher warm runtime token SD (±144 tok vs. ±4 tok). Every run improvised a slightly different polling loop, state dictionary, and try/except wrapper.
    • Catastrophic Prose-Drift Fragility (1 omitted return-type detail → 50% first-try failure, Cohort 3): Because the prose SKILL.md is the orchestrator, omitting a single implementation detail (that actions.py helpers return a (output_dict, message_str) 2-tuple rather than a bare dict) caused 5/5 live Arm B subagents to crash mid-flight on Task 1 Stage 2 (TypeError: 'tuple' object is not subscriptable) after Stage 1 had already mutated disk state — forcing manual cleanup and re-execution of Stage 1 (5/10 first-try workflow compliance).
  2. Why Lightflow Gives Us Consistency "Mainly for Free" (Cohort 1): With Lightflow, the agent never reads lightflow.yaml or actions.py and never synthesizes Python orchestration glue. It loads one universal 1,097-token runner skill once per session and issues uniform declarative CLI commands (start → resume → resume):
    • Zero-Variance Execution (10/10 first-try pass, ±4 tok runtime SD): Across all N=5 live LLM sessions (10/10 workflow runs), Lightflow achieved 100% first-try compliance, 0 duplicate upstream executions during failure recovery, and near-zero runtime variance (7,389 ± 15 chars / ~1,847 ± 4 tok).
    • Consistency Comes Free — And Actually Cheaper (-49% to -50% Session Tokens): Instead of paying a token tax for deterministic state checkpointing (passport.json), JSON Schema gate validation, and automatic rollbacks, Lightflow cuts total session tokens in half (~2,944 ± 4 tok vs. ~5,904 ± 144 tok on the live 2-workflow suite; ~5,570 tok vs. ~10,992 tok across the 7-workflow ladder) because 1 universal runner skill replaces N verbose per-workflow SKILL.md manuals (-84% skill context) and declarative CLI calls replace inline Python glue (3.8x less generated code).
  3. Candid Crossover Point: On a single 3-stage linear workflow in a one-off session (hn_digest), a bespoke SKILL.md + /tmp/state.json is ~760 tokens lighter (~732 vs. ~1,495 tok). Lightflow breaks even by the 2nd–3rd workflow in a session or on any single workflow with >= 8 stages, conditional branches, polling loops, or rollbacks.

1. Experimental Design: Two Steelmanned Arms

To ensure neither arm is strawmanned:

  • Shared Python Actions (examples/<name>/actions.py): Both Arm A and Arm B call the exact same uninstrumented Python functions in examples/, verified via passport.json stamps, on-disk sandbox state, and a non-invasive sys.setprofile hook injected by benchmarks/run_benchmark.py when LIGHTFLOW_BENCH_TRACE_FILE is set.
  • Arm A (Lightflow DAG):
    • Reads the single, workflow-agnostic skill skills/run_lightflows/SKILL.md (4,388 chars / ~1,097 tokens) once per session and reuses it across all 7 workflows (0 additional skill tokens on workflows 2-7).
    • Invokes python3 -m lightflow start|resume|cleanup --lightflow=examples/<name> --log_id=<id>.
    • Persists all stage history and payload.outputs.<stage> on disk in passport.json.
  • Arm B (Steelmanned Traditional Skill: SKILL.md + actions.py):
    • Reads a dedicated, comprehensive per-workflow benchmarks/baseline_skills/<name>/SKILL.md (507 to 1,877 tokens each; 28,300 chars / ~7,075 tokens across all 7 workflows) that explicitly documents every function signature, return tuple ((output_dict, message_str)), payload["outputs"] key, branch condition, polling loop, approval gate, and rollback function so the agent never has to read actions.py.
    • Executes python3 -c snippets importing examples.<name>.actions and persists intermediate state to a local /tmp/.../state.json file between turns (steelmanned Arm B0 state file pattern, avoiding command-line JSON state-bus bloat).

2. Live Zero-Context LLM Subagent Benchmark (N=5 Trials per Cohort = 15 Sessions / 30 Workflow Runs)

To measure real-world consistency and variance (Mean ± SD) alongside prose-to-code drift sensitivity, we spawned 15 independent zero-context LLM subagent sessions (N=5 trials in each of 3 controlled cohorts) executing both new >5-stage failure-recovery workflows back-to-back in a single session (30 live workflow executions total):

  1. examples/blue_green_release (8 stages): Parallel validation diamond (run_security_scan ‖ run_integration_suite) → 2-tick warm_green_environment poll → approve_canary_cutover gate → shift_and_verify_canary SLO breach (exit 1) + automatic revert_canary_traffic rollback → surgical recovery (Stages 6-8 only) → promote_green_to_prod → decommission_old_slot.
  2. examples/incident_db_failover (9 stages): detect_primary_outage → elect_failover_candidate → mutually exclusive run_if WAL replay poll (replay_missing_wal_segments vs. skipped verify_zero_loss_sync) → trigger_rule: ALL_DONE fence_old_primary join → approve_replica_promotion IC gate → one-way promote_standby_replica (promotion_count == 1) → cutover_pooler_and_verify_writes canary write failure (exit 1) + automatic revert_pooler_to_maintenance rollback → surgical recovery (Stages 8-9 only, preserving promotion_count == 1) → publish_failover_ledger.

The 3 Controlled Cohorts (N=5 Sessions Each)

  • Cohort 1 — Arm A (Lightflow DAG, N=5): Reads skills/run_lightflows/SKILL.md (1,097 tok) once and runs lightflow start

// HOW IT'S BUILT

KEY FILES

benchmarks/SKILL.mdREADME.md

// REPO STATS

12 stars