⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
benchmark-lightflows
Run and interpret the unified A/B Lightflow benchmark (benchmarks/run_benchmark.py) comparing Arm A (Lightflow DAG) against Arm B (Traditional Skill) across 7 matched domain workflow pairs (3 to 22 stages).
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
// README
Lightflow Unified A/B Benchmark: Lightflow DAG (Arm A) vs. Traditional Skill (Arm B)
Executive Summary: The Goal Is Execution Consistency — And Token Efficiency Comes Mainly for Free
When evaluating Lightflow (Arm A) against its natural baseline — a
Steelmanned Traditional Skill (Arm B: SKILL.md + actions.py) where the
Python business logic is already factored out into actions.py — the
conversation often starts with token count. However, evaluating both arms across
7 matched domain workflow pairs (3 to 22 stages, N=5 trials per
workflow in run_benchmark.py) and 15 independent live zero-context LLM
subagent sessions (N=5 per cohort = 30 live workflow runs) shows a deeper
engineering conclusion:
The primary reason to use a declarative DAG is execution consistency and side-effect safety — and token efficiency (
2.0xlower session tokens) comes along for free as a byproduct of moving orchestration out of the LLM's improvisation loop.
- Why Traditional Skills Drift Even When Python Work Is Factored Out: Even
with domain logic cleanly factored into
actions.py, a traditionalSKILL.mdstill relies on the LLM to act as the runtime state machine across turns: writing ad-hocpython3 -cglue scripts, persisting intermediate outputs between turns, evaluating branch conditions, catching exceptions to trigger compensating rollbacks, and resuming from the exact point of failure without re-executing completed upstream stages.- Generative Glue Variance (even with 100% synchronized docs,
Cohort 2): AcrossN=5live zero-context LLM sessions given a meticulously synchronizedSKILL.md, Arm B agents generated3.1xmore code (~1,565 ± 97 tokvs.~502 ± 4 tok) with23xhigher code-generation SD (±389 charsvs.±17 chars) and38xhigher warm runtime token SD (±144 tokvs.±4 tok). Every run improvised a slightly different polling loop, state dictionary, andtry/exceptwrapper. - Catastrophic Prose-Drift Fragility (1 omitted return-type detail →
50%first-try failure,Cohort 3): Because the proseSKILL.mdis the orchestrator, omitting a single implementation detail (thatactions.pyhelpers return a(output_dict, message_str)2-tuple rather than a baredict) caused5/5live Arm B subagents to crash mid-flight on Task 1 Stage 2 (TypeError: 'tuple' object is not subscriptable) after Stage 1 had already mutated disk state — forcing manual cleanup and re-execution of Stage 1 (5/10first-try workflow compliance).
- Generative Glue Variance (even with 100% synchronized docs,
- Why Lightflow Gives Us Consistency "Mainly for Free" (
Cohort 1): With Lightflow, the agent never readslightflow.yamloractions.pyand never synthesizes Python orchestration glue. It loads one universal1,097-token runner skill once per session and issues uniform declarative CLI commands (start→resume→resume):- Zero-Variance Execution (
10/10first-try pass,±4 tokruntime SD): Across allN=5live LLM sessions (10/10workflow runs), Lightflow achieved100%first-try compliance,0duplicate upstream executions during failure recovery, and near-zero runtime variance (7,389 ± 15 chars/~1,847 ± 4 tok). - Consistency Comes Free — And Actually Cheaper (
-49%to-50%Session Tokens): Instead of paying a token tax for deterministic state checkpointing (passport.json), JSON Schema gate validation, and automatic rollbacks, Lightflow cuts total session tokens in half (~2,944 ± 4 tokvs.~5,904 ± 144 tokon the live 2-workflow suite;~5,570 tokvs.~10,992 tokacross the 7-workflow ladder) because1universal runner skill replacesNverbose per-workflowSKILL.mdmanuals (-84%skill context) and declarative CLI calls replace inline Python glue (3.8xless generated code).
- Zero-Variance Execution (
- Candid Crossover Point: On a single 3-stage linear workflow in a one-off
session (
hn_digest), a bespokeSKILL.md+/tmp/state.jsonis ~760 tokens lighter (~732vs.~1,495 tok). Lightflow breaks even by the 2nd–3rd workflow in a session or on any single workflow with>= 8stages, conditional branches, polling loops, or rollbacks.
1. Experimental Design: Two Steelmanned Arms
To ensure neither arm is strawmanned:
- Shared Python Actions (
examples/<name>/actions.py): Both Arm A and Arm B call the exact same uninstrumented Python functions inexamples/, verified viapassport.jsonstamps, on-disk sandbox state, and a non-invasivesys.setprofilehook injected bybenchmarks/run_benchmark.pywhenLIGHTFLOW_BENCH_TRACE_FILEis set. - Arm A (
Lightflow DAG):- Reads the single, workflow-agnostic skill
skills/run_lightflows/SKILL.md(4,388 chars/~1,097 tokens) once per session and reuses it across all 7 workflows (0additional skill tokens on workflows 2-7). - Invokes
python3 -m lightflow start|resume|cleanup --lightflow=examples/<name> --log_id=<id>. - Persists all stage history and
payload.outputs.<stage>on disk inpassport.json.
- Reads the single, workflow-agnostic skill
- Arm B (
Steelmanned Traditional Skill:SKILL.md+actions.py):- Reads a dedicated, comprehensive per-workflow
benchmarks/baseline_skills/<name>/SKILL.md(507to1,877tokens each;28,300 chars/~7,075 tokensacross all 7 workflows) that explicitly documents every function signature, return tuple ((output_dict, message_str)),payload["outputs"]key, branch condition, polling loop, approval gate, and rollback function so the agent never has to readactions.py. - Executes
python3 -csnippets importingexamples.<name>.actionsand persists intermediate state to a local/tmp/.../state.jsonfile between turns (steelmannedArm B0state file pattern, avoiding command-line JSON state-bus bloat).
- Reads a dedicated, comprehensive per-workflow
2. Live Zero-Context LLM Subagent Benchmark (N=5 Trials per Cohort = 15 Sessions / 30 Workflow Runs)
To measure real-world consistency and variance (Mean ± SD) alongside
prose-to-code drift sensitivity, we spawned 15 independent zero-context
LLM subagent sessions (N=5 trials in each of 3 controlled cohorts) executing
both new >5-stage failure-recovery workflows back-to-back in a single session
(30 live workflow executions total):
examples/blue_green_release(8 stages): Parallel validation diamond (run_security_scan‖run_integration_suite) → 2-tickwarm_green_environmentpoll →approve_canary_cutovergate →shift_and_verify_canarySLO breach (exit 1) + automaticrevert_canary_trafficrollback → surgical recovery (Stages 6-8only) →promote_green_to_prod→decommission_old_slot.examples/incident_db_failover(9 stages):detect_primary_outage→elect_failover_candidate→ mutually exclusiverun_ifWAL replay poll (replay_missing_wal_segmentsvs. skippedverify_zero_loss_sync) →trigger_rule: ALL_DONEfence_old_primaryjoin →approve_replica_promotionIC gate → one-waypromote_standby_replica(promotion_count == 1) →cutover_pooler_and_verify_writescanary write failure (exit 1) + automaticrevert_pooler_to_maintenancerollback → surgical recovery (Stages 8-9only, preservingpromotion_count == 1) →publish_failover_ledger.
The 3 Controlled Cohorts (N=5 Sessions Each)
- Cohort 1 — Arm A (
Lightflow DAG,N=5): Readsskills/run_lightflows/SKILL.md(1,097 tok) once and runslightflow start
// HOW IT'S BUILT
KEY FILES