⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

v1.0.0

monte-carlo-analyze-root-cause

@monte-carlo-data⭐ 94 stars

|

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐ 94 on GitHubGitHub ↗

Growing

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

Analyze Root Cause Skill

Investigate data incidents and find root causes using Monte Carlo's observability data. Guides the agent through systematic investigation: alert lookup, lineage tracing, ETL checks, query analysis, and data profiling.

What it does

  • Investigates freshness delays, volume anomalies, schema changes, ETL failures, query regressions, and field metric drift
  • Maps blast radius using table and field-level lineage
  • Traces bad data upstream to find the source
  • Correlates changes (query modifications, volume shifts, ETL failures) with incident timeline
  • Profiles actual data when a database MCP connector is available
  • Matches findings against a catalog of known root cause patterns

MCP Tools Required

Connect to Monte Carlo's MCP server (integrations.getmontecarlo.com/mcp). The skill uses these tools:

ToolPurpose
get_alertsFetch incident/alert details
searchFind tables by name
get_tableTable metadata and fields
get_asset_lineageTable-level lineage
get_field_lineageField-level lineage (trace to source column)
get_table_freshnessUpdate/freshness history
get_table_size_historyRow count and size history
get_queries_for_tableRead/write query history
get_query_changesDetect SQL text modifications
get_query_rcaFailed/futile/missed query analysis
get_change_timelineUnified change timeline
get_etl_issuesETL pipeline issues (Airflow, dbt, Databricks) — pass platform param
get_etl_jobsFind ETL jobs writing to tables (Airflow, dbt, Databricks) — pass platform param
get_github_prsRecent GitHub PRs (via MC's GitHub integration)
get_jobs_performanceJob runtime stats, failure rates, trends
alert_assessmentOptional ~2-min triage of an incident (HIGH/MEDIUM/LOW confidence + impact)
run_troubleshooting_agentStarts the Troubleshooting Agent (TSA) on an incident; auto-invoked when an incident UUID is present
get_troubleshooting_agent_resultsPolls TSA results for an incident

Credits: alert_assessment and run_troubleshooting_agent consume Monte Carlo credits the same way the Troubleshooting Agent does when launched from the Monte Carlo UI.

Optional: A database MCP server (Snowflake, BigQuery, Redshift) for direct SQL queries.

Example prompts

  • "Investigate alert 12345"
  • "Why is the orders table stale?"
  • "Row count dropped 50% on analytics.prod.revenue — what happened?"
  • "Debug this freshness issue on our daily pipeline"
  • "The dashboard shows yesterday's data — can you find out why?"

Investigation flow

Intake (alert ID or user description)
    ↓
Auto-invoke TSA (if incident UUID + not opt-out + not narrow check)  ─┐
    ↓                                                                  │
Map blast radius (upstream + downstream lineage)                       │ TSA runs
    ↓                                                                  │ async in
Investigate by issue type (freshness / volume / schema / ETL / query)  │ parallel
    ↓                                                                  │
Check upstream causes (walk lineage chain)  ── poll TSA #1 ────────────┤
    ↓                                                                  │
Profile data (if DB connector available)                               │
    ↓                                                                  │
Check code changes (GitHub MCP or MC query changes)                    │
    ↓                                                                  │
Synthesize: root cause + evidence + impact + fix  ── poll TSA #2 ─────┘
                                                    + merge findings

When intake has no incident UUID, when the user explicitly opts out, or when the request is a narrow scoped check (e.g. "is X stale right now?"), TSA is skipped and the manual flow runs alone.

Reference files

FileDescription
references/freshness-investigation.mdFreshness delay playbook
references/volume-investigation.mdVolume anomaly playbook
references/schema-investigation.mdSchema change playbook
references/etl-failure-investigation.mdETL failure playbook
references/query-change-investigation.mdQuery modification playbook
references/field-anomaly-investigation.mdField metric drift playbook
references/data-exploration.mdSQL patterns for data profiling
references/intake-no-incident.mdIntake flow when no incident ID
references/common-root-causes.mdCatalog of known root cause patterns

// HOW IT'S BUILT

KEY FILES

skills/analyze-root-cause/SKILL.mdREADME.md

// REPO STATS

94 stars