<< All versions

Skill v1.0.0

currentAutomated scan96/100
krzemienski/validationforge/retrospective-validation
──Details
PublishedSeptember 27, 2026 at 09:58 PM
Content Hashsha256:230bf6f1b5626659...
Git SHAed1270bc1cbe
──Files
Files (1 file, 7.3 KB)
SKILL.md7.3 KBactive
SKILL.md · 217 lines · 7.3 KB

version: "1.0.0" name: retrospective-validation description: "Use to evaluate whether your validation process itself is working — not whether a single run passed, but whether the team's validation discipline is catching real bugs before they ship. Analyzes past validation results, deployments, and incidents to compute false PASS rate (validations that said PASS but bug shipped), false FAIL rate (validations that said FAIL but nothing was actually broken), revert frequency, and confidence score. Outputs specific process change recommendations. Reach for it on phrases like 'did our validation approach work', 'retrospective on process', 'quality of our validation', 'post-incident review of QA', or quarterly when reviewing QA discipline." triggers:

  • "retrospective validation"
  • "validate methodology"
  • "historical validation"
  • "did our approach work"
  • "post-mortem analysis"
  • "quality of our validation"
  • "false pass rate"
  • "qa retrospective"

context_priority: reference


Retrospective Validation

Validate whether a methodology, process, or technical approach actually worked by analyzing historical evidence. Uses past validation results, deployment outcomes, and incident history to assess effectiveness.

When to Use

  • After a sprint/milestone to evaluate the validation approach used
  • When deciding whether to continue or change a methodology
  • When comparing two approaches (A/B validation strategy)
  • Post-incident to assess if validation should have caught the issue
  • When building a case for adopting or abandoning a practice

Four-Phase Process

Phase 1: COLLECT Phase 2: ANALYZE Phase 3: CORRELATE Phase 4: CONCLUDE
Historical Categorize & Map causes to Confidence score
evidence quantify outcomes effects & recommendations

Phase 1: Collect Historical Evidence

Gather all available evidence from past validation runs, deployments, and incidents.

Evidence Sources

SourceWhat to CollectLocation
Past validation reportsPASS/FAIL verdicts, evidence qualitye2e-evidence/*/report.md
Git historyDeploy frequency, revert frequencygit log --oneline
Incident reportsProduction issues post-deployIssue tracker, post-mortems
Build logsBuild failure rate over timeCI/CD logs
User feedbackBug reports, complaintsIssue tracker
bash
mkdir -p e2e-evidence/retrospective
# Collect past validation reports
find e2e-evidence -name "report.md" -not -path "*/retrospective/*" \
| sort | while read f; do
echo "=== $f ===" >> e2e-evidence/retrospective/step-01-past-reports.txt
head -20 "$f" >> e2e-evidence/retrospective/step-01-past-reports.txt
echo "" >> e2e-evidence/retrospective/step-01-past-reports.txt
done
# Collect deploy history
git log --oneline --since="30 days ago" --grep="deploy\|release\|revert" \
> e2e-evidence/retrospective/step-01-deploy-history.txt
# Collect revert rate
TOTAL_DEPLOYS=$(git log --oneline --since="30 days ago" --grep="deploy\|release" | wc -l)
REVERTS=$(git log --oneline --since="30 days ago" --grep="revert" | wc -l)
echo "Deploys: $TOTAL_DEPLOYS, Reverts: $REVERTS" \
> e2e-evidence/retrospective/step-01-revert-rate.txt

Phase 2: Analyze Outcomes

Categorize and quantify the collected evidence.

Metrics to Calculate

MetricFormulaGoodConcerning
Validation PASS ratePASS journeys / total journeys>90%<80%
False PASS rateProduction bugs that validation missed / total PASSes<5%>10%
False FAIL rateInvestigations that found no real bug / total FAILs<10%>20%
Revert rateReverts / deploys<5%>10%
Mean time to detectTime from deploy to bug detection<1h>24h
Evidence qualityReports with cited evidence / total reports>95%<80%

Analysis Template

markdown
## Outcome Analysis
**Period:** YYYY-MM-DD to YYYY-MM-DD
**Total validation runs:** N
**Total deploys:** N
### Validation Effectiveness
-PASS rate: N%
-FAIL rate: N%
-False PASS (bugs in prod that validation missed): N
-False FAIL (unnecessary investigations): N
### Deploy Quality
-Deploys: N
-Reverts: N (N%)
-Production incidents: N
-Mean time to detect: Xh

Save to e2e-evidence/retrospective/step-02-analysis.md.

Phase 3: Correlate Causes and Effects

Map validation practices to outcomes.

Correlation Matrix

For each production incident or revert, trace back:

markdown
## Incident Correlation
### Incident: {description}
**Date:** YYYY-MM-DD
**Severity:** CRITICAL/HIGH/MEDIUM/LOW
**Root cause:** {technical cause}
**Validation gap analysis:**
-Was this feature validated before deploy? YES/NO
-If YES, what type of validation? {build gates only / e2e / visual / etc.}
-If YES, why didn't validation catch it? {missing journey / wrong criteria / flaky flow / etc.}
-If NO, why was validation skipped? {time pressure / oversight / no plan / etc.}
**Would ValidationForge have caught it?**
-Which skill would apply? {skill name}
-What journey would detect it? {journey description}
-Confidence: HIGH/MEDIUM/LOW

Pattern Detection

Look for patterns across multiple incidents:

PatternIndicates
Most incidents in UI renderingVisual inspection is insufficient
Most incidents in API integrationIntegration validation needs strengthening
Most incidents after "urgent" deploysProcess shortcuts are the problem
Low false FAIL rate but high false PASSValidation criteria are too lenient
High false FAIL rateValidation is too brittle (flaky flows)

Phase 4: Conclude

Confidence Formula

Confidence = Detection × Accuracy × Longevity
Where:
Detection (D) = 1 - (false_pass_rate) # How often we catch real bugs
Accuracy (A) = 1 - (false_fail_rate) # How often our FAILs are real
Longevity (L) = days_since_last_miss / 30 # Stability over time (capped at 1.0)
Confidence ScoreRatingAction
>0.8HIGHMethodology is working — maintain it
0.5-0.8MEDIUMMethodology has gaps — identify and fix
<0.5LOWMethodology is ineffective — redesign required

Report Template

markdown
# Retrospective Validation Report
**Methodology evaluated:** {description}
**Period:** YYYY-MM-DD to YYYY-MM-DD
**Confidence score:** X.XX ({HIGH/MEDIUM/LOW})
## Key Findings
1.{finding with evidence}
2.{finding with evidence}
3.{finding with evidence}
## What's Working
-{practice} — evidence: {cite specific outcome}
## What's Not Working
-{practice} — evidence: {cite specific failure}
## Recommendations
### Continue
-{practice to maintain}
### Start
-{new practice to adopt}
### Stop
-{practice to abandon}
## Evidence Files
[list all files in e2e-evidence/retrospective/]

Save to e2e-evidence/retrospective/report.md.

Integration with ValidationForge

  • Retrospective evidence goes to e2e-evidence/retrospective/
  • Results inform updates to validation plans (create-validation-plan)
  • False PASS analysis reveals missing journeys to add to future validation
  • False FAIL analysis reveals flaky flows to fix or quarantine
  • Confidence score tracks overall validation system health over time
  • The verdict-writer agent can reference retrospective findings in meta-verdicts
All versions