Skip to content

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Conference: NeurIPS 2026 โ€” Evaluations & Datasets Track
arXiv: 2609.36686
Code: https://github.com/cruiseresearchgroup/DecompRCA
Area: Causal Inference
Keywords: root cause analysis, retrieval coverage, conditional reranking, fault propagation, domain knowledge

TL;DR

The paper separates root cause ranking accuracy into whether the cause enters the candidate set and whether it ranks highly once retrieved, audits six suites from four datasets, and uses multi-signal retrieval plus single-call LLM reranking to show why the two failures require different remedies rather than simply more causal structure or reasoning capacity.

Background & Motivation

Root cause analysis (RCA) identifies the origin of an anomaly among many monitored metrics, rather than merely the metric with the largest change. Statistical methods such as BARO rank deviations from normal behavior, whereas PC/FCI combined with CIRCA, PageRank, or random walks rely on learned structure. Direct faults in microservices often make the source conspicuous; in systems with propagation, feedback, and amplification, downstream symptoms can be stronger than the origin. The first detected anomaly also need not be the first actual deviation: a weak source signal may cross a detection threshold later.

Conventional top@k checks only whether the final first k entries contain a cause. It merges a cause excluded from the manageable candidate pool with a retrieved cause ranked incorrectly. The first failure calls for better sensing coverage, window segmentation, or retrieval signals; the second calls for better evidence integration and ranking. Evaluation restricted to microservice faults with conspicuous origins risks treating nearly saturated retrieval coverage as a universal cross-system condition.

Core Idea: treat candidate coverage as a hard ceiling on ranking, measure retrieval and conditional reranking separately, and diagnose bottlenecks using both actual candidate pools and explicitly oracle-controlled pools containing the cause.

Method

Overall Architecture

The main contribution is an evaluation decomposition and empirical audit, not a claimed production-ready system. The validation pipeline receives a normal reference window, a fault window, and a supplied detection time. It selects a bounded candidate set using magnitude, onset, and discrete state-change signals, then submits structured candidate evidence to a pretrained LLM for one-call reranking. Optional domain knowledge (DK) describes system roles and propagation relationships without providing the current scenario's fault labels.

Two levels must be distinguished. Truncating an existing method's ranking at K enables a post-hoc retrieval analysis; it does not establish that the method physically contains a separate retrieval module. The paper's validation pipeline does explicitly retrieve and then rerank, with no new candidates allowed at reranking. Ground truth is used only for metric computation and explicitly labeled controlled-pool experiments, not ordinary inference.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Normal reference + fault window<br/>Supplied detection time"] --> B["Coverage decomposition"]
    B --> C["Multi-signal retrieval"]
    C --> D["Evidence-based reranking"]
    K["Optional system document<br/>No scenario fault labels"] -.-> D
    D --> O["Candidate ranking"]
    O --> E["Controlled bottleneck diagnosis"]
    G["Ground truth<br/>Evaluation only"] -.-> E
    N["Statistics fit on normal reference<br/>No fault classifier training"] -.-> C

The first three designs connect evaluation interpretation with inference; the final design is an offline diagnostic intervention, not a deployment stage that reveals the cause. Normal reference data estimate statistics; no new model is trained with fault-label supervision.

Key Designs

1. Coverage decomposition: expose different bottlenecks behind identical total accuracy

Let \(\mathcal C\) be a scenario's candidate set and \(\mathcal G\) its set of valid causes, allowing multiple ground-truth causes. Retrieval succeeds when the sets intersect, not only when every cause is found. Retrieval@K is the fraction of scenarios with such a hit. Rerank@k is the fraction of those retrieved scenarios whose final first k entries contain any valid cause. Provided that ranking stays within the pool and uses consistent ground-truth mapping and hit rules:

\[ \mathrm{top@}k=\mathrm{Retrieval@}K\times\mathrm{Rerank@}k. \]

Failures must use the same denominator of all scenarios: retrieval failure is \(1-\mathrm{Retrieval@}K\), and retrieved-but-misranked failure is \(\mathrm{Retrieval@}K-\mathrm{top@}k\). The second is an absolute failure fraction, not the conditional failure rate \(1-\mathrm{Rerank@}k\); only the two absolute terms sum to total error. If retrieval has no hits, conditional reranking has no estimable denominator and should not be presented as an ordinary valid accuracy.

For an existing full-ranking method, treating its first K entries as the candidate set makes Retrieval@K equal to its top@K; conditional reranking is top@k divided by top@K. At K equal to all observable metrics, coverage usually saturates, but unobservable or preprocessed-away causes can still be absent. Different K values answer different questions: whether the cause is visible within a practical diagnostic budget, versus whether ranking all metrics works. These are not the same experimental condition.

Graph methods can lose coverage through missing nodes, candidate-selection paths, or structural errors. However, the paper's implication that one missing edge excludes the cause is an illustration of a particular path-selection mechanism, not a universal theorem for every graph algorithm. Appendix C further states that CIRCA appends unscored sensors to its final list, whereas PageRank/RandomWalk rank graph nodes only. Absence from the graph and absence from the final list therefore require separate interpretation.

2. Multi-signal retrieval: cover heterogeneous anomaly signatures within a fixed total budget

Magnitude fits a RobustScaler on the normal window and takes the largest robust standardized deviation in the fault window. Appendix A defines the retrieval magnitude as a signed maximum; Appendix C specifies that CPS BARO uses the maximum absolute value, while RCAEval keeps the original signed maximum implementation. These differences mean the retrieval magnitude rule and BARO rows across domains cannot unconditionally be treated as the same algorithm.

Onset finds the first fault-window row whose absolute standardized value exceeds 1.5 and ranks earlier occurrences first; normal-window data provide the mean and standard deviation. If normal variance is zero, it instead uses the first deviation from the normal mean exceeding \(10^{-4}\). Metrics with no threshold crossing are excluded from that signal. This measures the first detectable deviation, not an oracle's true physical fault onset.

Discrete state-change first checks whether a metric has at most 5 unique values in either window, then requires different fault and normal modes. It orders candidates by the first departure from the normal mode exceeding \(10^{-2}\). This contributes discrete-state evidence without asserting that every actuator transition is a cause; normal operating transitions can also interfere.

The total budget K is divided equally among active signals, with remainders allocated to earlier signals. The three-signal allocations are 5/5/5 at K=15, 4/3/3 at K=10, and 2/2/1 at K=5. Their lists are merged in magnitudeโ€“onsetโ€“state-change order and deduplicated, so the actual pool can contain fewer than K items. Adding signals replaces magnitude candidates rather than expanding the pool for free. Coverage is consequently not monotone: WADI gains nothing at K=15, and HVAC's three-signal configuration is slightly worse than its two-signal configuration.

3. Evidence-based reranking: compare origins and symptoms rather than merely fuse scores

Each candidate becomes one record containing its name, magnitude score, first-deviation offset from detection, normal and fault means, and absolute and relative shifts. The LLM reads the complete list in one call and returns a ranking restricted to supplied candidates. Records do not state which retriever selected the item, so ranking relies on heterogeneous statistics and semantics rather than selection labels.

Without DK, the system prompt provides a generic RCA role. With DK, it additionally supplies system documentation covering component functions, naming conventions, propagation pathways, and general operational priors. The authors report drafting from public documentation and revising descriptive quality without optimizing against test fault outcomes, excluding scenario labels. Nevertheless, the document can contain human-provided priors: label-free is not human-knowledge-free and does not mean every external fact was learned unsupervised from telemetry.

Candidates are actually presented in retrieval-signal blocks, although the prompt calls them โ€œranked by deviationโ€; Appendix A explicitly discloses this mismatch. Position and wording can influence the model. Parsing removes out-of-pool names and appends omitted candidates in retrieval order; unparsable responses in the main end-to-end experiment fall back to retrieval order. This preserves the pool boundary but does not eliminate position bias.

4. Controlled bottleneck diagnosis: separate evidence value from coverage

Same-candidate controls rerank the exact lists seen by the LLM using individual signals and equal-weight Borda fusion, testing whether improvements come only from a better pool. Borda sums each candidate's one-indexed ranks under all three signals; an unscorable item receives the worst rank plus one. It is a fixed-rule control, not a learned fusion model.

A separate experiment reserves a ground-truth slot when retrieval misses the cause, replacing the last candidate and inserting the cause at a seeded random position. Every method receives the same small pool containing the cause; an all-candidates control supplies every sensor. The reserved-slot experiment uses evaluation labels and measures who ranks better if coverage is solved, not realistic end-to-end performance. One WADI and one SWaT cause fall outside the evaluated sensor set and cannot be inserted, changing controlled denominators to 13/35 while main results retain 14/36.

This intervention also changes evidence composition and input position. Appendix G's random non-cause pool stress test includes cases where a baseline beats the LLM, including SWaT CIRCA-FCI outperforming statistical baselines. An advantage on controlled retriever pools therefore does not imply victory on arbitrary pools, and graph-method weaknesses in the main setting do not establish a universal limitation of structural RCA.

A Worked Example

The following abstract example explains the metrics and is not a measured fault case from the paper. Suppose a fixed-budget pool contains at least one valid cause in 64 of 100 scenarios. The reranker puts any valid cause first in 32 of those scenarios.

Retrieval@K=0.64, Rerank@1=0.50, and end-to-end top@1=0.32. The cause is absent in 36% of all scenarios; a further 32% have a retrieved but misranked cause. Conditional ranking failure among retrieved scenarios is 50%, not 32%.

In one retrieved scenario, abstract metric A changes only slightly but early, B changes strongly but later, and C exhibits a discrete state change. Retrieval can admit evidence from all three signatures. The reranker combines the normal reference and system description to assess which item resembles an origin, rather than automatically selecting the largest deviation B.

If A never enters the pool, even a stronger reranker cannot return it. Inserting A with ground truth can test ranking ability, but is permitted only in the diagnostic experiment and cannot count as A being recovered by real inference.

Loss & Training

The paper introduces no new training loss, does not fine-tune the reranker with fault labels, and does not train a fault classifier. Normal data estimate scaling, means, variances, and modes; a pretrained LLM performs inference. Graph baselines separately learn structure from short windows or normal corpora, which is not a training module in the validation pipeline.

The main reranker is gpt-oss-120b with a shared configuration across datasets, temperature 1.0, and a 4096-token output limit. Each configuration is repeated 3 times without a supplied random seed. Reported means and sample standard deviations capture API decoding variability across runs, not three training seeds or confidence intervals over the scenario population.

Benchmarks supply the windows: WADI/SWaT use labeled fault start times and the preceding 30-minute normal reference; RCAEval uses windows around injection, typically about 6 minutes and at most 10 minutes per side; HVAC uses the occupied fault day and a matching seasonal normal occupied day of 900 rows. Known timing and reference selection remain prerequisites, so the method has not solved online anomaly detection and fault segmentation end to end.

Key Experimental Results

Main Results

Four datasets yield six suites: WADI has 14 scenarios, SWaT 36, HVAC 48, and each RCAEval RE1-OB/SS/TT suite has 125. The first three use metric-level ground truth. RCAEval maps metrics to service prefixes, folds database instances into their owning services, and deduplicates predictions before service-level scoring. Scores at different granularities do not represent identically difficult tasks.

The following top@1 selection comes from Table 3. The best statistical baseline is epsilon-Diagnosis on HVAC and BARO elsewhere; the best graph baseline is selected from six PC/FCI ranking configurations. LLM values are means ยฑ sample standard deviations over three inference runs.

Suite Best statistical baseline Best graph baseline LLM without DK LLM with DK
WADI 0.214 0.214 0.333 ยฑ 0.082 0.310 ยฑ 0.041
SWaT 0.194 0.111 0.213 ยฑ 0.016 0.148 ยฑ 0.016
HVAC 0.146 0.104 0.153 ยฑ 0.043 0.278 ยฑ 0.012
RE1-OB 0.784 0.576 0.875 ยฑ 0.009 0.888 ยฑ 0.008
RE1-SS 0.856 0.632 0.872 ยฑ 0.024 0.944 ยฑ 0.008
RE1-TT 0.560 0.328 0.653 ยฑ 0.024 0.685 ยฑ 0.009

The no-DK configuration's mean top@1 is at least as high as the best baseline on all six suites. However, HVAC's 0.153 versus 0.146 and SWaT's 0.213 versus 0.194 are narrow margins, not statistically established universal victories. DK is not uniformly beneficial end to end: SWaT drops from 0.213 to 0.148, whereas HVAC rises from 0.153 to 0.278.

HVAC's main rows reuse predictions from controlled pools but count all 18/48 scenarios originally missed by retrieval as failures. The pools are unchanged for the 30 retrieved scenarios. Appendix F gives this scoring-equivalence argument; the controlled table's 0.326 must not be presented as realistic HVAC end-to-end accuracy.

Ablation Study

The fixed-K=15 retrieval analysis comes from Table 1. Three signals share one total budget; each retriever does not independently receive 15 slots.

Suite Magnitude only Magnitude + onset Three signals Note
WADI 0.64 0.57 0.64 Three signals ultimately tie magnitude
SWaT 0.53 0.61 0.67 State evidence adds coverage
HVAC 0.35 0.65 0.63 Onset dominates; the third signal slightly hurts
RE1-OB 1.00 0.99 0.98 Replacing magnitude candidates slightly hurts
RE1-SS 1.00 1.00 1.00 Retrieval is saturated
RE1-TT 0.98 0.98 0.91 Three signals reduce coverage

The next table uses Table 4's retrieval-controlled pools, forced to contain a valid cause, with all baselines rerun on identical pools. It isolates ranking and is not an end-to-end table; WADI/SWaT have 13/35 scenarios.

Suite Best same-pool baseline LLM without DK LLM with DK With-DK gain over best baseline
WADI 0.308 0.410 ยฑ 0.044 0.436 ยฑ 0.089 +12.8 pp
SWaT 0.200 0.229 ยฑ 0.029 0.267 ยฑ 0.017 +6.7 pp
HVAC 0.188 0.188 ยฑ 0.055 0.326 ยฑ 0.012 +13.9 pp
RE1-OB 0.784 0.883 ยฑ 0.005 0.891 ยฑ 0.012 +10.7 pp
RE1-SS 0.856 0.859 ยฑ 0.009 0.928 ยฑ 0.000 +7.2 pp
RE1-TT 0.560 0.688 ยฑ 0.024 0.741 ยฑ 0.009 +18.1 pp

Gains retain the source's report, including HVAC's +13.9 pp, while displayed 0.326โˆ’0.188 gives +13.8 pp. Unreported precision may explain the difference, but the cache does not confirm this; the values are not silently reconciled. Cross-table discrepancies should first be checked against pool composition, denominators, and stochastic baseline settings.

Key Findings

  • Table 2's no-DK decomposition gives conditional Rerank@1 of 0.52/0.32/0.24 for the CPS suites and 0.90/0.87/0.72 for microservices. CPS both loses more causes at retrieval and ranks retrieved causes less reliably.
  • SWaT's Table 2 retrieval failure is 0.33 and retrieved-but-misranked failure is 0.45. Neither should be confused with conditional ranking failure of approximately 0.68.
  • Retrieval can conceal DK's value. Across 33 scenario-runs from 11 originally missed SWaT scenarios in the controlled experiment, no-DK ranks the cause first 0/33 times and with-DK 6/33 times. They still all score zero under ordinary end-to-end evaluation.
  • Timestamp sensitivity measures retrieval: shifting detection 5 minutes earlier reduces HVAC Retrieval@15 by 27.1 pp and SWaT by 22.2 pp. RCAEval's maximum change at tested ยฑ1/ยฑ2-minute offsets is at most 3.2 pp.
  • The best long-normal-window graph top@1 is 0.214/0.083/0.083 on WADI/SWaT/HVAC, leaving the main-setting conclusion unchanged. RE1-TT has no global-graph experiment, so long-window results do not cover all six suites.

Highlights & Insights

  • Candidate coverage bounds every pool-restricted reranker. This makes stronger LLMs and more observable evidence separately evaluable interventions rather than attributing all progress to reasoning capacity.
  • Fixed-total-budget ablations avoid mistaking more candidates for a better signal. RE1-TT's coverage loss particularly shows that heterogeneous evidence is not automatically beneficial; retrieval should match how a system manifests faults.
  • Natural-language DK avoids requiring a precise causal graph before supplying knowledge. It also introduces conflicting priors and documentation-quality risks, making both real-coverage and controlled-ranking evaluations necessary.

Limitations & Future Work

  • CPS coverage is largely limited to water and building systems, with only 14 WADI scenarios. Three decoding repetitions do not establish deployment stability or statistical significance; settings are fixed on the same benchmarks rather than independently validated systems.
  • Known fault windows, labeled start times, and normal references remain strong assumptions. HVAC's occupied-schedule selection particularly illustrates manual reference design. Future audits should jointly examine detection, segmentation, and coverage instead of only the final ranker.
  • Controlled insertion position affects the LLM. In Appendix G, inserted HVAC causes receive hits in the first five slots but none outside them. Removing those hits lowers no-DK from 0.188 to 0.153, below controlled RCD's 0.188, weakening a robust interpretation of โ€œno-DK is never worse.โ€
  • De-identification preserves type, unit, and structural semantics. It mitigates dependence on public benchmark names but cannot prove the absence of all pretrained memorization or leakage sources.
  • Broad source claims require qualification: โ€œthe LLM is never worseโ€ on controlled retriever pools does not include Appendix G's random-pool counterexamples or HVAC's no-DK comparison after insertion-position hits are discounted. Table 4's HVAC gain of +13.9 pp differs from +13.8 pp obtained by subtracting displayed values; the original report is retained rather than silently normalized. The prompt's description of deviation-ranked inputs also conflicts with actual multi-signal presentation order.
  • Industrial use should remain assistive diagnosis with a short candidate list, supporting evidence, and human review, not automatic control based on an LLM ranking. This note provides no procedures for inducing faults or damaging facilities.
  • vs BARO / epsilon-Diagnosis: statistical deviations work well for direct faults, but strong downstream symptoms can obscure propagation origins. The paper does not invalidate statistical methods; it motivates extra signals and conditional ranking evaluation under finite candidate budgets.
  • vs RCD / CIRCA / PC-FCI graph ranking: although grouped with statistical comparators, RCD includes localized causal search; CIRCA uses structure and regression tests. The audit concerns specified learned graphs, preprocessing, and windows, not expert graphs, lag-aware models, or all causal RCA.
  • vs RankGPT / SpecRCA: listwise reranking and hypothesis-then-verification can both be evaluated through this decomposition. The reusable idea is logging candidates actually considered and separating misses from misranking, rather than copying a fixed prompt. Coverage-aware detection is a research direction, not a system implemented here.

Rating

  • Novelty: 4/5 โ€” The decomposition is a conditional-probability relation; its novelty lies in an actionable RCA audit perspective.
  • Experimental Thoroughness: 4/5 โ€” Cross-domain, same-pool, forced-coverage, and robustness controls are useful, but CPS samples are small and some claims are protocol-dependent.
  • Writing Quality: 3/5 โ€” The main narrative is clear; universal phrasing, subset explanations, and table rounding require careful reading.
  • Value: 4/5 โ€” The framework has practical value for detector selection, retrieval-bottleneck diagnosis, and knowledge-benefit evaluation.