
LLM proposes, code decides: the self-evolving causal DAG
Telco Technical deep-dive · The Trust Problem (2 of 2)
How do you let a system rewrite its own knowledge every week and guarantee it can never corrupt itself? This deep-dive on the self-evolving causal DAG shows the discipline — LLM proposes causal links, deterministic code decides what commits — with the five-phase weekly loop, the bootstrap sources, and the auto-rollback that enforces "only improve, never silently degrade."
Key terms in this article
· DAG — Directed Acyclic Graph — a graph of one-way links with no cycles; here, alarm cause→effect.
· F1 — F1 score — an accuracy measure balancing precision and recall (0–1).
· LLM — Large Language Model — the AI model that does the semantic reasoning.
· RAN — Radio Access Network — the cell-tower / radio layer of a mobile network.
· RCA — Root Cause Analysis — working backwards from symptoms to the single underlying fault.
The trust article ended on a promise: a system that learns new failure patterns every week while guaranteeing it can never corrupt its own reasoning. This is the mechanism — the Causal DAG, a directed acyclic graph of alarm causality that is the beating heart of the platform's safety story.
What the Causal DAG is (and is not)
The Causal DAG captures relationships of the form "Alarm X causes Alarm Y with confidence Z." It is stored as a named graph in Amazon Neptune, alongside the live topology. Two properties make it special:
· It is pre-computed. The graph is built and refined offline. At inference time, narrowing an alarm storm to candidate roots is a deterministic Gremlin traversal — under 10 milliseconds, no LLM call. This is design principle P4: no LLM on the inference hot path. The reasoning about what causes what has already happened; inference just walks the answer.
· It is not a probabilistic correlation table. Correlation says two alarms co-occur. The Causal DAG asserts a directed, validated causal edge with a confidence bound and a minimum-evidence threshold behind it. That directionality is what lets the system traverse upstream from symptom to cause instead of drowning in everything that merely happened at the same time.
Each CAUSES edge carries the metadata that makes safe evolution possible:
| Edge property | Role |
|---|---|
| confidence (0.0–1.0) | ADD threshold ≥ 0.60; REMOVE threshold < 0.30; only edges ≥ 0.60 are traversed at inference |
| evidence_count | Resolved incidents supporting the edge — minimum 3 to add a new one |
| source | expert_rule / granger / causalif / evocause / human_validated — provenance for every edge |
| batch_id | Tags the weekly batch that wrote it, enabling batch-level rollback |
| active (bool) | Soft-delete flag — retired edges are excluded from traversal but retained for audit |
The core principle, made mechanical
LLM proposes, code decides is not a slogan here — it is the literal control flow of the weekly evolution pipeline. The LLM adds semantic reasoning about why alarm A might cause alarm B, the domain knowledge statistics alone miss. Then deterministic code validates every proposal: acyclicity, confidence bounds, domain consistency, minimum evidence. Nothing the LLM emits touches Neptune unchecked.

Figure 1 — The weekly evolution loop. The LLM only ever writes to a proposal queue; deterministic code validates every change before commit, and Phase 3 can auto-roll-back.
The five-phase evolution pipeline
The graph improves through a disciplined cycle across five phases:
Phase 0 — Bootstrap (one-time). The initial graph is seeded from three independent sources so it is never dependent on a single method:
· Expert rules extracted from vendor documentation (Ericsson/Nokia/Huawei alarm guides).
· Historical correlation via Granger causality over the historical alarm corpus.
· Structure learning (CausalIF) to discover edges the other two miss.
| Bootstrap source | Method | Initial yield |
|---|---|---|
| Expert rules | Vendor docs + NOC engineers; committed at 0.90–0.98 confidence | ~100–200 edges |
| Granger causality | Bivariate test (p < 0.01, ≥10 co-occurrences) over 12+ months; STCR ontology filter cuts ~39M pairs to ~35K testable | ~200–500 edges |
| CausalIF | awslabs/causalif — LLM semantic priors + Hill-Climb/BDeu + bootstrap stability (>0.70); 0.83 directed F1 on benchmarks | ~100–300 edges |
That yields a ~400–1,000 edge starting graph over ~2,000–5,000 alarm-type nodes — a working baseline from day one, before a single incident is learned from.
Phase 1 — Evidence collection (continuous). Every resolved incident becomes evidence. When an engineer confirms or corrects an RCA, an extraction pipeline records support (or contradiction) for the relevant causal edges into a staging table — with recency weighting so the graph tracks a changing network.
Phase 2 — Weekly evolution batch. This is where proposing and deciding meet:
· Step 2.1 — Aggregation. Evidence is aggregated per (source_alarm, target_alarm) pair into candidate edges with a net evidence score and current graph state (exists / absent).
· Step 2.2 — LLM proposal. A frontier model — Claude Sonnet via Amazon Bedrock, deliberately *not* the inference-time Qwen3-32B — reviews each candidate and proposes one of five actions with a rationale. Because this is a weekly, latency-insensitive background task, semantic quality matters more than inference cost (~$2–5 per batch, ~$10–20/month). Separating the offline proposer from the online reasoner is itself part of the safety design: the model that suggests graph changes never touches the live inference path.
| Action | When the model proposes it |
|---|---|
| ADD | Net positive evidence and the model can explain the causal mechanism (new edge) |
| STRENGTHEN | Existing edge, accumulating positive evidence (confidence +0.05–0.15, capped at 1.0) |
| WEAKEN | Negative or conflicting evidence (confidence −0.05–0.15, floored at 0.0) |
| REMOVE | Net evidence ≤ −5 and existing confidence < 0.30 (soft-delete: active=false) |
| NO_CHANGE | Insufficient evidence or uncertain mechanism — the default when in doubt |
Evidence is recency-weighted (exponential decay, 0.95^weeks-old) so a changing network is tracked, not averaged into the distant past.
· Step 2.3 — Deterministic validation. Every proposal is checked by code: would it create a cycle? Is the confidence within bounds? Is there enough evidence? Does it violate a domain rule? Failures are rejected outright.
· Step 2.4 — Neptune commit. Only validated changes are written, atomically, to the graph.
Phase 3 — Evaluation & safety (weekly, post-commit). After committing, the system re-scores itself against the benchmark. Automatic rollback fires if any trigger trips — for example, RCA accuracy dropping more than ~3% week-over-week. This is design principle P6, monotonic improvement, made operational: the graph can only improve or stay neutral, never silently degrade.
Phase 4 — Advanced evolution (Month 6+). Once the base loop is stable, higher-order refinements switch on: conditional-dependence testing to prune spurious edges, temporal decay so stale causal links fade, cross-domain causal discovery (RAN↔Transport↔Core), counterfactual confidence calibration for high-confidence edges, and periodic CausalIF revalidation.
Inference-time usage
At incident time the flow is boringly deterministic — exactly what you want:
1. Symptom alarms arrive from the storm detector.
2. A Gremlin traversal walks the Causal DAG upstream from those symptoms to candidate roots (<10ms).
3. The ranked candidates are handed to the Supervisor agent as a pre-narrowed hypothesis space — 47 alarms become 3 candidate roots before the reasoning model spends a single token.
That hand-off is the quiet efficiency win of the whole architecture. The LLM does not search; it adjudicates a short, pre-validated list. Fast, cheap, and — because the list came from a validated graph, not a generative guess — safe.
Why this beats the alternatives
A pure-LLM approach reasons from scratch every time: slow, expensive, and non-deterministic. A pure-statistical approach (correlation rules) cannot explain why and cannot incorporate engineer knowledge. The Causal DAG takes the semantic strength of the LLM and the guarantees of deterministic systems and refuses to compromise either — the LLM never gets to corrupt the graph, and the graph never has to guess.
That is the technical substance behind "you can't let an LLM touch a live network — unless it can't hallucinate." The LLM here genuinely cannot, because the only thing it is allowed to touch is a proposal queue that code inspects before anything becomes real.
Next: The Hidden Cost of Renting Someone Else's Brain — why a self-hosted model beats a metered API for a 24/7 NOC — the cost case.
Further reading — external sources
· NetCause — graph-temporal counterfactual root-cause analysis
https://arxiv.org/html/2606.13543v1
· Agentic Diagnostic Reasoning — the case against hard-coded correlation (counterpoint)
https://arxiv.org/pdf/2601.07342v1
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article