
You can't let an LLM touch a live network — unless it can't hallucinate
Telco Thought leadership · The Trust Problem (1 of 2)
Every exec asks the same question 30 seconds into the demo: "you're letting it change the live network?" — and they're right to, because LLMs hallucinate structurally, not occasionally. This article makes the reframe that turns autonomy from reckless to deployable: separate proposing from deciding*, so the model's creativity is harnessed but its mistakes physically cannot reach the network.*
Key terms in this article
· DAG — Directed Acyclic Graph — a graph of one-way links with no cycles; here, alarm cause→effect.
· GDPR — General Data Protection Regulation — EU/UK data-protection law.
· LLM — Large Language Model — the AI model that does the semantic reasoning.
Every executive who has watched a convincing NOC demo asks the same question about thirty seconds in: "And you're letting it change the live network?" It is the right question. It is, in fact, the only question that matters, because it separates the pilots that reach production from the ones that quietly die in a sandbox.
Large language models hallucinate. Not occasionally — structurally. They are trained to produce plausible continuations, and a plausible-sounding root cause delivered with total confidence is exactly the failure mode you cannot tolerate on infrastructure that carries emergency calls. A NOC that auto-remediates based on a hallucinated diagnosis has not automated the operation; it has automated the outage.
So the industry has mostly split into two unsatisfying camps. One keeps the LLM firmly in an advisory box — it suggests, a human always decides — which caps the value at "faster copy-paste." The other lets the model act and crosses its fingers, which works right up until the night it doesn't. Neither is autonomy. Autonomy that you can actually deploy requires a third path: an architecture where the model's creativity is harnessed but its mistakes cannot reach the network.
The reframe: separate proposing from deciding
The pattern that unlocks this is deceptively simple to state and demanding to build: the LLM proposes; deterministic code decides.
The insight is that LLMs and traditional software fail in opposite ways. An LLM is brilliant at semantic reasoning — understanding why a fibre cut on an aggregation switch would produce this particular pattern of downstream alarms — and terrible at guarantees. Deterministic code is the reverse: it cannot reason about novel situations, but when it says a graph is acyclic or a confidence score is above threshold, it is right, every time, with no probability attached.
Put them in their correct roles and the weaknesses cancel:
· The LLM adds the semantic layer — hypotheses, causal explanations, the domain knowledge that pure statistics miss.
· Deterministic code validates every proposal before it can affect anything real — acyclicity checks, confidence bounds, domain-consistency rules, minimum-evidence thresholds.
· No LLM output reaches the live system unchecked. No LLM call sits on the inference hot path.
This is not a philosophical stance; it is an architectural boundary you can point to in a diagram and test in CI.
Trust is built from four concrete things
"Trustworthy AI" is a phrase that has been drained of meaning by overuse. In a NOC it decomposes into four things you can actually verify:
1. Confidence gates. Below a hard threshold, the system does not act — it escalates to a human. Auto-remediation on critical elements requires explicit approval. Novel patterns trigger full escalation by default.
2. Determinism where it counts. The knowledge the system reasons from — the causal relationships between alarms — is pre-computed and validated offline. At the moment of an incident, finding the root cause is a database traversal, not a generative guess. It is fast (milliseconds) and it is auditable.
3. Monotonic improvement. The system updates itself, but it can only get better or stay flat — never silently worse. If this week's accuracy drops beyond a small tolerance, it rolls back automatically.
4. An audit trail a regulator can read. Every decision, every piece of supporting evidence, every confidence score, retained for twelve months to satisfy Ofcom and GDPR. Trust that cannot be inspected is just marketing.
The market's answer, and why mine is stricter
The market has largely chosen a different answer: "glass box" autonomy — make the AI's reasoning transparent and keep a human approving anything that touches the control plane. That is a real, defensible model, and it is the one the major platform vendors have shipped. My argument is deliberately stricter. Transparency explains a mistake, usually after the fact; deterministic validation prevents it from ever reaching the network. Explainability makes the human approval step meaningful, but it still routes every high-stakes action through a tired human at 2am — the exact bottleneck autonomy is meant to remove. I am arguing the harder line on purpose: not that the market already agrees, but that a glass box you have to watch is weaker than a gate that cannot let a bad action through.
Why this is the real moat
Here is the part that surprises people. The hard, defensible part of an autonomous NOC is not the model. Anyone can fine-tune a capable open-weight model. The moat is the safety architecture — the validated causal knowledge base, the deterministic guardrails, the evaluation harness that proves the thing is honest. That is months of careful engineering that a competitor cannot shortcut by swapping in a bigger model.
It is also what turns a board-level objection into a board-level asset. "We let AI touch the live network" is terrifying. "We have an AI that proposes actions, a deterministic layer that validates every one, confidence gates that escalate anything uncertain, and a twelve-month audit trail" is a governance story — the kind that gets a deployment approved.
The next article opens the box on the mechanism that makes this real: the self-evolving causal graph that lets the system learn new failure patterns every week while guaranteeing it never poisons its own knowledge.
Next: LLM Proposes, Code Decides — inside the self-evolving Causal DAG, the weekly evolution pipeline, and the deterministic validation that makes learning safe.
Further reading — external sources
· TM Forum — the AI trust gap: 72% believe, 14% can prove it
https://www.tmforum.org/news-insight/newsroom/tm-forum-reveals-ai-trust-gap-72-of-csps-say-their-ai-is-trustworthy,-only-14-can-prove-it
· Nokia — the glass-box imperative for governing AI-native networks
https://www.nokia.com/blog/the-glass-box-imperative-governing-intent-based-automation-in-ai-native-networks/
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article