Breaking things is how you learn what holds.
Anyone who has spent enough time in business continuity knows this instinctively. Resilience is not about following the plan. It is about knowing what to do when the plan fails. That mindset — test everything, trust nothing, push until it breaks — does not stay inside the office. It becomes a way of operating, even in our personal life.
So when generative AI showed up, it went straight into the work. Not as an experiment. Not after reading a whitepaper about digital transformation. It went into real client engagements, tested against real deliverables, challenged with real problems.
“Write me an ISO 22301 gap analysis for a mid-size financial institution.” Forty seconds. Three pages. Perfect structure. Confident language.
Then you read it with practitioner eyes. And you realize half of it is elegant garbage.
But before getting into what happened, it is worth pausing on what these tools actually are — because most professionals using them daily have no idea what they are interacting with.
So what is a LLM?
A Large Language Model is not intelligent. It does not understand. It does not reason. It does not know your business, your standard, or your client.
What it does is predict the next word.
That sounds trivial. It is not. When you train a neural network on trillions of words — every book, every website, every manual, every forum post, every academic paper available on the internet — the statistical patterns it absorbs become extraordinarily powerful. It learns that after “the Recovery Time Objective for critical processes should not exceed” the most probable next words are something like “four hours” rather than “banana.” It learns the structure of professional documents, the rhythm of executive communication, the vocabulary of ISO standards.
The result is a machine that produces text that looks like it was written by someone who knows what they are talking about. The formatting is perfect. The vocabulary is precise. The confidence is absolute.
And that confidence is the trap.
The model is not retrieving facts from a database. It is not consulting ISO 22301 and checking its answer against the clauses. It is generating the most statistically probable sequence of words given the prompt. When the training data contains enough correct examples, the output is correct. When it does not — or when the question requires judgment that no text corpus captures — the model does not say “I don’t know”. It generates the most probable-sounding answer with the same absolute confidence.
That is the famous infamous hallucination. Not a bug. A feature of the architecture.
But there is something worse than hallucination: sycophancy.
Ask a Large Language Model to write a business continuity plan, it will write one.
Then tell it “the recovery strategies are too generic.” It agrees, apologizes, rewrites.
Tell it “actually, the first version was better.” It agrees again, apologizes again, reverts.
Tell it “are you sure this meets ISO 22301 clause 8.4?” It assures you it does.
Tell it “I don’t think it does.” It agrees with your concern and modifies the output.
The model is not evaluating its own work. It is agreeing with whatever the last message says. It is optimized to be helpful, which in practice means it is optimized to say yes.
For a BCM professional, this is poison. The entire value of an expert is the ability to say no. No, that RTO is unrealistic. No, this plan will not survive a real incident. No, your board is wrong about the risk appetite. A tool that validates everything does not augment expertise. It flatters it.
As you certainly did it too, I tested many models like Gemini, Claude or ChatGPT. Different architectures, different strengths, different failure modes. But the same pattern emerged across all of them.
Feed the AI a well-structured standard like ISO 22301, and it performs brilliantly on anything derivable from rules: compliance matrices, checklist generation, clause mapping, gap identification. The output is fast, accurate, and often better structured than what a consultant would produce manually.
Feed it a messy, political, context-dependent problem — “should this organization invest in a third data center?” — and it produces eloquent nonsense. Confident, well-formatted, professionally structured nonsense.
Ask it to draft a post-incident report: excellent first pass. Ask it to decide whether to activate the crisis team at two in the morning based on incomplete data: meaningless output. Ask it to know that the operations director will reject page three because the RTO contradicts what he told the CEO last month: impossible.
No model carries the scar tissue of a failed recovery. No model has sat through the meeting where recovery priorities were decided not by impact data but by political leverage. No model reads body language, navigates egos, or knows when to push and when to shut up.
But — and this matters just as much — a huge portion of daily BCM, risk, cybersecurity, DORA work does not require scar tissue. It requires structure. Formatting compliance matrices. Populating templates. Computing impact cascades. Generating first drafts of documents that will be rewritten anyway. Tasks where experience adds nothing because the rules are explicit and the output is deterministic.
One conclusion became unavoidable after months of testing across every model available: AI in a professional discipline should only be used by someone who is competent in that discipline.
Not because the AI is bad. Because when it is wrong — and it will be wrong — only a practitioner can catch it. A junior who does not know what a valid RTO looks like will accept whatever the model produces. A manager who has never written a real continuity plan will not notice that the recovery strategy collapses under pressure. The AI does not come with a warning label on the paragraphs it hallucinated.
Never trust it blindly. Never use it for something serious if you cannot correct it yourself. And never mistake a fluent output for a competent one.
Two realities coexist in every working day. Tasks where AI is dangerous because it mimics competence without possessing it. And tasks where practitioners are inefficient because the work demands structure, tons of writing, not judgment.
The question that we should ask ourselves: how and where exactly could the AI help us in our expert disciplines?
Next: Article 2 — Human-AI collaboration in expert professional disciplines.