The Chatbot did it – managing a new wave of crises

What to do when your AI agent goes off script

We’ve written about crises and crisis management comms before, and focused on the slow burners, and those that spring from nowhere. But guess what – predictably, AI has entered the room. Yep, while it can streamline, improve efficiencies, and work 24/7 it appears it can also get you into a spot of bother.

AI agents are great. They can answer a customer query at 2am and approve a refund over the weekend. But they can also quote a policy that doesn’t exist, and by the time anyone on the crisis management team (CMT) sees it – it’s too late.

Pass this off as a technical error and say sorry about that ‘n all, fix the bug and move on, right? Wrong. Because a wrong answer from your chatbot is still your company talking.

Learning from the Canadians

The case everyone in comms and legal now cites is Air Canada’s, where a customer asked the airline’s chatbot about bereavement fares. The bot invented a policy, telling the customer they could book at full price and claim a retroactive refund. When the customer tried to claim it, human agents pointed to the real policy on the website and refused.

Air Canada’s defence, taken to the BC Civil Resolution Tribunal, was that the chatbot was effectively a separate legal entity responsible for its own words. The tribunal member hearing the case called that argument “remarkable,” and found the airline hadn’t taken reasonable care to keep its chatbot accurate, was simply part of the company’s website, and it doesn’t get to disown what it said any more than a badly briefed spokesperson does.

This isn’t a precedent in English law, but it’s a warning to us all. The UK Jurisdiction Taskforce (UKJT) – who sets out how English law applies to new technology – has said there’s no English authority yet on this exact point, but liability for negligent misrepresentation will generally attach where a business held a chatbot out as speaking on its behalf. The UKJT’s own commentary points straight back to the Air Canada case as the example. 

The same discipline applies as in any other crisis – know who’s accountable before the incident, not after – only it’s applied to a spokesperson that never sleeps, can talk to thousands of customers simultaneously, and will continue to do so until it’s told not to.

It’s all about thresholds

So, be prepared and if you have a crisis management plan, speak to a digital PR, or crisis management agency to update it now. That means your CMT should be asking, before anything goes wrong:

  • Who is the named human owner? For every AI system that can act or speak without direct sign-off, this should be one person who can be found at 7am on a Saturday.
  • What’s the escalation threshold? – This is the point at which the agent should stop and hand a decision to a human, rather than act alone?
  • Is there a decision log? This shows not just what the AI did, but what it knew and why it acted, in language your comms and legal teams can actually use in a statement?
  • Who speaks for the business? If the story breaks – your usual spokesperson, or someone with the technical grounding to explain what happened without sounding like they’re reading a disclaimer, should be fully briefed throughout any planning.

The ‘escalation threshold’ is possibly the most important part of any of this. As AI and machine learning evolved over the past decade or so, and like most technology, via the military – there would always be a threshold. AI would examine and assess the data, come to a very informed decision, and then hand it over to a ‘commander’ – someone who could make a mission-critical decision. 

And here’s the crux, because getting your thresholds right in the first place can hugely influence behaviours downstream, and avert crises. The threshold is also the point at which you formulate your Q&As, because if they aren’t met, or they are compromised, you better have a robust response prepared for the media, your stakeholders and the public.

Your existing CMT structure barely needs to change to cover this. Add “AI / agent oversight” as a line next to your IT lead’s data breach procedure, and make sure whoever holds it is in the room when the crisis plan gets rehearsed.

The three Rs revisited

Review – What the agent was authorised to do, what it actually did, and where the gap between the two opened up.

That gap is your story, whether you tell it or someone else does.

Respond – But resist the instinct to explain the technology before you address those affected. Air Canada’s mistake wasn’t building a chatbot that got something wrong, It was trying to argue, in public, that the mistake wasn’t theirs to own.

Acknowledge it, fix it, and then explain the failings.

Repair – This is where AI crises differ slightly. The fix requires tighter escalation rules, a named owner, a documented boundary on what the agent can decide alone.

Stakeholders will want to see the change, and by conveying this over time will help to re-build trust.

An AI agent doesn’t get nervous, doesn’t go off script under pressure, and doesn’t slip up in a live interview the way a person might. That’s exactly why it’s dangerous from a reputational standpoint –  it will say the wrong thing with complete, unshakeable confidence, at scale, before anyone has a chance to intervene.

It may not be a person who made the mistake, but it’s still your responsibility. Building AI oversight into your crisis planning now, is a great deal cheaper than working it out for the first time in front of a court – Canadian or English.

If you’d like to talk through how AI accountability fits into your crisis management planning, get in touch with Catalyst PR.