Skip to content
Bowtie risk templates

Bowtie — unsafe or non-compliant LLM response to a customer

Barrier-based bowtie for a customer-facing LLM assistant emitting unsafe or non-compliant output, with preventive barriers against injection, jailbreak, stale retrieval and ungrounded generation and recovery barriers for regulatory breach and PII disclosure, used in an AI safety case or Consumer Duty review.

Template previewBowtie risk
Assistant emits an unsafe or non-compliant response to a customerHAZARDA general-purpose LLM with authority to state poli…THREATSPREVENTIVE BARRIERSRECOVERY BARRIERSCONSEQUENCESClassifier not retrained after thelast two jailbreak familiesWeekly red-team corpus refreshgates the classifier releaseAttachment text stripped ofinstruction-like spans beforeprompting60% · preventionInjection classifier on allretrieved and uploaded text85% · detectionUntrusted content rendered tothe model inside a quoted,non-executable block50% · controlSystem prompt hardening plusrefusal training on the340-case suite90% · preventionPer-session jailbreak scoringwith step-up to a human agent70% · detectionPolicy team publishes outside theCMS, so no reindex event firesMonthly reconciliation of CMSdocument IDs against index IDsIndex freshness SLA: reindexwithin 4 h of a policypublish80% · preventionEffective-date filter onevery retrieval query75% · controlAbstain rule: no answer belowa 0.62 retrieval score80% · controlCitation-required decoding,answers without a cited spanare dropped70% · preventionBlocking eval gate on the 1200-case regression suite inCI85% · preventionShadow traffic comparison for72 h before any promotion60% · detectionOutput classifier blocksregulated-advice languagebefore send80% · detectionSampled QA review of 2% oftranscripts by the complianceteam30% · detectionBreach register and 72 hregulator notificationrunbook25% · recoveryAnswer carries a citation thecustomer can open and check35% · mitigationGoodwill remediation andcomplaint handling within 8weeks50% · recoveryPII detector and redaction onthe outbound response90% · detectionTenant-scoped retrievalfilter, index partitioned percustomer85% · controlData-breach response: notifythe DPO within 24 h, ICOwithin 72 h20% · recoveryKill switch drops theassistant to a scriptedfallback in under 5 min55% · controlPrepared holding statementand press escalation path40% · recoveryIndirect prompt injection inan uploaded statement orticket attachment6/yr → 0.18/yrDirect jailbreak of the chatsurface by an end user40/yr → 1.2/yrRetrieval returns asuperseded policy document24/yr → 1.2/yrModel answers with nosupporting retrieval(ungrounded generation)90/yr → 5.4/yrRegression introduced by amodel or prompt change12/yr → 0.72/yrRegulatory breach —unsuitable advice givenunder Consumer Dutysev 5 · risk 4.57Customer acts on wronginformation and suffersfinancial losssev 4 · risk 11.31PII of another customerdisclosed in the responsesev 5 · risk 0.522Incident reaches media orsocial platformssev 3 · risk 7.05Unsafe ornon-compliantresponse…8.7/yrtop-event frequencyTop event8.7/yrinherent 172/yrResidual risk23.45inherent 2,924Risk reduction99.2%threat side 94.9%Barriers215 threats · 4 consequencesDominant threatModel answers with no supporting retrie…62.1% of the top eventResidual risk 23.45 against an inherent 2,924 — the barriers remove 99.2% of it.Barrier criticality — residual risk if that one barrier were removed1. Abstain rule: no answer below a 0.62 retrieval sco…+58.21 ×3.52. Citation-required decoding, answers without a cite…+33.96 ×2.43. System prompt hardening plus refusal training on t…+29.11 ×2.24. Output classifier blocks regulated-advice language…+18.27 ×1.85. Index freshness SLA: reindex within 4 h of a polic…+12.94 ×1.66. Goodwill remediation and complaint handling within…+11.31 ×1.5FindingsEach of the 4 consequences is credited the full top-event frequency of 8.7/yr — the model here is that one top event produces all of them, so the residual score of 23.45 adds 4 outcomes to one event. That is right when they happen together and up to 4× too much when …

Make it your own.

title "Assistant emits an unsafe or non-compliant response to a customer"
hazard "A general-purpose LLM with authority to state policy to 1.2 M retail banking customers"
top "Unsafe or non-compliant response reaches a customer"
unit "/yr"

threat "Indirect prompt injection in an uploaded statement or ticket attachment" likelihood: 6.0 {
  barrier "Attachment text stripped of instruction-like spans before prompting" effectiveness: 0.6 type: prevention
  barrier "Injection classifier on all retrieved and uploaded text" effectiveness: 0.85 type: detection {
    escalation "Classifier not retrained after the last two jailbreak families" {
      control "Weekly red-team corpus refresh gates the classifier release"
    }
  }
  barrier "Untrusted content rendered to the model inside a quoted, non-executable block" effectiveness: 0.5 type: control
}

threat "Direct jailbreak of the chat surface by an end user" likelihood: 40.0 {
  barrier "System prompt hardening plus refusal training on the 340-case suite" effectiveness: 0.9 type: prevention
  barrier "Per-session jailbreak scoring with step-up to a human agent" effectiveness: 0.7 type: detection
}

threat "Retrieval returns a superseded policy document" likelihood: 24.0 {
  barrier "Index freshness SLA: reindex within 4 h of a policy publish" effectiveness: 0.8 type: prevention {
    escalation "Policy team publishes outside the CMS, so no reindex event fires" {
      control "Monthly reconciliation of CMS document IDs against index IDs"
    }
  }
  barrier "Effective-date filter on every retrieval query" effectiveness: 0.75 type: control
}

threat "Model answers with no supporting retrieval (ungrounded generation)" likelihood: 90.0 {
  barrier "Abstain rule: no answer below a 0.62 retrieval score" effectiveness: 0.8 type: control
  barrier "Citation-required decoding, answers without a cited span are dropped" effectiveness: 0.7 type: prevention
}

threat "Regression introduced by a model or prompt change" likelihood: 12.0 {
  barrier "Blocking eval gate on the 1 200-case regression suite in CI" effectiveness: 0.85 type: prevention
  barrier "Shadow traffic comparison for 72 h before any promotion" effectiveness: 0.6 type: detection
}

consequence "Regulatory breach — unsuitable advice given under Consumer Duty" severity: 5 {
  barrier "Output classifier blocks regulated-advice language before send" effectiveness: 0.8 type: detection
  barrier "Sampled QA review of 2% of transcripts by the compliance team" effectiveness: 0.3 type: detection
  barrier "Breach register and 72 h regulator notification runbook" effectiveness: 0.25 type: recovery
}

consequence "Customer acts on wrong information and suffers financial loss" severity: 4 {
  barrier "Answer carries a citation the customer can open and check" effectiveness: 0.35 type: mitigation
  barrier "Goodwill remediation and complaint handling within 8 weeks" effectiveness: 0.5 type: recovery
}

consequence "PII of another customer disclosed in the response" severity: 5 {
  barrier "PII detector and redaction on the outbound response" effectiveness: 0.9 type: detection
  barrier "Tenant-scoped retrieval filter, index partitioned per customer" effectiveness: 0.85 type: control
  barrier "Data-breach response: notify the DPO within 24 h, ICO within 72 h" effectiveness: 0.2 type: recovery
}

consequence "Incident reaches media or social platforms" severity: 3 {
  barrier "Kill switch drops the assistant to a scripted fallback in under 5 min" effectiveness: 0.55 type: control
  barrier "Prepared holding statement and press escalation path" effectiveness: 0.4 type: recovery
}