Bowtie — unsafe or non-compliant LLM response to a customer
Barrier-based bowtie for a customer-facing LLM assistant emitting unsafe or non-compliant output, with preventive barriers against injection, jailbreak, stale retrieval and ungrounded generation and recovery barriers for regulatory breach and PII disclosure, used in an AI safety case or Consumer Duty review.
Make it your own.
title "Assistant emits an unsafe or non-compliant response to a customer"
hazard "A general-purpose LLM with authority to state policy to 1.2 M retail banking customers"
top "Unsafe or non-compliant response reaches a customer"
unit "/yr"
threat "Indirect prompt injection in an uploaded statement or ticket attachment" likelihood: 6.0 {
barrier "Attachment text stripped of instruction-like spans before prompting" effectiveness: 0.6 type: prevention
barrier "Injection classifier on all retrieved and uploaded text" effectiveness: 0.85 type: detection {
escalation "Classifier not retrained after the last two jailbreak families" {
control "Weekly red-team corpus refresh gates the classifier release"
}
}
barrier "Untrusted content rendered to the model inside a quoted, non-executable block" effectiveness: 0.5 type: control
}
threat "Direct jailbreak of the chat surface by an end user" likelihood: 40.0 {
barrier "System prompt hardening plus refusal training on the 340-case suite" effectiveness: 0.9 type: prevention
barrier "Per-session jailbreak scoring with step-up to a human agent" effectiveness: 0.7 type: detection
}
threat "Retrieval returns a superseded policy document" likelihood: 24.0 {
barrier "Index freshness SLA: reindex within 4 h of a policy publish" effectiveness: 0.8 type: prevention {
escalation "Policy team publishes outside the CMS, so no reindex event fires" {
control "Monthly reconciliation of CMS document IDs against index IDs"
}
}
barrier "Effective-date filter on every retrieval query" effectiveness: 0.75 type: control
}
threat "Model answers with no supporting retrieval (ungrounded generation)" likelihood: 90.0 {
barrier "Abstain rule: no answer below a 0.62 retrieval score" effectiveness: 0.8 type: control
barrier "Citation-required decoding, answers without a cited span are dropped" effectiveness: 0.7 type: prevention
}
threat "Regression introduced by a model or prompt change" likelihood: 12.0 {
barrier "Blocking eval gate on the 1 200-case regression suite in CI" effectiveness: 0.85 type: prevention
barrier "Shadow traffic comparison for 72 h before any promotion" effectiveness: 0.6 type: detection
}
consequence "Regulatory breach — unsuitable advice given under Consumer Duty" severity: 5 {
barrier "Output classifier blocks regulated-advice language before send" effectiveness: 0.8 type: detection
barrier "Sampled QA review of 2% of transcripts by the compliance team" effectiveness: 0.3 type: detection
barrier "Breach register and 72 h regulator notification runbook" effectiveness: 0.25 type: recovery
}
consequence "Customer acts on wrong information and suffers financial loss" severity: 4 {
barrier "Answer carries a citation the customer can open and check" effectiveness: 0.35 type: mitigation
barrier "Goodwill remediation and complaint handling within 8 weeks" effectiveness: 0.5 type: recovery
}
consequence "PII of another customer disclosed in the response" severity: 5 {
barrier "PII detector and redaction on the outbound response" effectiveness: 0.9 type: detection
barrier "Tenant-scoped retrieval filter, index partitioned per customer" effectiveness: 0.85 type: control
barrier "Data-breach response: notify the DPO within 24 h, ICO within 72 h" effectiveness: 0.2 type: recovery
}
consequence "Incident reaches media or social platforms" severity: 3 {
barrier "Kill switch drops the assistant to a scripted fallback in under 5 min" effectiveness: 0.55 type: control
barrier "Prepared holding statement and press escalation path" effectiveness: 0.4 type: recovery
}