An AI assistant may read customer messages, internal documents and external content while also following instructions from its operator. Prompt injection occurs when untrusted content tries to make the system ignore those intended instructions, expose information or take an unauthorized action.

01

The risk enters through content the system is meant to read

A malicious instruction can appear in a user message, uploaded document, retrieved webpage, email or connected data source. Because a language model processes both instructions and content as text, it may not reliably recognize which words are authoritative on its own.

Indirect prompt injection is especially important for assistants that search or summarize outside material. The person using the tool may never see the hidden instruction that influenced the response.

02

Architecture limits what a manipulated response can reach

Banks separate untrusted content from protected instructions and give the assistant only the data and tools needed for the approved use. Access controls are enforced by surrounding systems rather than trusting the model to remember every restriction.

High-impact actions such as changing account data, releasing a payment or sending a customer communication can require deterministic validation and human approval. A model may propose an action without having authority to commit it.

03

Grounding and output checks reduce unsupported behavior

Retrieval can be limited to approved, permission-aware sources, with citations that let the employee verify the answer. Input and output filters may identify known attack patterns, requests for secrets or content that falls outside the assistant's purpose.

No filter is complete. Controls therefore combine source trust, data minimization, tool restrictions, transaction limits and refusal paths instead of treating one classifier or system prompt as a security boundary.

04

Adversarial testing follows the real workflow

Testing includes direct and indirect injection attempts, conflicting documents, encoded instructions, poisoned retrieval content and requests that combine several individually permitted steps into an impermissible result. The test should use the same tools and permissions available in production.

Teams record which attacks succeeded, whether protected data or actions were reachable and how reliably the system escalated uncertainty. Changes to the model, prompt, knowledge source or tool connection can require renewed testing.

05

Monitoring and containment assume attempts will continue

Logs connect retrieved content, model output, tool requests, approvals and final actions without unnecessarily exposing sensitive data. Unusual tool sequences, repeated refusals or attempts to access unrelated information can trigger review.

The bank maintains a way to disable a connector, revoke credentials, isolate a compromised source and route work to a person. Prompt-injection defense is therefore an operating capability, not a one-time instruction added before launch.

Sources

Read the primary material

Banking Explained prioritizes regulators, official publications and first-party announcements.