How do we restrict AI authority, detect failures and preserve a reliable path to containment?

AI security has two related dimensions. Organizations can use AI to assist defensive work, and they must secure the AI systems they deploy. A model that reads documents, accesses customer data or operates tools becomes part of the business's security boundary. Its permissions and outputs need the same operational discipline as other software, plus controls for AI-specific failure modes.

For a first project, improve visibility and restrict authority before adding more autonomy. Knowing which agents exist, what they can access and who owns them is a practical prerequisite for trust.

Recent developments and their limits

Google introduced Gemini 3.8 Flash Cyber on 2 September 2026, positioning it for vulnerability discovery and defensive patching. Access is restricted to trusted defenders through the Fairwind Program. This should not be described as a universally available security tool, and vendor benchmark results do not establish performance on an organization's codebase. Google announcement.

Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, describing options involving customer-controlled storage and misuse review. The announcement says rollout will occur in phases later in the fall; some controls are opt-in. It is an announced architecture and rollout, not evidence that every enterprise account already has the same configuration. Anthropic announcement.

OWASP's 2025 guidance identifies prompt injection as a major application risk, including instructions embedded in external content. It also explains that RAG and fine-tuning do not fully remove that vulnerability. OWASP guidance.

Our assessment is that stronger models increase the value of secure integration. A system that can take more useful actions can also cause more damage when it misinterprets a request or follows untrusted content.

The main concepts in plain language

Prompt injection occurs when content the model reads attempts to redirect its behavior. For example, an uploaded document may contain text telling an assistant to ignore its assigned task. The document is data; it is not an authorized operator.

Excessive authority is a separate problem. If a summarization assistant holds credentials that can export an entire customer database, a single mistake can have a large effect. A warning in its prompt does not remove that permission.

Assurance means collecting evidence that the complete system behaves acceptably for its intended use. It includes testing, monitoring, ownership and incident response. It is not the same as a vendor's safety benchmark or a model declaring that its answer is safe.

Practical example: secure an agency research assistant

An illustrative agency uses an assistant to summarize public competitor pages and internal client notes. A security review discovers that the same integration token can read every client folder and send external emails.

The redesign separates clients, limits retrieval to the requesting user's permissions and removes email-sending access from the research stage. A separate publication service receives only an approved final artifact. Its approval record identifies the exact content and intended recipient.

Testing includes a harmless simulated document that asks the assistant to perform an out-of-scope action. The expected result is no external action and an auditable rejection or escalation. The test uses synthetic data in a controlled environment; it is not an attempt against a third party.

The outcome is a smaller and more inspectable risk boundary. This scenario is an original example, not a claim about an actual security incident.

Implementation plan

  1. Inventory the systems. Record each AI application, model provider, owner, data source, connector and action permission. Include browser extensions and experiments that may have entered routine use.
  2. Classify the consequences. Separate drafting from external communication, record modification, financial action and access administration. Apply stronger controls where errors are harder to reverse.
  3. Reduce permissions. Give each service only the tools and data it needs. Keep credentials outside prompts and model-generated content. Avoid shared accounts spanning clients.
  4. Validate actions in software. Check identity, allowed operation, target, parameters and current approval before execution. Reject unexpected fields or destinations.
  5. Build a focused test set. Include unauthorized data requests, misleading source content, malformed tool results, stale approvals and repeated execution after timeout.
  6. Monitor meaningful events. Record tool use, state changes, approval decisions and exceptions with appropriate redaction. Logs should support investigation without becoming an uncontrolled copy of sensitive data.
  7. Rehearse containment. Demonstrate how to stop the workflow, revoke its access, preserve evidence and return to a manual process. Assign a person who can make that decision.

These steps are original operational recommendations. The precise control set should be proportionate to the application rather than copied mechanically into every low-risk assistant.

A useful permission specification

CapabilityResearch stageApproved publication stage
Read public sourcesAllowed through controlled retrievalUsually unnecessary
Read client recordsLimited to authorized client and userOnly approved final content
Send external messagesNot allowedExplicit recipient and content approval
Change access rightsNot allowedNot allowed
Store long-term memoryReviewed facts onlyPublication audit record only

This is a design example. Permission enforcement belongs in the application and identity systems, not solely in the model's instruction text.

Evaluating defensive AI

If AI is used to review code or summarize security alerts, measure the quality of verified findings. More alerts are not necessarily better. A high false-positive rate can consume analysts' attention and hide important issues.

Use an approved test environment and known cases. Evaluate whether a suggested fix resolves the defect, preserves intended behavior and introduces no new issue in the tested scope. Keep final release authority with the responsible engineering team.

Distinguish three outcomes: valid finding, invalid finding and unresolved finding. The unresolved category prevents uncertain model output from being forced into an unjustified yes-or-no judgment.

Metrics and release gates

MeasurePractical use
Unauthorized-action attempts blockedConfirms that application controls reject out-of-scope operations
Cross-client access testsChecks isolation of confidential material
Valid finding rateReveals analyst workload from false alarms
Verified remediation timeMeasures time to a tested fix, not a generated suggestion
Incident containment timeTests whether the team can stop and investigate a failure

For a pilot, define a release gate that requires every known critical boundary test to pass. This is a necessary condition, not proof that all attacks are impossible. Include fresh tests as the application gains new tools or data sources.

Security evaluation also has a cost. Budget for the system owner, reviewers, monitoring and incident handling. Avoid treating these as optional overhead after a successful demonstration.

Common mistakes

A private deployment is not automatically secure. It can still leak data through logs, unsafe connectors or excessive internal permissions. Likewise, a provider's retention policy does not determine how the customer's own application stores transcripts.

Human approval can become ceremonial if reviewers cannot see the exact action or are asked to approve too many low-value prompts. Present clear changes and consequences, and reserve approval for decisions that benefit from it.

Do not use a model's confidence score as the only security gate. A confidently generated action can still be unauthorized. Authorization is a rule to enforce, not a prediction to make.

First-month recommendation

Use week one to inventory applications and remove unnecessary access. In week two, implement action validation and client isolation. Use week three for controlled tests and log review. In week four, run an incident exercise and document the release decision.

The goal is a system whose authority is limited, behavior is visible and failures can be contained. Stronger model safeguards are helpful, but the organization remains responsible for the workflow it connects to the model.

Sources and scope

Evidence cutoff: 10 September 2026. This report focuses on defensive architecture and operational assurance. It does not claim compliance with a specific law or certification. Product access and future rollout claims are explicitly qualified.

  • Tulsee Doshi and Raluca Ada Popa, Google. “Introducing Gemini 3.8 Flash and 3.8 Flash Cyber.” 2 September 2026. Source.
  • Anthropic. “Developing Enterprise Frontier Safeguards with our customers.” 1 September 2026. Source.
  • OWASP Gen AI Security Project. “LLM01:2025 Prompt Injection.” 2025 edition; accessed 10 September 2026. Source.

Turn the research into operating value.

Connect the use case, architecture, evidence, controls and operating model around a decision that matters.

Discuss the decision