How can we automate multi-step work without losing control, evidence or accountability?

An AI agent can interpret a goal, choose a tool, inspect the result and decide what to do next. That makes it useful for work that changes from case to case: researching a supplier, preparing a campaign brief or investigating a support request. The business opportunity is substantial, but the first implementation should have a narrow job, a named owner and a clear stopping point.

The central recommendation is to automate a bounded process before attempting an autonomous department. Give the system enough flexibility to handle information, while keeping authority over money, publication and sensitive records in explicit application controls.

Recent developments and what they establish

On 9 September 2026, ABI Research described persistent, collaborative agents as an emerging enterprise direction, highlighting memory, model routing and orchestration. This is an analyst interpretation of industry development, not evidence that autonomous operations are already reliable across businesses. ABI Research announcement.

Alibaba launched Qwen Cloud for global markets in Singapore on 26 May 2026, according to its 28 May announcement. The platform brings model access and agent-oriented interfaces together. Its launch illustrates investment in agent infrastructure across Asia as well as Western markets; it does not establish that any particular workflow will deliver a return. Alibaba Cloud announcement.

Mistral announced Workflows in public preview on 27 April 2026, describing support for persistent execution, traceability and approval pauses. These are relevant operational capabilities, but the announcement's preview status should not be silently converted into a claim of general availability. Mistral Workflows.

Our assessment is that the important change is the surrounding execution system. A capable model becomes more useful when work can resume after failure, permissions remain enforceable and a person can inspect the outcome.

How the technology works

A conventional workflow follows predefined steps. An agent chooses some steps dynamically. A practical implementation often combines both: software decides which records the user may access, the model extracts and compares information, and software checks the proposed output before any action is taken.

For example, a research agent may decide to inspect a competitor's product page after finding a launch announcement. It should not decide to buy a subscription, contact that competitor or publish a conclusion unless the application explicitly allows that action. Reading and acting are different permission categories.

The essential components are a model, a set of narrowly defined tools, a record of workflow state, an evidence store and an approval interface. Memory is optional. Persistent memory should hold verified facts with an owner and expiry date, rather than every statement the model previously produced.

Practical example: a weekly competitor briefing

Consider an illustrative marketing agency monitoring five competitors for one client. Today, an analyst checks websites, records changes and prepares a slide. The proposed agent reads only approved public sources and produces a draft with links, dates and uncertainties. An account director approves the client-facing version.

The input is a competitor list, approved URLs, geographic scope, a comparison period and the client's product categories. The output contains three kinds of entries: verified change, unresolved claim and no relevant change. This prevents the system from inventing news merely to fill a weekly template.

Suppose a vendor advertises a new feature but its documentation still describes a limited preview. The correct entry is: feature announced; access restrictions remain; commercial availability needs confirmation. The analyst should see both sources. A polished sentence without that distinction is a failure.

The first version should save a draft in a review queue. It should not send emails or edit the client's website. This is a proposed implementation example, not a reported 8i client engagement.

Implementation plan

  1. Specify the decision. Write a one-page brief naming the audience, five competitors, permitted sources and the action the briefing will inform. Exclude unrelated news and speculative market-size estimates.
  2. Establish a baseline. Record how long three manual briefings take, which errors occur and how reviewers judge usefulness. Preserve examples of strong and weak outputs.
  3. Create an evidence record. Save the source URL, page title, publication date when available, retrieval time and exact supporting passage or permitted short excerpt. Keep announcement date separate from planned release date.
  4. Constrain tool access. Expose a read-only search and retrieval interface. Enforce URL and data permissions in software. Do not give the model a general-purpose credential with access to unrelated client systems.
  5. Validate the draft. Require each factual entry to map to evidence. Check dates, duplicate stories and unsupported numbers. Allow the model to say the evidence is insufficient.
  6. Add human approval. Show proposed conclusions alongside sources. Record the approver and changes. If eventual publication is enabled, bind approval to the exact version being published.
  7. Handle interrupted work. Persist completed steps, cap retries and make repeated execution safe. A restarted run must not create duplicate briefings or duplicate downstream actions.

Use a single agent initially. Add specialist agents only if a measured bottleneck justifies their coordination cost. Several models agreeing with each other do not constitute independent evidence.

A reusable task specification

Prepare a competitor-change briefing for the defined period using only the approved sources. For every material claim, include its source and date. Separate announced, preview and available capabilities. Distinguish fact from interpretation. If sources conflict, show the conflict. Do not contact anyone or publish anything. Return a draft for review and list unresolved questions.

This instruction improves clarity, but it is not a security boundary. Permissions and action restrictions must also exist outside the model.

Evaluation and economics

Evaluate completed, useful work rather than the number of tool calls. Count a briefing as successful only if it meets the agreed coverage, evidence and accuracy requirements. Report missed significant changes separately from invented changes: they create different business risks.

MeasureHow to calculate or assess it
Supported-claim rateClaims backed by correctly matched evidence divided by claims checked
Important-change recallRelevant changes found divided by a human-reviewed reference set
Review effortMedian reviewer minutes per accepted briefing
Cost per accepted briefingModel, retrieval, hosting and review cost divided by accepted briefings
Unauthorized actionsAny action outside the approved scope; investigate every occurrence

A worked planning example: 20 briefings at 90 manual minutes each require 30 hours. If a pilot actually measures 35 minutes of review and correction per briefing, the gross time released is about 18.3 hours. Subtract setup, maintenance and operating costs before calculating value. These numbers are illustrative assumptions, not a forecast or a vendor benchmark. Released hours become financial savings only if the business can redeploy capacity or avoid real expense.

Failure modes and controls

Untrusted pages can contain instructions intended to divert an agent. Treat source content as evidence to inspect, never as authorization. Stale pages can also produce wrong conclusions without any attack. Keep timestamps and recheck the underlying product documentation when a change matters.

Long-running jobs need time and spending limits. Set a maximum number of retrieval attempts and a clear escalation state. A run that exceeds its budget should return its evidence and unresolved question, rather than continue indefinitely.

Avoid automatic memory updates from draft findings. A false conclusion saved as permanent knowledge can contaminate later work. Require review before durable facts are added and provide a way to correct or remove them.

A realistic first month

Spend the first week on scope, source access and the manual baseline. In week two, run the system on historical periods whose answers are already known. Use week three for shadow operation alongside the analyst. In week four, review quality, review time and total cost before enabling any additional action.

Proceed when the output is consistently useful and the failure cases are visible and recoverable. Keep the process manual if exceptions dominate, evidence cannot be obtained lawfully or nobody can own the result. The first deliverable should be a dependable research process, not a demonstration of unrestricted autonomy.

Sources and scope

Evidence cutoff: 10 September 2026. This report combines dated public announcements with original implementation recommendations. Vendor capabilities and access conditions should be rechecked during procurement. Illustrative examples and targets are not measured customer outcomes.

  • ABI Research. “AI Claws Signal the Next Phase of Enterprise AI: Persistent, Collaborative Agents.” 9 September 2026. Source.
  • Alibaba Cloud Community. “Alibaba Cloud Launches Qwen Cloud for Global Markets.” 28 May 2026; launch occurred 26 May. Source.
  • Mistral AI. “Workflows for work that runs the business.” 27 April 2026. Source.

Turn the research into operating value.

Connect the use case, architecture, evidence, controls and operating model around a decision that matters.

Discuss the decision