Where should AI change the product, and what architecture and controls make that value dependable?

An AI-native product is not an established workflow with a chat box attached. It is designed around capabilities such as interpretation, generation, prediction or adaptive action, while acknowledging that outputs are probabilistic and may fail in unfamiliar ways. The commercial opportunity comes from changing the economics or quality of an outcome—not from displaying AI.

Autonomy should be earned progressively. Begin with assistance where the system proposes and a person decides; increase delegated action only when value, evaluation, observability, reversibility and accountability justify it.

Choose the right opportunity

Assess candidate use cases on two axes:

  • Value potential: frequency, delay, expertise bottleneck, variability and economic consequence.
  • dependability requirement: error severity, reversibility, data sensitivity, explainability, regulation and affected-party vulnerability.

Good early opportunities have meaningful repetitive knowledge work, accessible context, verifiable output and safe fallback. High-value but consequential decisions may still be appropriate, but they demand stronger evidence and human authority.

Avoid automating a broken process. Redesign the service so that AI, deterministic software and people each perform the work suited to them.

Product architecture beyond the model

The dependable product is a system containing:

  1. experience and consent: clear capability, limitations and control;
  2. context and retrieval: governed, current and permissioned information;
  3. model layer: selected for quality, latency, cost and deployment constraints;
  4. tools and actions: least-privilege access with confirmation for consequential steps;
  5. policy and orchestration: boundaries, workflow state and escalation;
  6. evaluation: task-specific offline and live evidence;
  7. observability: traces, outcomes, incidents, cost and drift;
  8. human operation: review, exception handling, support and redress.

Model capability is only one dependency. Supplier concentration, changing behaviour, intellectual property, security and inference economics belong in the business case.

A responsible autonomy ladder

  • Inform: retrieve or summarise; the user interprets.
  • Recommend: propose options with evidence; the user selects.
  • Draft: produce a reversible artefact for review.
  • Act with confirmation: prepare an action and require explicit approval.
  • Act within bounds: execute limited tasks under policies and monitoring.
  • Coordinate: plan and use multiple tools, with interruption and audit controls.

Promotion requires evidence at the current level plus a safety case for the next. The more consequential and irreversible the action, the stronger the confirmation, monitoring and redress.

Evaluation as product management

Build an evaluation set from real tasks, languages, edge cases, adversarial cases and known failures. Measure outcome quality, factual or evidential support, instruction adherence, harmful outputs, consistency, latency, cost and human correction. Analyse results by relevant population and market. Define release thresholds and regression tests.

NIST's AI RMF provides four connected functions—Govern, Map, Measure and Manage—for incorporating trustworthiness into design, development, use and evaluation.[1] The NIST Generative AI Profile extends the framework to generative AI risks.[2] These are management frames, not certifications that eliminate risk.

Microsoft's Responsible AI Standard illustrates how principles can be translated into goals, requirements and tools across a product lifecycle.[3] Its transparency reporting also reinforces that governance needs organisational accountability and continuing disclosure.[4]

Human control and service recovery

“Human in the loop” is meaningful only if the reviewer has information, competence, time, authority and a viable alternative. Design escalation triggers, pause and shutdown mechanisms, correction, appeal and remedy. Tell users when AI materially shapes an interaction or outcome and avoid anthropomorphic cues that overstate understanding.

For agents, apply least privilege, isolate credentials, restrict tools and destinations, validate inputs and outputs, require confirmation for financial, legal, external communication or destructive actions, and retain proportionate audit evidence.

Asia-Pacific and Hong Kong implications

Evaluation must cover languages and code-switching actually used in the service. Performance in English cannot establish performance in traditional Chinese, Cantonese-influenced input or regional terminology. Cross-border data, sector rules, model hosting and vendor terms may change architecture and market viability. Hong Kong can be a regional business hub, but governance must identify which obligations attach to the user, operator, data and affected market.

Economics and strategic advantage

Model total cost per successful outcome: inference, retrieval, evaluation, human review, exception handling, support and supplier risk. Falling unit model cost does not rescue a product that lacks distribution, proprietary context, workflow integration or trust. Sustainable advantage is more likely to come from a learning system—permissioned data, evaluated workflows, customer access and operational excellence—than from access to a general model.

Sources

  1. NIST, AI Risk Management Framework. https://airc.nist.gov/airmf-resources/airmf/
  2. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. https://doi.org/10.6028/NIST.AI.600-1
  3. Microsoft, “Microsoft's framework for building AI systems responsibly” and Responsible AI Standard resources. https://blogs.microsoft.com/on-the-issues/2022/06/21/microsofts-framework-for-building-ai-systems-responsibly/
  4. Microsoft, Responsible AI Transparency Report. https://www.microsoft.com/en-us/corporate-responsibility/topics/responsible-ai/reports/transparency-report/
  5. OECD and Eurostat, Oslo Manual 2018. Used to retain the distinction between invention and implemented innovation. https://doi.org/10.1787/9789264304604-en

Turn the research into a product decision.

Connect customer evidence, commercial logic and responsible delivery around the next commitment.

Discuss the decision