What is the cheapest credible evidence needed before the next commitment?
Experimentation is a capital-allocation discipline. Its purpose is not to prove a favoured concept but to reduce a consequential uncertainty before more money, time or reputation is committed. A prototype is valuable only when its fidelity matches the question and its observed evidence can change a decision.
The strongest product systems combine qualitative discovery, prototypes, commercial tests, operational pilots and controlled experiments. No single method answers desirability, usability, viability, feasibility and responsibility at once.
Start with the decision and uncertainty
Frame every experiment with:
- the decision that will follow;
- the assumption that could invalidate the concept;
- the observable result and threshold;
- the population and context;
- the exposure, ethical and operational risks;
- the action for positive, negative and ambiguous outcomes.
This reverses the common sequence of building something and then deciding what the resulting numbers mean. Pre-specified criteria reduce motivated interpretation.
Match evidence to the question
Use an evidence ladder:
- Problem evidence: observation and interviews establish the problem in context.
- Comprehension evidence: concept artefacts test whether people understand the idea.
- Behavioural intent: realistic choices, sign-ups or deposits provide stronger evidence than stated enthusiasm.
- Usability evidence: interactive prototypes reveal whether intended users can complete critical tasks.
- Delivery evidence: concierge or limited pilots test operational feasibility and willingness to pay.
- Causal evidence: controlled experiments estimate the effect of a change at sufficient scale.
- Outcome evidence: longitudinal measures test retention, economics, safety and realised value.
Moving upward costs more. Teams should stop as soon as the current decision has adequate evidence, yet avoid treating low-cost signals as proof of high-stakes outcomes.
Trustworthy experimentation
Microsoft's experimentation guidance stresses that trustworthy results depend on data quality, statistical practice and an experimentation culture, not merely an A/B testing tool.[1][2] Booking.com's published account similarly describes organisational experimentation at very large scale.[3] Netflix's experimentation platform shows how metrics, analysis and platform design combine to support decisions.[4]
For business leaders, five controls matter:
- a primary outcome and guardrail metrics are defined in advance;
- sample, duration and stopping rules are appropriate;
- instrumentation and assignment are checked;
- novelty, seasonality and segment effects are considered;
- results include uncertainty and practical significance, not only a binary statistical label.
Not every question permits randomisation. Quasi-experimental analysis, phased rollout and comparative pilots can be useful, but causal claims must match the design.
Experiment portfolio and governance
Maintain an experiment register containing hypothesis, decision, owner, evidence class, exposure, safeguards, result and reusable learning. Review the portfolio for duplication and blind spots. Failed hypotheses are useful when the test was valid and the learning is retained; repeated unrecorded experiments are waste.
Use higher approval thresholds for experiments involving vulnerable users, personal data, pricing, credit, employment, health or automated consequential decisions. Quality assurance must be continuous across development and operation, as the UK Government Service Manual recommends.[5]
Asia-Pacific and Hong Kong implications
Evidence can be highly context-dependent. A successful test in one economy, language or channel should not be assumed to transfer across Asia-Pacific. Run small replication tests in priority markets and record local differences in acquisition channels, device use, trust, payments and regulation. In Hong Kong, bilingual stimuli and locally realistic customer journeys are essential when language changes comprehension or behaviour.
AI-native experimentation
AI can generate prototype variants and analyse research more quickly, increasing both learning capacity and false-positive risk. Log model/version, prompt or policy, evaluation dataset and human review. Test non-deterministic systems repeatedly, include adversarial and edge cases, and measure factuality, harmful outcomes, override and escalation—not only task completion. The NIST AI RMF's Govern, Map, Measure and Manage functions provide a practical risk frame.[6]
Sources
- Microsoft Research, “Online Experimentation at Microsoft.” https://www.microsoft.com/en-us/research/publication/online-experimentation-at-microsoft/
- Microsoft Research, “Patterns of Trustworthy Experimentation: During-Experiment Stage.” https://www.microsoft.com/en-us/research/articles/patterns-of-trustworthy-experimentation-during-experiment-stage/
- Booking.com, “Online Experimentation at Booking.com.” https://arxiv.org/abs/1710.08217
- Netflix Technology Blog, “It's All A/Bout Testing: The Netflix Experimentation Platform.” https://medium.com/netflix-techblog/its-all-a-bout-testing-the-netflix-experimentation-platform-4e1ca458c15
- UK Government Service Manual, “Quality assurance: testing your service regularly.” https://www.gov.uk/service-manual/technology/quality-assurance-testing-your-service-regularly
- NIST, AI Risk Management Framework. https://airc.nist.gov/airmf-resources/airmf/
Turn the research into a product decision.
Connect customer evidence, commercial logic and responsible delivery around the next commitment.
