Design a business experiment that could change a decision
Use only information supplied in the conversation and this skill text. Do not browse, call tools, inspect files, execute code, create artifacts, or take external actions. Return the work directly in chat. Attribute material claims to supplied source labels or quotations; a pasted URL is a source label, not evidence that its contents were checked. Distinguish supplied facts, reasonable interpretations, proposals, and unknowns.
Inputs and scope: Gather the decision, proposed change, target population, current process, evidence, practical success threshold, available participants, time, constraints, and any reported results. Numerical interpretation requires supplied results. If a decision-changing input is absent, ask the smallest useful question and complete the portions supported by available material. State assumptions explicitly; do not manufacture facts, approvals, dates, or completion evidence.
Method 1. State the decision and testable hypothesis, including the mechanism and observations that would contradict it. Separate a feasibility pilot from a causal effectiveness test and from a commercial rollout. 2. Define the unit of assignment and analysis, treatment, comparison, and eligible population. Consider random allocation when feasible; identify selection bias, spillovers, timing effects, and confounders when it is not. 3. Specify a primary outcome aligned with the business decision, measurement window, practical importance threshold, and guardrails. Separate leading indicators from final outcomes and planned analyses from exploratory questions. 4. Describe implementation fidelity, data capture, review responsibilities, and treatment of missing or incomplete observations. Blinding may reduce some biases when practicable; it is not a universal guarantee of validity. 5. Plan the comparison and decision rule before results. Explain which inputs a qualified analyst would need for sample-size or power planning; do not prescribe universal minimum samples, choose tests solely by small n or normality, or compute complex statistics. 6. Define how results will be interpreted alongside cost, uncertainty, adverse effects, and generalizability. A p-value concerns incompatibility with a specified null model under assumptions, not the probability an effect exists; Bayesian analysis also remains vulnerable to selective choices.
Output: Return a test brief, hypotheses and disconfirmers, design and comparison, outcome definitions, proposed decision criteria, validity risks, and information needed before execution.
Quality checks: Do not claim experiments ran, diagnostics passed, or statistical significance from unsupplied data. Use formal evidence-grading standards only when their criteria are supplied; prioritize decision usefulness over reference or figure quotas.
Worked example: A service team wants to test a revised reminder email. Propose comparing eligible accounts over the same period, measuring payment within seven days and complaint rate. If managers choose which accounts receive it, flag selection bias and limit causal claims. A favorable walkthrough shows plausibility, not measured uplift; keep exploratory segment comparisons separate from the primary outcome.