GTM Experiment Decision Engine
Register one experiment, run the pre-flight feasibility check, and see what the evidence actually licenses — adopt, continue, revise, or stop.
Explore the adversarial examples or run one bounded custom read. The custom path deliberately holds the design architecture constant so this remains an evidence-decision demonstration, not a substitute for experiment design or statistical analysis.
This differs from the Decision-Ready Readout’s “Adopt” reading, which follows from period-over-period movement with no causal claim. An “Adopt” reading here requires an identified, mature, and adequately informative benefit — a stricter evidence standard for a different kind of question, not a generally stronger one.
Each scenario is a separate, self-contained fixture — switching does not carry anything over.
Executive decision record
A compact handoff for a budget, campaign, or operating review. It is derived from the same engine result below, not a second recommendation layer.
Decision
Adopt for closed-lost accounts eligible for re-engagement
Evidence
Assigning X caused an increase in new opportunities within 120 days, per assigned account of [2.66 to 6.34] (126 of 1200 versus 72 of 1200), which is at least the pre-declared meaningful effect Δ (2).
Here X is the treatment: SDR outreach sequence.
Why now
The evidence licenses this, with the named conditions attached.
Decision boundary
Revenue; the win rate of the incremental opportunities this created; whether the effect transfers to colder, future cohorts; or which part of the outreach-plus-offer bundle is doing the work.
Experiment contract
- What's being decided
- Fund a permanent closed-lost SDR motion, or notAlternatives: Fund permanently / Do not fund
- Who's in the test
- Closed-lost accounts eligible for re-engagement (2,400 accounts)
- What the treatment group got
- SDR outreach sequence
- What the comparison group got instead
- Business as usual: newsletter, inbound handling
- Unit of assignment
- Accounts
- How the comparison was built
- Account-level randomized test
- What was measured
- New opportunities within 120 days, per assigned account
- How big a change would matter
- 2 percentage points, on a baseline rate of 6%
- How long results were given to show up
- 120 days after assignment
- What must not get worse
- Do-not-contact requests per assigned account — must not rise by more than 2 percentage points
Validity
Design and identification
VALID
The design supports this reading as stated, with no unresolved validity problem capping it.
Validity findings
5 findings; none of them capped this reading.
Causal language
This reading's design tier permits a causal statement.
Guardrails
- HeldDo-not-contact requests per assigned account
Effect readout
4.5percentage points
- Scale
- [2.7, 6.3]
- Basis
- Randomized comparison — Account-level randomized test
- Threshold
- Δ = 2.0 percentage points
Estimated effect: 4.5 percentage points, likely between 2.7 and 6.3. The threshold that would count as meaningful is 2.0.
Dashed line marks Δ, the pre-declared meaningful effect.
A meaningful benefit is supported
This outcome is a pipeline metric.
Decision record
Evidence conclusion
Assigning X caused an increase in new opportunities within 120 days, per assigned account of [2.66 to 6.34] (126 of 1200 versus 72 of 1200), which is at least the pre-declared meaningful effect Δ (2).
Here X is the treatment: SDR outreach sequence.
Recommended action
Adopt for closed-lost accounts eligible for re-engagement
Why
The evidence licenses this, with the named conditions attached.
Conditions
- Keep reading the metric this decision actually rests on, not just the leading indicator.
- Keep a comparison group even after adopting, so the effect keeps being measured.
What this does not prove
Revenue; the win rate of the incremental opportunities this created; whether the effect transfers to colder, future cohorts; or which part of the outreach-plus-offer bundle is doing the work.
What would change the decision
A month-12 read where incremental opportunities close below roughly the 19% lower-bound break-even, or a holdout on the next cohort whose interval includes zero.
How it decided
Design class
Account-level randomized test
Rule that decided
Rule 15 of 17. A meaningful benefit is supported, precautions apply, and the change is cheap to reverse, so it can be adopted with conditions.
Feasibility
This design is underpowered to detect an effect of the minimum meaningful size.
- Standard error
- 0.970
- Minimum detectable effect
- 2.411
- Minimum detectable effect ÷ threshold Δ
- 1.21
- Expected events (treatment / comparison)
- 72.0 / 72.0
All validity findings
- Some assigned units never actually received the intervention.Changes what the estimate measures; resolved, does not change this reading.
- Some of the comparison group were exposed to the same thing as the test group.Changes what the estimate measures; resolved, does not change this reading.
- Sales knew which arm a unit was in and could act on that knowledge.Changes what the estimate measures; detected.
- The outcome's horizon is no longer than the intervention's own exposure window.Changes what the estimate measures; detected.
- This design was underpowered going in, so an exciting-looking result is more likely to be noise.Disclosed alongside the reading; detected.
What this instrument is
Every scenario here is a portfolio-built demonstration: an authored fixture, not a live experiment or real campaign data. What it produces is a modeled reading of that fixture, not a record of something that happened. This instrument shows how a decision-grade evidence standard would be applied to a designed comparison — it is not production causal-inference software, and the actions it recommends (“Adopt”, “Continue”, “Revise”, “Stop”) are not claims that any real intervention has been proven to work. They describe what one experiment’s evidence licenses for one population, nothing more.
Nothing leaves your browser: there is no signup, storage, or analytics here, same as every other instrument. The point estimates and intervals a scenario declares are accepted as inputs, not computed from raw data — this instrument does not estimate effects itself.
A recommendation here is not an authority assignment (it says nothing about who may act or approve), not a budget allocation, not a portfolio decision among competing experiments, and not an instruction to a production system. It is not the decision itself — a human decision may reasonably differ for reasons outside the evidence, such as strategy, capacity, or timing.
Every specific number a scenario shows (Δ, tolerance margins, the low-information cutoff) is that fixture’s own declared value. The engine’s own thresholds are illustrative pending further calibration, not validated statistical standards.