Method
How this study is designed
1. Two factors, not one list
An explanation condition is a pair. The content factor sets what is explained. The delivery factor sets how that material reaches the participant. Interactivity, adaptation and conversation are not kinds of explanation, they are ways of handing the same material over. Keeping them on their own axis is what makes it possible to say whether an effect came from the information or from the way it was given.
Factor A: content
- No explanation
- The recommendation on its own. The control.
- Why
- Attribution: how much each answer counted, as exact Shapley values for the network and as exact weights for the scorecard.
- What would change it
- Contrastive: the smallest single change to the situation that flips the outcome, found by re-running the advisor.
- How sure
- Uncertainty: the calibrated probability of the outcome, with the probability of every other outcome.
- All three
- The three contents above, together.
Factor B: delivery
- Static
- A panel, read as it is.
- Interactive
- The inputs can be moved and the advice reacts live, with a why-not selector.
- Adaptive
- Plain sentences or the detailed version, chosen by the measured financial literacy score.
- Conversational
- A language model in the browser retells the computed facts and answers follow-up questions.
A third factor runs alongside them: the advisor itself, a neural network whose explanations are computed after the decision, or an interpretable scorecard whose explanations are exact. It is assigned at random and logged, which turns explanation faithfulness into a measured variable rather than an assumption.
2. The cells this pilot fills
Crossing five contents with four deliveries gives twenty cells, far more than a pilot can fill. This is a fractional design: five cells vary content while delivery is held static, four cells vary delivery while the content is held at all three, and the all-three static condition sits in both arms as the hinge that ties them together.
| Content | Static | Interactive | Adaptive | Conversational |
|---|---|---|---|---|
| No explanation | No explanation | Interactive only | not run | not run |
| Why | Why | not run | not run | not run |
| What would change it | What would change it | not run | not run | not run |
| How sure | How sure | not run | not run | not run |
| All three | All three | Interactive with all three | Adaptive to literacy | Conversational |
Nine cells. Eight of them are in the random pool. The conversational cell is offered by choice only, because it needs a WebGPU browser and downloads a model, so assigning it at random would fail for part of the sample.
3. Which comparisons are interpretable
Clean
- Why, What would change it, How sure and All three, each against No explanation. Delivery is static throughout, so a difference is attributable to content.
- Interactive with all three, Adaptive and Conversational, each against All three static. The content is identical, so a difference is attributable to delivery.
- Interactive only against No explanation. Neither shows written explanation content, so this asks whether exploration on its own can do the work an explanation does.
Not reported
- Interactive only against All three static. The two differ on content and on delivery at the same time, so any difference between them cannot be assigned to either factor.
- Any comparison that crosses both arms without passing through the hinge condition.
4. The outcome is an interaction
Half of the recommendations each participant sees are deliberately shifted in the wrong direction while the explanation keeps describing the advisor's real reasoning. Appropriate reliance is following the sound ones and overriding the flawed ones. A condition that raises following on both is producing compliance, not calibration, which is why the quantity of interest is the condition by scenario interaction and never a main effect on trust.
Measures per trial
Decision (follow, adjust, reject, ask a human), the direction and size of an adjustment, trust, understanding, decision confidence, mental demand, decision time, the time spent reading the case, and the interaction traces: what-if moves, why-not questions and conversational turns.
Moderators
Financial literacy measured with the Big Three questions, self-rated financial knowledge, and the language the session ran in.
5. How the two strands fit together
This is a convergent design with an embedded qualitative strand. The experiment is the core, and the qualitative material is collected inside the same session rather than in a separate study, so every free-text answer is attached to a decision whose condition and scenario are known.
Quantitative core
Between participants: explanation condition and advisor. Within participants: six cases, half sound and half flawed. Outcome: appropriate reliance, with trust, understanding, confidence, demand and time alongside it.
Embedded qualitative
One reason per decision written in the moment, two reflective questions at the end, and the full transcripts of the conversational condition, which record what people ask when they can ask anything.
Integration
A joint display with one row per condition, reliance numbers beside the themes from that condition's reasons. The interesting findings live where the two disagree, for example a condition that raises trust and self-reported understanding while reliance gets worse.
6. Analysis plan and what gets excluded
- Appropriate reliance is modelled with mixed-effects logistic regression, condition by scenario, with random intercepts for the participant and for the case, and literacy entered as a moderator.
- Sessions where the explanation style was assigned at random are the experiment. Sessions where a participant chose their style are a separate stratum, analysed as a preference signal and never pooled with the random one.
- Excluded from the experimental analysis: sessions that fail the attention check, repeat sessions from the same browser, and responses on the advisor pages, which are marked as tryouts because the person chose their own style and wrote their own profile.
- This deployment is a pilot for developing the instrument. It is not powered for confirmatory tests, and pilot data is not for publication before a formal ethics review. The confirmatory plan is to run the content arm first and the delivery arm afterwards at the content level that arm selects.
See it from the inside
The fastest way to understand the design is to sit in one of its cells for ten minutes.