AdviceIT

Method

How this study is designed

AdviceIT varies two things independently: what an explanation says, and how it reaches you. This page states the design, the cells this pilot fills, which comparisons are interpretable, and how the numbers and the free text are meant to be read together.

1. Two factors, not one list

An explanation condition is a pair. The content factor sets what is explained. The delivery factor sets how that material reaches the participant. Interactivity, adaptation and conversation are not kinds of explanation, they are ways of handing the same material over. Keeping them on their own axis is what makes it possible to say whether an effect came from the information or from the way it was given.

Factor A: content

No explanation
The recommendation on its own. The control.
Why
Attribution: how much each answer counted, as exact Shapley values for the network and as exact weights for the scorecard.
What would change it
Contrastive: the smallest single change to the situation that flips the outcome, found by re-running the advisor.
How sure
Uncertainty: the calibrated probability of the outcome, with the probability of every other outcome.
All three
The three contents above, together.

Factor B: delivery

Static
A panel, read as it is.
Interactive
The inputs can be moved and the advice reacts live, with a why-not selector.
Adaptive
Plain sentences or the detailed version, chosen by the measured financial literacy score.
Conversational
A language model in the browser retells the computed facts and answers follow-up questions.

A third factor runs alongside them: the advisor itself, a neural network whose explanations are computed after the decision, or an interpretable scorecard whose explanations are exact. It is assigned at random and logged, which turns explanation faithfulness into a measured variable rather than an assumption.

2. The cells this pilot fills

Crossing five contents with four deliveries gives twenty cells, far more than a pilot can fill. This is a fractional design: five cells vary content while delivery is held static, four cells vary delivery while the content is held at all three, and the all-three static condition sits in both arms as the hinge that ties them together.

ContentStaticInteractiveAdaptiveConversational
No explanationNo explanationInteractive onlynot runnot run
WhyWhynot runnot runnot run
What would change itWhat would change itnot runnot runnot run
How sureHow surenot runnot runnot run
All threeAll threeInteractive with all threeAdaptive to literacyConversational

Nine cells. Eight of them are in the random pool. The conversational cell is offered by choice only, because it needs a WebGPU browser and downloads a model, so assigning it at random would fail for part of the sample.

3. Which comparisons are interpretable

Clean

  • Why, What would change it, How sure and All three, each against No explanation. Delivery is static throughout, so a difference is attributable to content.
  • Interactive with all three, Adaptive and Conversational, each against All three static. The content is identical, so a difference is attributable to delivery.
  • Interactive only against No explanation. Neither shows written explanation content, so this asks whether exploration on its own can do the work an explanation does.

Not reported

  • Interactive only against All three static. The two differ on content and on delivery at the same time, so any difference between them cannot be assigned to either factor.
  • Any comparison that crosses both arms without passing through the hinge condition.

4. The outcome is an interaction

Half of the recommendations each participant sees are deliberately shifted in the wrong direction while the explanation keeps describing the advisor's real reasoning. Appropriate reliance is following the sound ones and overriding the flawed ones. A condition that raises following on both is producing compliance, not calibration, which is why the quantity of interest is the condition by scenario interaction and never a main effect on trust.

Measures per trial

Decision (follow, adjust, reject, ask a human), the direction and size of an adjustment, trust, understanding, decision confidence, mental demand, decision time, the time spent reading the case, and the interaction traces: what-if moves, why-not questions and conversational turns.

Moderators

Financial literacy measured with the Big Three questions, self-rated financial knowledge, and the language the session ran in.

5. How the two strands fit together

This is a convergent design with an embedded qualitative strand. The experiment is the core, and the qualitative material is collected inside the same session rather than in a separate study, so every free-text answer is attached to a decision whose condition and scenario are known.

Quantitative core

Between participants: explanation condition and advisor. Within participants: six cases, half sound and half flawed. Outcome: appropriate reliance, with trust, understanding, confidence, demand and time alongside it.

Embedded qualitative

One reason per decision written in the moment, two reflective questions at the end, and the full transcripts of the conversational condition, which record what people ask when they can ask anything.

Integration

A joint display with one row per condition, reliance numbers beside the themes from that condition's reasons. The interesting findings live where the two disagree, for example a condition that raises trust and self-reported understanding while reliance gets worse.

6. Analysis plan and what gets excluded

  • Appropriate reliance is modelled with mixed-effects logistic regression, condition by scenario, with random intercepts for the participant and for the case, and literacy entered as a moderator.
  • Sessions where the explanation style was assigned at random are the experiment. Sessions where a participant chose their style are a separate stratum, analysed as a preference signal and never pooled with the random one.
  • Excluded from the experimental analysis: sessions that fail the attention check, repeat sessions from the same browser, and responses on the advisor pages, which are marked as tryouts because the person chose their own style and wrote their own profile.
  • This deployment is a pilot for developing the instrument. It is not powered for confirmatory tests, and pilot data is not for publication before a formal ethics review. The confirmatory plan is to run the content arm first and the delivery arm afterwards at the content level that arm selects.

See it from the inside

The fastest way to understand the design is to sit in one of its cells for ten minutes.