Two advisors, same expert data
A neural network and an interpretable scorecard, both trained on the ILS-Bench expert consensus. One opaque, one transparent, so explanation faithfulness becomes a variable.
AdviceIT is an open study on which explanations help people follow AI advice when it is sound and push back when it is not. Try the advisors, then give ten minutes to the research.
No account, no personal data, hypothetical cases only.
Recommended outcome
Growth
What drove this
How a session works
Random assignment keeps the research clean. Choosing one yourself is allowed, and recorded as your choice.
Each case comes with the advisor's recommendation and the explanation style you were given.
Follow it, adjust it, reject it or send it to a human adviser. Those decisions are the data.
The question
Robo-advisors already manage real money, and explainable AI promises to make their advice understandable. But an explanation is only useful if it calibrates trust: helping you follow advice when it is sound and push back when it is flawed. Which explanation styles actually do that is an open question, and it is the question behind this instrument, built as a follow-up to a systematic literature review on trust and algorithm aversion in AI financial advice (SSRAAI 2026).
Two advisors power the study, both trained on ILS-Bench, a benchmark of 400 investor cases validated by a panel of four financial-domain experts. One is a neural network whose explanations must be computed after the fact. One is an interpretable scorecard whose explanations are exact. Comparing them turns explanation faithfulness itself into something we can measure.
Your part
Models can be benchmarked automatically. Trust cannot. Whether an explanation helps a person rely on advice appropriately can only be learned from people making decisions, which is exactly what the study session records: you read short hypothetical cases, see the advisor's recommendation with one explanation style, and tell us what you would do.
Everything is anonymous. No name, email or account data is collected, the cases are hypothetical, no real money is involved, and some recommendations are deliberately altered so that appropriate reliance can be measured at all. You are told which ones at the end.
Inside the instrument
A neural network and an interpretable scorecard, both trained on the ILS-Bench expert consensus. One opaque, one transparent, so explanation faithfulness becomes a variable.
The experts refused to automate almost half of the cases. Both advisors learned that, and can answer Human review instead of a portfolio.
Move the inputs, switch inputs off, ask why not another outcome. Every preview is a real re-run of the model, nothing is faked.
An open-weight language model runs on your GPU through WebLLM, grounded only on the computed facts, and can read a free-text description into the form.
Some study recommendations are deliberately shifted the wrong way while the explanation stays honest. Noticing the mismatch is the skill the study measures.
The training data, the cross-validated results, the confusion matrix and every one of the 400 cases are on the Training data page. The scorecard's weights are printed in full.
Six short cases, one explanation style, a debrief at the end.
Take part in the study