AdviceIT by Radit

Explainable AI investment advice and how people rely on it

The AI advisor

About the training data

This page runs the neural-network advisor: a small multilayer perceptron trained on ILS-Bench, 400 investor cases whose suitability labels and recommended outcome were validated by a panel of four financial-domain experts. It maps the three suitability labels (risk tolerance, risk capacity, liquidity need) plus age to one of six outcomes, the five portfolios or Human review, with percent cross-validated accuracy and calibrated probabilities. Its weights are not human-readable, so the explanations on this page are computed post hoc (Shapley values, model-agnostic counterfactuals, calibrated probability). The interpretable rule-based advisor is a scorecard fitted on the same data whose explanations are exact, so the same explanation styles can be compared between an opaque and a transparent model learned from the same expert judgements.

The interpretable rule-based advisor

About the training data

This page runs a scorecard derived from data: a multinomial logistic regression fitted on ILS-Bench, the same 400 expert-validated cases and the same twelve inputs as the neural network, with percent cross-validated accuracy and calibrated probabilities. It is rule-based in the sense that every outcome has one weight per input and the outcome with the largest total wins, exactly like the points-based suitability scorecards used in practice, except that its weights were learned from expert judgements rather than set by hand. Because the weights are visible, the contribution of every input to the recommended outcome is read directly from them, exact and additive. It is the transparent counterpart of the AI advisor: same data, same inputs, no black box.

The whole scorecard: what every input is worth, in points, for every outcome

These are the model's fitted weights, shown as they are. For any profile, exactly one row applies from each group. Add the applying rows, the age effect and the starting points in each column: the outcome with the largest total wins, and the totals are turned into the probabilities shown in the app. Points are log-odds, so a difference of about 1 point between two outcomes means roughly 3 to 1 odds. The "Why" explanation under a recommendation is this same table applied to one profile, relative to the neutral baseline.

Researcher controls

Customise
Content
Delivery

Content is what is explained. Delivery is how it is shown: as is, with steering controls, adapted to the literacy score, or as a grounded conversation. A preset is one combination, and the log records both parts.

Scenario

Mode

Researcher mode. Open a participant link (one trial) or open the full study flow (consent, literacy check, six trials, attention check, debrief).

Investor profile

Answer the questions a robo-advisor would ask.

18 to 80.

How many years until you expect to need this money.

Risk tolerance
Emergency fund covering at least 6 months of expenses
Income stability
Significant debt or large fixed obligations

For example high-interest debt, or fixed expenses that leave little room. Lowers risk capacity.

Could you need this money much sooner than planned?

For example rent, tuition or a tax bill within a year or two. Makes the liquidity need urgent whatever the horizon.

Financial knowledge, self-rated

Recorded only. This does not affect the recommendation. In a study it is a measured moderator variable (financial literacy).

Financial literacy check (three questions)

The "Big Three" questions of Lusardi and Mitchell. Recorded only, they do not affect the recommendation. They give the literacy score that the adaptive delivery uses to choose the plain or the detailed explanation, so this check appears with that delivery. In the full study flow it is asked at the start for every participant, as the moderator variable.

1. Suppose you had 100 in a savings account and the interest rate was 2 percent per year. After 5 years, how much do you think you would have in the account if you left the money to grow?

2. Imagine that the interest rate on your savings account was 1 percent per year and inflation was 2 percent per year. After 1 year, how much would you be able to buy with the money in this account?

3. Buying a single company's stock usually provides a safer return than a stock mutual fund.

Literacy score: not answered.

Or describe your situation in your own words

Optional. The in-browser language model reads the description and fills in the fields above, following the ILS-Bench language-to-suitability procedure. Check the fields before continuing.

AI recommendation

Recommended outcome

Balanced

    Your response

    Would you follow this advice?
    A few more questions about this decision

    Session

    0 response(s) recorded in this browser. Stored locally so a reload does not lose them.

    Time Participant Condition Advisor Scenario Profile Shown Score Margin Trust Decision Time (ms)

    No responses yet.

    Study design

    Research question

    Does explanation style (none, feature-based, counterfactual, confidence) affect appropriate reliance on AI investment advice, and does financial literacy moderate the effect?

    Design

    Between-subjects on explanation condition, defined by two parts: content (why, what would change it, how sure, in any combination or none) and delivery (static, interactive what-if, adaptive to literacy, conversational). Eight named presets cover the common combinations. Advisor type (neural network or interpretable rule-based advisor, both learned from the same expert data) is a second between-subjects factor that manipulates explanation fidelity with the origin of the rules held constant. A study picks a subset of conditions, for example the four single-content presets and the hybrid, by two advisors. Each participant completes sound and flawed advice trials in a seeded random order (the study flow uses six fixed hypothetical cases, half sound and half flawed, with one attention check). Financial literacy, measured with the Lusardi and Mitchell "Big Three", is a moderator. A within-subjects design with counterbalanced condition order is a viable alternative if recruitment is limited.

    Variables

    Independent variables: explanation content, explanation delivery, advisor type.

    Dependent variables: appropriate reliance (follow rate on sound advice, override rate on flawed advice, where override is Adjust, Reject or Ask a human adviser), trust rating (1 to 7), decision time. When the participant adjusts, the chosen portfolio and the number of steps from the shown one are recorded, so reliance is also available on a scale. Secondary: perceived understanding, decision confidence, mental demand (one NASA-TLX style item), a free-text reason, and in the interactive condition the number of what-if moves and "why not" questions. Human review is a possible outcome of every advisor and "Ask a human adviser" a possible decision, so deferral is measured on both sides.

    How appropriate reliance is measured

    The Scenario control switches between sound advice (the model as is) and flawed advice (the recommendation shifted two portfolios in the wrong direction). Appropriate reliance means following sound advice and overriding flawed advice. Every logged row records the scenario, the shown portfolio and the sound portfolio, so both rates can be computed per condition.

    Analysis sketch

    Mixed-effects models with participant as random effect (several trials per participant), explanation style and advisor type as fixed effects, financial literacy as covariate, separately for sound and flawed trials. The Analytics page gives the descriptive rates per condition. Short interviews, or the free-text reasons, for why participants followed or overrode.

    Ethics note

    A real study needs ethics approval, informed consent, hypothetical scenarios only, no real financial advice, and a debrief about the flawed trials. The study flow implements a consent screen, fixed hypothetical cases, an attention check and a debrief, to be adapted to the approving committee's requirements.

    Provenance

    AdviceIT extends my systematic literature review of trust and algorithm aversion in the choice between human and AI financial advisors (SSRAAI 2026). The review described the problem of miscalibrated trust. AdviceIT by Radit is a first step toward the design-side answer.

    Limitations, plainly

    The neural network is trained on ILS-Bench, 400 expert-validated but synthetic narratives, not on real client records. No participant data collected yet. Model portfolios are stylised, not calibrated to a regulatory suitability standard. LLM explanations and narrative reading depend on the participant's browser and GPU.

    How it works

    AdviceIT has two advisors, one per page, both learned from the same expert-validated data, and a set of explanation conditions built from two choices: what is explained (content) and how it is delivered (form). The advisor produces the recommendation. The explanation condition decides what the participant sees next to it. The scenario decides whether the advice shown is sound or deliberately flawed. Both advisors take the same profile, describe it with the same three suitability labels, and answer with one of the same six outcomes.

    Shared vocabulary: suitability labels

    Both advisors describe a profile with the three labels used by ILS-Bench, derived from the form by three documented rules. Risk tolerance is the stated Low / Medium / High (Moderate), or Inconsistent when a written description shows conflicting attitudes. Risk capacity counts what could force selling at a loss: no emergency fund, variable income, significant debt or obligations. None of them: High. One: Moderate. Two or more: Low. Liquidity need follows the horizon (1 to 2 years Urgent, 3 to 5 High, 6 to 10 Moderate, 11 or more Low), and a concrete near-term need makes it Urgent whatever the horizon. Age is the fourth input. These rules follow the dataset's codebook, which defines capacity by income, savings, debt and obligations, and liquidity by the need to access the funds soon.

    Advisor 1: the AI advisor (neural network)

    A multilayer perceptron (12 encoded inputs: the three suitability labels one-hot encoded plus age. The participant sets seven form fields, the label rules compress them to the labels, and the encoding expands the labels to 12 numbers. Two hidden layers of 16 units, 6 outputs) trained with ml/train_model.py on ILS-Bench (Bonelli 2026, Mendeley Data, CC BY 4.0): 400 investor narratives whose suitability labels and recommended outcome were validated by a panel of four financial-domain experts. The six outcomes are the five portfolios and Human review. Cross-validated accuracy is percent (5-fold, three repeats), on par with the dataset author's own draft labels and with a lookup table over the label combinations, and its probabilities are calibrated with temperature scaling. It runs in the browser from the exported weights in ml_weights.js. Its weights are not human-readable, so its explanations must be computed post hoc: exact Shapley values of the probability of the recommended outcome, the same counterfactual search, and calibrated probability for confidence.

    Advisor 2: the interpretable rule-based advisor

    A multinomial logistic regression fitted on the same ILS-Bench cases and the same twelve inputs as the network, percent cross-validated accuracy, calibrated with temperature scaling. It is a scorecard derived from data: one weight per input and outcome, the outcome with the largest total wins, and the contribution of each input to the evidence for the recommended outcome is read directly from the weights, exact and additive in log-odds. Same data and inputs as the network, but transparent. That is what makes explanation fidelity a factor: the same explanation styles are exact here and estimated on the AI advisor page.

    Language to suitability

    The ILS-Bench procedure starts from what an investor writes. The profile form therefore accepts a free-text description: the in-browser language model reads it into the form fields (age, horizon, tolerance, emergency fund, income) and flags an Inconsistent risk attitude when the text asks for high returns while saying that a loss would cause serious stress. The researcher can load any of the 400 benchmark cases and compare the advisors' outcome with the expert consensus for that case.

    Explanation conditions: content and delivery

    An explanation condition is a combination of content (what is explained: why, what would change it, how sure, in any combination or none) and delivery (how it is shown: static, interactive what-if, adaptive to literacy, or conversational). The Explanation control offers named presets for the common combinations, and Customise for any other. Both parts are recorded in the log.

    The response

    Under every recommendation the participant rates trust (1 to 7) and chooses Follow, Adjust, Reject or Ask a human adviser. Adjust asks which portfolio they would go for instead, and the log records it together with the number of steps from the shown portfolio. Three optional ratings (understanding, decision confidence, mental demand) and a free-text reason follow. Decision time runs from the moment the recommendation was shown.

    The study flow

    A participant link with flow=study runs the whole procedure: consent, the three literacy questions, six fixed hypothetical cases (half with sound and half with flawed advice, in an order seeded by the participant ID so it is reproducible), one attention check, and a debrief that names the flawed cases. The Analytics page then reads the collected responses.

    Scenarios: sound and flawed advice

    The Scenario control (researcher mode) or the scenario URL parameter (participant links) decides which advice is shown. Sound advice is the advisor's real outcome. Flawed advice takes that outcome and deliberately shifts it two portfolios in the wrong direction: conservative outcomes are pushed toward Aggressive growth, growth outcomes toward Capital preservation, and a Human review outcome is replaced by an automated portfolio two steps more aggressive than the score would justify. Two steps was chosen because one step is often still defensible, while two is a mistake a careful reader can catch.

    The explanations always describe the advisor's real reasoning, so in the flawed scenario the explanation and the recommendation do not fit together. Whether participants notice, and whether the explanation style helps them notice, is exactly what the study measures. Every logged response records the scenario, the portfolio that was shown and the sound portfolio, so appropriate reliance can be computed per condition: the follow rate on sound advice and the override rate (Adjust or Reject) on flawed advice. Following flawed advice is over-reliance, rejecting sound advice is under-reliance. In researcher mode an amber note marks a flawed recommendation. Participants never see that note, and a real study debriefs them about the flawed trials afterwards.

    The full source is readable in model.js (label rules), ml_model.js, logit_model.js, explanations.js, llm.js and ml/train_model.py.