Snapshot
- Role
- Product design
- Focus
- Four transparency features for a retrieval-augmented research tool, validated with 50 financial professionals
- Published
- CHI EA '25, co-authored with Renuka Sindhgatta
The problem
Financial analysts are expected to use AI tools they cannot audit. Regulatory compliance requires tracing every AI-generated insight back to its source, and retrieval-augmented systems don't expose enough of their reasoning to make that possible.
The trust problem runs both ways. Over-reliance is dangerous in high-stakes decisions; blanket scepticism means the tool goes unused. The design target isn't more trust, it's calibrated trust.
Before designing anything, I ran informal interviews about what people actually want from a research tool when they're deciding where to put money. Those conversations shaped what the four features were trying to do.
Four features, four bets
Confidence levels
Confidence levels. LOW, MEDIUM, HIGH. My bet, and my idea: analysts should know how sure the system is before weighting its answer.
Source highlighting
Source highlighting. Show the exact passage behind an answer, so reasoning is traceable rather than asserted.
Branching
Branching. Let analysts pursue parallel lines of inquiry without losing context.
Feedback
Feedback. Categorical and open-ended, so the system can be corrected rather than only accepted or ignored.
Four features, four bets. One of them was mine, and the study disproved it.
The study
Method and measures
50 financial professionals. Pre- and post-task trust surveys, Wilcoxon signed-rank for the quantitative comparisons, thematic analysis for the qualitative.
What it proved me wrong about
Confidence levels did nothing. On their own they failed to significantly increase trust (p = 0.080). The most intuitive feature, the one that looks like transparency, was the one that didn't work. It was also mine.
The reason, I think: a confidence score is a claim about reliability, not evidence of it. It asks an analyst to trust the system's self-assessment, which is precisely the thing they cannot verify. It restates the problem in smaller type.
Highlighting worked on both measures. Combined with confidence levels it significantly increased trust (p = 0.0036); alone it significantly improved understanding (p = 0.0021). Analysts could trace a claim back to a specific passage and judge it themselves.
Trust and understanding, measured across the four features.
Branching mattered more than either. 27 of 50 named it their most valuable control feature, ahead of sources (15), feedback (26), and personalisation (8). I had designed it as a convenience. Participants used it to verify reasoning before acting.
What follows for design
- A system's self-reported confidence is not transparency. Showing what it used is.
- Traceability beats summary: analysts want the passage, not a description of the passage.
- Control features get repurposed. Branching was built for convenience and used for verification, worth watching for, because it tells you what users actually needed.
What I'd design now
If I rebuilt this tool with what the study produced, I'd cut confidence levels entirely and spend the space on retrieval. Highlighting won because it let analysts check the work themselves, so the interface should show more of what the system read and less of what it thinks about what it read. Branching moves from a convenience feature to the primary surface, since that's what participants actually used to verify reasoning before acting.
Which would you trust more: a system that tells you how sure it is, or one that shows you what it used?
Outcome
Published at CHI EA '25. The findings apply wherever AI-assisted decisions carry accountability, legal, healthcare, consulting, anywhere a recommendation has to be defensible.
Divya Ravi and Renuka Sindhgatta. 2025. Exploring Trust and Transparency in Retrieval-Augmented Generation for Domain Experts. CHI EA '25. DOI: 10.1145/3706599.3719985