All data in the prototype is synthetic and not affiliated with any financial institution.
A concept prototype on trust, transparency, and graceful recovery in AI-powered financial investigation.
AI Data Investigator lets a risk analyst ask questions about transaction data in plain English. Instead of a paragraph, the AI returns a structured analysis:
An experienced risk analyst that's fluent in data but not necessarily in SQL. A risk analyst's goal is to identify, evaluate, and minimize potential threats that could harm an organization's financial health.
Solo product designer and builder. I defined the problem and AI UX patterns to establish trust and transparency. I directed development with an AI coding agent, and tested the output to find and fix where the AI got things wrong.
| Type | Self-initiated concept |
|---|---|
| Timeline | 3 days |
| Tools | Claude Code, React, TypeScript, Tailwind CSS, Netlify |
| Focus | Explainability · Confidence indicators · Error recovery · RAG · Permission-aware AI · Prototyping in code |
A chatbot connected to a database gives answers analysts can’t verify. In regulated finance, an answer you can’t verify is a liability.
Answers separate data from interpretation, explain their confidence, expose their evidence and sources, fail honestly, and respect access permissions.
A working, hosted prototype over 1,000 synthetic transactions: 5 AI response examples covering multiple scenarios.
Risk analysts spend their day answering why:
The answers are in the data, but getting to them usually means writing SQL or waiting on someone who can. Can AI solve this? The common pattern, a chat window connected to a database, returns a fluent paragraph with no way to verify it.
In a consumer app, a wrong answer is an annoyance. In a bank, an analyst who escalates or clears a case on a wrong AI answer creates regulatory exposure and real harm to customers.
So the design question wasn’t “how do I make AI answer questions?” It was “how do I make its answers verifiable?”
The analyst asks a question, the AI analyzes the data, and the answer is structured so the analyst can understand it, verify it, and dive deeper.
Instead of a mysterious black box that spits out a response to a prompt, the loading spinner and progress bullets show that the AI is reading the data, generating SQL, and analyzing the results. The user can see the AI is working on their question and what it’s doing.
The actual SQL query results the AI is reading from are displayed for the user. The AI can answer questions based on these results, but the distinction between the real query results and AI interpretation is maintained. The two never blend.
Every answer carries a High, Moderate, or Low rating plus a written rationale stating what the confidence doesn’t cover.
The language intentionally avoids definitive statements. The geographic analysis goes further, rating its own result Low and arguing against it.
The generated SQL, relevant transactions, and reasoning are showcased in every AI response.
Policy questions retrieve from a library of fictional policy documents. The answer shows the exact passage it relied on, labelled as a source, and View source opens the full document.
The AI is trained to refuse questions it can’t answer, and to explain why. Two common cases are uncertainty and permissions:
Suggested next steps are labelled "Advisory. Nothing here has been carried out." Labels like "Interpreted as..." and use of italics to distinguish all AI-generated writing allow the analyst to make up their own mind based on not only the recommendations, but more importantly, the real data.
I built a working React prototype rather than static mockups, directing development with an AI coding agent. Code let me test what a mockup can’t: how the design behaves against real (synthetic) data.
Mistake 1
The first dataset reported a 50.7% failure rate for a single day. This would be unrealistic to analysts. The data was rebalanced so the spike day reads as a genuine spike (about 22%), with a cause the analysis explains.
Lesson: In a domain tool, whether the numbers are believable is part of the UX.
Mistake 2
I initially had all AI text displayed in italics. But what about the fact that real data was being selected by the AI?
I needed to explain that the sampled data was real, but still selected by the AI. Now the AI explains what data it chose to display and why.
Lesson: Being transparent about AI-generated writing is crucial. But transparency about all AI decisions is essential.
Live
A live, shareable prototype tested end to end: ask, analyze, inspect, verify, follow up.
Outputs
5 AI response examples and 3 response types (answer, an honest “can’t answer” for impossible questions, and access-restricted responses for when it doesn't have access necessary to answer a question).
Every answer
Carries confidence, rationale, underlying transactions, generated SQL, and cited sources in the form of policy documents.
WHAT’S NEXT
Shadow 3–5 risk analysts, and test whether separating data from interpretation actually changes what they write in case notes.
Confidence and risk currently rely on color. Add non-color cues and run a WCAG 2.2 AA audit.
Turn the patterns (confidence meter, evidence item, source citation, refusal card) into reusable components for other AI features or make use of existing design system components.
The LLM would turn questions into queries and write the explanation around the computed results. It would never generate the numbers itself, so the SQL and evidence views become the verification layer.
REFLECTION
It’s about making it honest about what it knows, clear about what it doesn’t, and easy to verify, so the person accountable for the decision can trust the right answers and catch the wrong ones.