arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

インドの銀行業務向けAIの安全性と信頼性を測る評価集

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Suvradip Paul, Chandra Bhushan, Harsh Sharma, Nitin Kukreja, Yatharth Dedhia, Keyur Doshi, Prashant Devadiga

この論文をやさしく読む

ひとことで言うと

銀行業務のAIが毎回安全に、正しい口座と情報を使って対応できるかを測る評価集です。

何に役立つ?

銀行向け支援AIの選定や改善で、返答だけでは見えない操作上の誤りを調べる用途が考えられます。

この研究の面白いところ

同じ事例を三回試して全部成功した割合を測ると43.7~58.2%で、一回でも成功した割合60~74%との差が大きくなりました。

どこまで分かった?

数値はこの評価集における11モデルの結果です。実際の銀行サービスでの事故率やすべての国の銀行業務への適用結果ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

銀行業務の支援AIは、口座ごとの情報を使って依頼に答え、多くの場合はツールを通じて操作も行う必要がある。最終的な返答だけを評価すると重要な誤りを見逃す。すでに持っている情報を再び尋ねたり、古い文脈に頼ったり、口座を取り違えたり、正しい値を述べながら書き込みには誤った値を使ったりする場合がある。そこで、インドの個人向け銀行業務を対象とする799事例の評価集IndicBankBenchを提案する。五つの業務領域、能力と拒否に関する領域、20の主要な評価軸を含む。各事例は、安全性、操作とツール利用、応答の適切さ、助言の質という四段階で評価する。ツール利用と大部分の安全性検査は決定的に判定する。書き込み前の確認が曖昧な事例だけを限定的な判定器が扱い、別の大規模言語モデルによる判定器が応答の意味的な適切さを評価する。各事例を三回実行し、すべて成功した場合だけ合格とする厳格なpass^3を報告する。評価した11モデルで、この厳格な信頼性は43.7~58.2%だった一方、少なくとも一回成功した割合は60~74%だった。この差は、一度でも成功した割合が銀行業務での安定した振る舞いを過大評価し得ることを示す。事例ごとの診断により、不要な質問をするシステムと、操作はするが顧客の文脈を照合できない、または依頼を完全には解決できないシステムも区別できる。事例、模擬環境、評価用の仕組みを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advisory quality. Tool use and most safety checks are deterministic. A narrow resolver handles only ambiguous confirmation-before-write cases, while a separate LLM judge evaluates semantic response adequacy. We run every case three times and report strict pass^3, which requires success on all trials. Across the eleven evaluated models, strict reliability ranges from 43.7% to 58.2%, whereas at-least-once success ranges from 60% to 74%. This gap shows that at-least-once success can overstate dependable banking behavior. The case-level diagnostics also distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. We release the cases, mock environment, and evaluation harness.

著者のコメント

16 pages, 4 figures

arXiv ID: 2609.29167 / 要約の誤りについて