ドメイン名の自己申告がAIの危険判定を揺らす
ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments
この論文をやさしく読む
ひとことで言うと
ドメイン名に含まれる『公式』などの言葉だけで、AIの危険判定が変わるかを調べた研究です。
何に役立つ?
ドメイン名の安全性をAIで判定する仕組みで、名前自体の自己申告に引きずられない評価を設計する際に役立ちます。
この研究の面白いところ
命令文を使わず、名前の一部を変えるだけで警告率が大きく上下し、公式ドメインを参照させても残る影響がありました。
どこまで分かった?
数値は構成したブランド類似の名前、64ブランド、五つのモデルでの判定結果です。現実の被害件数を示すものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
『フィッシングではない』や『公式』といった短い主張は、明示的なプロンプトインジェクション命令がなくても、大規模言語モデルによるドメイン名の判定を変え得る。本研究は、検査対象の名前自身が安全性や承認を主張する操作をClaimMirageとして調べる。64のブランドと五つの言語モデルについて、62万2080件の判定を分析し、作成したブランドに似た名前に含めた十種類の主張を、長さとハイフンの数を合わせた対照と比較した。自己主張はモデルと入力条件に応じて、警告を大きく減らす場合も増やす場合もあった。ある条件では、比較のために模倣の可能性があるブランド名とその公式ドメインをプロンプトに与える基本的な安全策があっても、登録可能な名前の部分に置いた危険を否定する語句は警告率を45.3ポイント下げた。同じモデルでこれらの参照情報がない場合、その位置に置いた承認を示す語句は警告率を65.6ポイント上げた。参照情報やドメイン構成要素の注釈により、警告率の低下が一部解消されたが、残るものやかえって大きくなるものもあった。結果は、検査対象の名前が自身について述べる主張への耐性を試し、ドメインを安全または承認済みと判断する前に独立した証拠を求める必要があることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.
著者のコメント
6 pages, 3 figures
arXiv ID: 2609.29130 / 要約の誤りについて