中断後の旅行日程を再計画・修復・編集する手法の比較
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
この論文をやさしく読む
ひとことで言うと
旅行予定が欠航や施設閉鎖で実行できなくなったとき、全体を作り直す方法と部分修正の方法を共通条件で比べています。
何に役立つ?
実行可能性の回復、既に受け入れた予定の維持、計算費用の釣り合いを考える際の評価です。
この研究の面白いところ
単一障害500件と複合障害200件を比較し、全再計画が複合障害で高い成功を示す一方、階層修復や局所修正は成功時に変更を少なく抑えます。
どこまで分かった?
評価したアダプターとベンチマークの中での比較です。IPyHOPPERはLLM推論を使わず、iTIMOは多くのトークンを使うなど実装条件も異なり、単純な成功率だけでは選べません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
旅行計画エージェントが作成した日程は、承認後に欠航、ホテルの空室不足、観光地の閉鎖によって実行不能になることがある。このような日程の改訂には、全面的な再計画、古典的な計画修復、LLMベースの旅行エージェントによる改訂が関わるが、課題の定式化と評価手順が異なるため、比較が難しい。本研究では、TREK由来の2つのベンチマークセットを用いた体系的な実証研究を行う。一つは実行可能・不可能な事例を含む単一中断500件、もう一つは実行可能な同時複合中断200件である。LLM-Z3による全面再計画、IPyHOPPERによる階層的修復、iTIMOによる局所改訂アダプターを、有効性、計画の安定性、計算コストについて比較した。 Geminiを用いたLLM-Z3は、複合中断で観測された成功率が最も高かった。IPyHOPPERは単一中断の全体成功率でその構成にほぼ匹敵し、成功した修復では、承認済みの日程を大幅に多く維持した。階層的修復と局所修復が成功した場合、全面再計画より編集回数が少なく、承認済みの約束を多く保持した。計算プロファイルも異なり、IPyHOPPERはLLM推論を使わず、評価したLLM-Z3アダプターは簡潔な1回の推論を使い、iTIMOアダプターは大幅に多くのトークンを消費した。本研究は、評価した設定の範囲で、実行可能性の回復、既存の約束の維持、計算コストのバランスを取るための実践的な指針を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration's single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.
著者のコメント
16 pages, 2 figures, 6 tables
arXiv ID: 2609.19654 / 要約の誤りについて