画像の内容と分けて感情表現を3軸で調整する
Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance
この論文をやさしく読む
ひとことで言うと
描く人物や風景の説明はそのままに、画像が与える感情的な印象を快・不快、覚醒度、支配性の3軸で調整する方法です。
何に役立つ?
考えられる用途は、同じ内容の画像について感情表現だけを段階的に変え、制作時に比較することです。研究では客観指標と人の評価を用いて、制御の正確さや変化の知覚可能性を調べています。
この研究の面白いところ
感情を表す単語を追加する方式と異なり、内容説明とは別の連続座標を設けます。データセットでも内容説明と感情評価を分けて収集し、学習でも両者の両立を扱います。
どこまで分かった?
要旨にはデータセット規模、評価者数、改善幅の具体値は示されていません。人が知覚し順序付けできる変化を報告していますが、すべての人や文化で同じ感情として受け取られるとまでは述べていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文章から画像を生成するモデルは対象や場面を正確に描けるようになったが、制作者が内容の説明を書き直さずに、画像で伝えたい細かな感情を指定することは依然として難しい。自然言語で感情を示唆することはできても、意味が安定し、強度の順序が定まった制御尺度は得られない。 本研究ではEMOTRANSを提案する。心理学に基づく快・不快、覚醒度、支配性(VAD)の座標を、内容を記述する文章から独立した生成条件へ変換し、ノイズ除去の各段階で調節することで、感情的な作風を細かく調整できる創作上の変数にする。この目的のために、客観的な内容説明と、それとは別に複数の評価者から収集した感情評価を組み合わせた絵画データセットEMOVADも構築する。さらに、共有モデルを用いる二分岐の学習によって、感情表現と内容の保持を両立させる。 客観評価と人による評価から、提案する枠組みは3次元の感情制御の正確さを改善し、文章との整合性と画像品質で競争力を維持しながら、知覚でき、順序付けできる連続的な変化を生み出すことが示された。本研究は、客観的な内容の描写から細かな感情調整へと画像生成を拡張する、実用的な感情駆動型の手法を提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We propose EMOTRANS, which transforms psychologically grounded valence-arousal-dominance (VAD) coordinates into generation conditions that are independent of the content text and modulated across denoising stages, making emotional style a finely adjustable creative variable. To support this goal, we construct EMOVAD, an art-painting dataset that pairs objective content descriptions with separately collected emotional ratings from multiple annotators. We also coordinate emotional expression and content preservation through dual-branch training with a shared model. Objective and human evaluations show that the framework improves the accuracy of three-dimensional emotion control and produces perceptible, orderable continuous changes while maintaining competitive text alignment and image quality. This work provides a practical emotion-driven approach to image generation that extends objective content depiction to fine-grained emotional adjustment.
著者のコメント
9 figures, 4 tables
arXiv ID: 2609.24215 / 要約の誤りについて