arXiv論文メモ
新着一覧
cs.HC · 査読状況未確認

文章の意味ではなくバイト列をダンスの動きに変換

The Choreographic Genome: Amplifying the Silent Structure of Text into Dance

Michael Li, Alison Ding

この論文をやさしく読む

ひとことで言うと

文章の意味を動作にするのではなく、文字を保存するバイト列を一定の規則でダンスへ変換する作品・可視化の方法です。

何に役立つ?

文章の符号化の違いを身体の動きとして見せる芸術表現や、文字体系と計算機表現の関係を考える用途が提示されています。

この研究の面白いところ

同じ意味かどうかではなく、符号化された構造によって動きが変わります。複数バイトで表される文字体系の違いが、1文字あたりの動きの量にも現れます。

どこまで分かった?

著者は動作生成性能のベンチマークではなく芸術的な可視化と明示しています。3倍近い動きという結果は意味の豊かさや言語の優劣を示す指標ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成AIの近年の進歩により、複雑な人間の動きを前例のない忠実さで合成できるようになった。しかし現在のテキストから動作への変換系は、言語の意味に厳密に依存する。入力が「私は両手を上げる」であれば、モデルは手を上げた姿勢を探し、文章の意味以外の性質はすべてノイズとして捨てられる。本研究では、その捨てられた構造を信号として扱う。文章が何を意味するかではなく、どのように組み立てられているかを増幅する、身体による可視化の道具を提示する。 まず主成分分析とK平均法によって、ダンスの運動学的特徴を256個の様式的な「領域」からなる動作コードブックへ量子化し、動きの主軸に沿って領域を並べる。次に、任意の入力文章の生のバイト表現をこのコードブックへ直接対応付け、文章の「振付ゲノム」と呼ぶ決定論的な領域列を作る。事前計算した動作の妥当性を表すグラフと物理的な平滑化処理の組合せが、この列を滑らかな全身運動へ変換する。こうして踊る身体は、意味を扱う系が無視するバイト単位の構造を表示する面になる。 シェイクスピアのソネット、機械のエラーログ、ソースコード、奴隷制廃止論者の問い、先住民の文字体系、デーヴァナーガリー文字を含む一連の芸術的事例を通じ、各文章が見た目に異なるダンスを生むことを示す。また、ASCII中心の計算環境で周縁化される文字体系では、1文字あたりの動きが3倍近くへ増幅されることを示す。本研究はこれを動作合成のベンチマークではなく、何を信号として数え、何を聞き取られないままにするかを問い直す、批評的で詩的な可視化として位置付ける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in generative artificial intelligence have enabled the synthesis of complex human motion with unprecedented fidelity. However, current text-to-motion systems rely strictly on linguistic semantics: if an input reads "I put my hands up", the model searches for a pose with raised hands, and every non-semantic property of the text is discarded as noise. In this work, we treat that discarded structure as the signal. We present an embodied visualization instrument that amplifies not what a text means, but how it is built. Our method first quantizes dance kinematics into a motion codebook of 256 stylistic "regions" using Principal Component Analysis and K-Means clustering, and orders those regions along the dominant axis of movement. We then map the raw byte representation of any input text directly onto this codebook, producing a deterministic sequence of regions that we call the text's "choreographic genome". A precomputed plausibility graph and a set of physics smoothing routines turn this genome into fluid, full-body movement, so that the dancing body becomes a display surface for the byte-level structure that semantic systems ignore. Through a series of artistic case studies, including a Shakespeare sonnet, a machine error log, source code, an abolitionist's question, and Indigenous and Devanagari scripts, we show that each text produces a visibly distinct dance, and that scripts marginalized by ASCII-centric computing are amplified into close to three times as much movement per character. We frame this not as a motion-synthesis benchmark, but as a critical and poetic visualization that asks what we choose to count as signal, and what we allow to go unheard.

著者のコメント

Accepted at IEEE VIS 2026 Arts Program

arXiv ID: 2609.22519 / 要約の誤りについて