arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

大規模言語モデルは古典アラビア語のマカーマを書けるか

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

AbdulRahman A. Morsy (1), Aya Zirikly (1 and 2) ((1) Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, (2) Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)

この論文をやさしく読む

ひとことで言うと

5種類のLLMに古典アラビア語のマカーマを書かせ、指示方法によって文体や構成の質がどう変わるかを評価した研究。

何に役立つ?

文化や文学形式に固有の制約を持つ文章を、LLMで評価・生成する方法を考える際の参考になる。

この研究の面白いところ

少数例は押韻散文の密度を安定して改善したが、5モデル全体の総合得点は例なしの指示が最高で、評価軸ごとに適した指示が異なった。

どこまで分かった?

評価は5モデルと指定された指示条件・評価軸での比較である。LLM判定だけでなく人の注釈も使っているが、文学的な価値を全面的に判定したものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は創作的な文章生成で高い性能を示すが、文化的な背景を持ち、文体上の制約が強い文学形式を作れるかは十分に調べられていない。従来の研究は現代語や詩に偏り、マカーマのような古典的散文の伝統はほぼ未研究である。マカーマは、押韻散文(サジュ)、密度の高い修辞的装飾、連作的な物語構造を特徴とする古典文学のジャンルであり、LLMが表面的な流暢さを超えて文学的能力を持つかを試す難しい題材となる。 本研究は、LLMによるマカーマ生成を初めて統制して評価したとする。5モデルを、例を示さない指示、少数の例を示す指示、規則を与える指示の条件で比較した。出力は人による注釈とLLM判定の双方を使い、修辞の豊かさ、サジュの密度、構造の一貫性、文体の真正さなどで評価した。結果として、指示の方法は文体の質に大きく影響した。少数例を示す方法はサジュの密度を最も一貫して改善したが、修辞と一貫性への影響はモデルにより異なった。これらの側面では、最も強いモデルであるGPT-4oとGPT-5.4-miniが規則による指示から最も恩恵を受けた一方、5モデル全体の総合得点は例を示さない指示が最も高かった。さらに、アラビア語マカーマの慣習との文体的な一致度にはモデル間で系統的な差があった。結果は、独立した二つ目のLLM判定器、対応のある統計的有意性検定、サジュに関するLLMを使わない代理指標でも裏付けた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.

著者のコメント

14 pages

arXiv ID: 2609.28245 / 要約の誤りについて