arXiv論文メモ
新着一覧
cs.SE · 査読状況未確認

組織の設計慣行からAI生成コードの不適合を検査する

A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form

Chien-Tsun Chen, Yu Chin Cheng

この論文をやさしく読む

ひとことで言うと

AIが一般的な書き方へ流れるのを防ぐため、組織固有の設計慣行を明示し、違反を検出・修正する開発方法です。

何に役立つ?

AI生成コードの検査項目と、人が判断すべき仕様の妥当性を整理する助けになります。Scrumシステムの構築・再構築を事例として報告しています。

この研究の面白いところ

コードを直す自動ループと、検査基準そのものが現実に合うかを決める人のループを分ける点が中心です。

どこまで分かった?

報告された証拠は64仕様から作った特定のシステムに関するものです。テスト件数やゲート数だけで仕様と現実の適合を保証するとはしておらず、そこには人の判断が残ります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)が生成したコードが、指定された要件を満たすとは仮定できない。レビュー、テスト、静的解析は依然として必要だが、十分な検査・制御の仕組みにはどれがどの役割で必要なのかは未解決である。そこで、クリストファー・アレグザンダーの形の理論に基づく設計理論と、その適用方法を提案する。アレグザンダーの説明では、形とその文脈の適合は、特定された不適合が存在しないことを通じて、否定的にのみ認識できる。本研究では組織の慣行を明示し、それと問題の分類から不適合を導く。LLMを、多数のコードベースで学習していても、どのコードベースにも根差していない、地域固有の様式になじみのない作り手としてモデル化する。その出力は、局所的な慣行よりも一般的な慣習へ流れやすい。 4つの仕組みを設計する。第1は、問題の明示的な表現であるジャクソンの問題フレームと、慣行の明示的な表現である4形式のパターン言語、第2は決定論的な不適合検出器、第3は修正ループ、第4は表現と検出器を管理し、人の承認を必要とする規則制定の回路である。得られた方法を、不適合に基づく開発(Misfit-Governed Development、MGD)と呼び、ハーネス・エンジニアリングの実践として位置付ける。二重ループの過程は、LLMが検査ゲートに照らして反復する自律的な内側のループと、仕様を現実に照らして判断する人による外側のループを分ける。これらが一体となって、S = P = T = Wという保証モデル、すなわち仕様・プログラム・テスト・現実のモデルを構成する。等号は同一性ではなく関係を表す。 64の問題フレーム仕様から、イベントソーシングによる4つの集約を持つScrumシステムを構築し、再構築した証拠を報告する。検証には約1300の生成テストと28の処理停止ゲートを用い、そのうち1つは188の規則を適用する。これはアレグザンダーが1996年のOOPSLAで提起した課題のうち、生成性の側面に対応する。一方、仕様がなお現実に適合しているかという道徳的な側面は、人の判断を必要とし、外側のループに属する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Code generated by large language models (LLMs) cannot be assumed to meet specified requirements. Reviews, testing, and static analysis still apply, but which of them a sufficient harness needs, and in what role, is open. We propose a design theory derived from Christopher Alexander's theory of form, and a methodology for applying it. In Alexander's account, fit between a form and its context can be perceived only negatively, through the absence of identified misfits. We make the organization's tradition explicit and derive the misfits from it and from the problem's classification. The theory models the LLM as a non-native vernacular builder, trained on many codebases but native to none, whose output tends to drift toward mainstream conventions rather than the local tradition. We engineer four pieces of machinery: explicit representations of the problem (Jackson's problem frames) and of the tradition (a four-form pattern language); deterministic misfit detectors; a fix loop; and a human-gated legislative circuit governing the representations and detectors. We call the resulting methodology, a practice of harness engineering, Misfit-Governed Development (MGD). Its dual-loop process separates an autonomous inner loop, where the LLM iterates against the gates, from a human outer loop, where specifications are judged against the world. Together they form the S = P = T = W assurance model (specification, program, tests, world), whose equals signs name relations, not identity. We report evidence from building and rebuilding a Scrum system of four event-sourced aggregates from 64 problem-frame specifications, verified by about 1,300 generated tests and 28 blocking gates, one applying 188 rules. This addresses the generativity dimension of Alexander's 1996 OOPSLA challenge. The moral dimension, whether the specification still fits the world, requires human judgment and belongs to the outer loop.

arXiv ID: 2610.01372 / 要約の誤りについて