小型LLMの実務完了率を実行基盤で改善するMingbird
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
この論文をやさしく読む
ひとことで言うと
小型LLMが実務を途中でやめたりループしたりする原因を、モデルだけでなく周辺の実行基盤にも求めた研究です。入力予算や完了確認など十の機構を持つMingbirdを比較しています。
何に役立つ?
ローカルで小型モデルを使うエージェントの、課題完了率を改善する設計の参考になります。モデルを替える前に、入力の詰め込み方や終了判定を点検する材料になります。
この研究の面白いところ
機械・モデル・予算・採点を固定して実行基盤だけの差を調べています。一方、単独機構の効果については反復の変動が大きいことを示し、因果的な解釈を抑えています。
どこまで分かった?
著者は自作ベンチマーク、単一マシン、単一試行採点を限界として明記しています。アブレーションの各差は最大0.069の再試行変動と同程度であり、各機構の効果を確定した結果としては扱えません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
20億〜90億パラメータの小型の重み公開モデルは一般的なノートPCで動作するが、クラウド規模を想定したエージェント実行基盤の下では、実際の課題をほとんど完了しない。ツールの事前入力がコンテキストをあふれさせ、自己修正が発散し、ツールの実演がループし、課題が黙って放棄される。本研究は、単一マシンでの統制比較と一つの第三者ベンチマークから、こうした失敗の相当部分がモデルではなく実行基盤に起因する証拠を示す。 WindowsとOllama向けのローカル優先のエージェント実行基盤Mingbirdを導入する。十の機構が小型モデルの失敗形態に個別に対応する。その代表的な三つは、バイト単位で差し引き増加をゼロにする事前入力予算、完了を受け入れる前に課題を読み直す完了ゲート、シグネチャ単位のループ検出である。 マシン、モデル、予算、採点を固定した統制比較LRABでは、4実行基盤×4公開モデル(20億〜350億パラメータ)×18の実務課題を、成果物の決定的な採点で評価した。Mingbirdは総合0.886を達成し、gooseの0.631、opencodeの0.479、agent-miniの0.405を上回った。全288条件を公開している。τ²-benchでは278課題、3条件、1プロトコルで、比較対象の0.791と0.737に対し総合0.856を得た。同じ18課題での最先端モデルの試験では、実行基盤間の結果が0.997から0.478まで広がる一方、適切に構成された実行の枠組み同士の差は0.072以内に収まった。 機構を一つずつ除くアブレーションは、方向性を示すものとしてのみ報告する。同じ条件を同じ夜に反復しても平均が最大0.069動き、これは単一試行で観察された各差と同程度だからである。バッチをそろえた唯一の比較、すなわち全機構と文章の再読だけの比較では、実行可能な完了確認機構は三回の反復を通じ、対応のある比較で+0.10を与えた。証拠には、自作ベンチマーク、単一マシン、単一試行の採点という明示された限界がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $\tau^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
著者のコメント
44 pages, 9 figures. Code, benchmark protocol, scoring code, and all 288 per-cell results: https://github.com/Mingbird/Mingbird-agent
arXiv ID: 2610.02001 / 要約の誤りについて