arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

少量のアイスランド手話動画で単独手話認識を評価

Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting

Finnur Ágúst Ingimundarson, Gu{\dh}ný Björk {\TH}orvaldsdóttir, Mathias Müller, Sarah Ebling

この論文をやさしく読む

ひとことで言うと

少ない動画しかないアイスランド手話の認識で、他の手話からの学習転移がどれほど効くか比較した。

何に役立つ?

データの少ない手話向け認識モデルの研究や、事前学習の選択に参考になる。著者は実用にはまだ遠いとしている。

この研究の面白いところ

849クラスの86%に例が2件しかない条件で、米国手話からの転移によりSPOTERの精度が14~24ポイント改善した。

どこまで分かった?

大きな語彙の課題でも精度は22.6%であり、実用水準を示した結果ではない。データは辞書由来の単独手話動画。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アイスランド手話の単独手話認識について、初めての実験を報告する。二言語のオンライン辞書から作られたÍTM SignWikiを使用した。動画は1,845本、クラスは849で、その86%には例が2件しかなく、全体として話者をまたぐほぼワンショット認識となる。オープンソースの認識枠組みOpenHandsとSPOTERを、語彙規模が22、117、849クラスの三つの課題で比較し、三つの姿勢推定器と二種類の言語間転移も評価した。アイスランド手話のデータだけでは、三課題すべてでSPOTERがOpenHandsを上回り、MediaPipeによる姿勢推定はAlphaPoseやSDPoseより良かった。最大の改善をもたらしたのは言語間転移である。米国手話データでSPOTERを事前学習してからアイスランド手話で微調整すると、精度は14~24パーセントポイント上がり、三課題でそれぞれ72.7%、47.9%、22.6%となった。他の六つの手話のデータを使った多言語学習では、全課題におけるOpenHandsの精度が1.41%から28.86%に上がった。実用にはまだ遠いが、データが豊富な手話からの転移が、極めて少量のデータしかない手話に有望であることを示唆する。両枠組みの改変版を公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (ÍTM). We use ÍTM SignWiki, a dataset derived from a bilingual Icelandic--ÍTM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the full task effectively one-shot recognition across signers. We compare two open-source ISLR frameworks, OpenHands and SPOTER, on three tasks of increasing vocabulary size (22, 117 and 849 classes), and evaluate three pose estimators and two forms of cross-lingual transfer. With ÍTM data alone, SPOTER outperforms OpenHands on all three tasks, and MediaPipe poses give better results than AlphaPose or SDPose. Cross-lingual transfer brings the largest gains: pretraining SPOTER on American Sign Language data before finetuning on ÍTM raises accuracy by 14--24 percentage points, to 72.7%, 47.9% and 22.6% on the three tasks, and multilingual training with data from six other sign languages lifts OpenHands from 1.41% to 28.86% on the full task. Although far from practical use, the results suggest that transfer from better-resourced sign languages is promising for very low-resource ones. We release our adapted versions of both frameworks.

arXiv ID: 2609.25862 / 要約の誤りについて