arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

検出結果が届く時点を予測するイベントカメラ方式

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim

この論文をやさしく読む

ひとことで言うと

イベントカメラの検出結果が計算後に届く時刻を見越し、物体の未来の状態を予測する方式である。

何に役立つ?

考えられる用途は、高速で移動するロボットや車両が古い検出結果で判断する問題を減らすことだ。

この研究の面白いところ

従来の観測時刻での評価を見直し、利用可能時刻で測った性能低下を、軽量な時間特徴の融合で回復した。

どこまで分かった?

性能は記載された運転・ドローン・飛行データでの評価であり、実機の意思決定や安全性の改善は要旨で検証されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

イベントカメラは高速なロボットシステムに低遅延の知覚をもたらすと期待される。わずかな遅れでも、検出結果が後段のロボットの判断に使われる頃には古くなり得る。しかし現代のイベント検出器でも、予測を利用できるまでに数十ミリ秒の計算が必要である。従来の評価は、予測を観測時刻の正解ラベルと比べるため、この遅れを無視する。予測が出る頃には場面が変化しているかもしれない。本研究は、イベントに基づく複数物体の検出で、観測時刻と利用可能時刻のずれを調べる。最先端のイベント検出器でも、観測時刻ではなく結果が利用できる時刻で評価すると性能が大きく落ちることを示す。 対策として、出力が利用可能になる時点での物体の状態を予測する、因果的な検出器ChronoFuseを導入する。ChronoFuseは、複数の尺度を持つ特徴階層で、現在の表現と保存した過去の特徴を時刻をまたいで因果的に融合し、未来の観測を使わずに短期的な時間の手掛かりを得る。この融合経路は軽量で、追加パラメータは17万、平均の端から端までの遅延増加は0.84ミリ秒にとどまる。1Mpxの運転データでは遅延によって失われた精度の71%を、急速に動くドローンのFREDでは90.8%を取り戻し、遅延がない場合の性能に近づいた。動きが極端なEV-FlyingではsAP 20.95に達し、最も強い通常のイベント検出器の2.25に対して9.3倍となった。高速で変化する場面で動作するロボットには、未来の状態を先読みすることが重要になり得る。自動運転、機敏な飛行、ロボットによる捕捉などが含まれる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Event cameras promise low-latency perception for high-speed robotic systems, where even short delays can render detections stale by the time they inform downstream robotic decisions. Yet modern event detectors still require tens of milliseconds of computation before their predictions become available. Conventional evaluation ignores this delay by comparing predictions with annotations at the observation timestamp, even though the scene may have changed by the time those predictions are produced. We study this observation-availability mismatch in event-based multi-object detection and show that state-of-the-art event detectors degrade substantially when evaluated at prediction availability rather than observation time. To address this, we introduce ChronoFuse, a causal availability-time detector that predicts object states for when its output becomes available rather than for when its input was observed. ChronoFuse performs causal cross-time fusion over a multi-scale feature hierarchy, combining current representations with cached temporal features to expose short-term temporal cues without using future observations. The fusion pathway is lightweight, adding only 0.17 million parameters and 0.84 ms of mean end-to-end latency overhead. ChronoFuse recovers 71% of the accuracy lost to latency on 1Mpx driving data and 90.8% under rapid drone motion on FRED, nearly restoring zero-delay performance. Under the extreme motion of EV-Flying, ChronoFuse reaches 20.95 sAP, compared with 2.25 for the strongest standard event detector (9.3x gain). These results show that predicting ahead can be critical for robots operating in fast-changing scenes, including autonomous driving, agile flight, and robotic interception.

arXiv ID: 2609.26919 / 要約の誤りについて