English

連絡先

研究室見学、大学院進学、共同研究のご相談など、お気軽にご連絡ください。

Email

kyoshioka47@keio.jp

Kentaro (Ken) Yoshioka

所在地

〒223-8522

神奈川県横浜市港北区日吉3丁目14−1 矢上キャンパス23棟

学生居室: 23-214, 14-305, 24-318 / PI: 23-216A

← プロジェクト一覧に戻る
arXiv 2026プレプリントCircuit

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

W. Zhang, J. Yin, K. Yoshioka

DetAS の概要図
論文 Figure 1 より

概要

物体検出を固定パイプラインではなく動的な意思決定として捉えるエージェント型フレームワーク。マルチモーダル大規模言語モデル(MLLM)が中心エージェントとなり、画像復元モジュールと専門検出器のツールボックスから、シーンに応じた処理手順をその場で組み立てる。少量の注釈データから意思決定の経験を蓄積して推論時に活用する DetAS-X により、6つのベンチマークで既存のMLLMベース検出器を平均F1で28.36%、暗所データセットDarkFaceでは最大37.01%上回った。

An agentic detection framework that treats object detection as a dynamic decision process rather than a fixed pipeline. A multimodal large language model acts as the central agent, composing a detection workflow per scene by choosing from a toolbox of restoration modules and specialized detectors.

Two components drive it: Self-Adaptive Image Restoration, which decides whether and how to enhance an image before detection, and Multi-Expertise Detection, which reconciles the predictions of several domain-specialized detectors through instance-level reasoning. DetAS-X extends this with Self-Evolving Experience Harvesting, accumulating node-level decision experience from a small annotated set so the system reasons from past decisions at inference time.

Across six challenging benchmarks DetAS-X outperforms existing MLLM-based detectors by 28.36% F1 on average, reaching a 37.01% gain on DarkFace.