
An agentic detection framework that treats object detection as a dynamic decision process rather than a fixed pipeline. A multimodal large language model acts as the central agent, composing a detection workflow per scene by choosing from a toolbox of restoration modules and specialized detectors.
Two components drive it: Self-Adaptive Image Restoration, which decides whether and how to enhance an image before detection, and Multi-Expertise Detection, which reconciles the predictions of several domain-specialized detectors through instance-level reasoning. DetAS-X extends this with Self-Evolving Experience Harvesting, accumulating node-level decision experience from a small annotated set so the system reasons from past decisions at inference time.
Across six challenging benchmarks DetAS-X outperforms existing MLLM-based detectors by 28.36% F1 on average, reaching a 37.01% gain on DarkFace.
