Robust object detection in optical remote sensing imagery via partial-channel self-attention and intensity-enhanced feature fusion
Optical remote sensing object detection is challenged by sparse object distribution, large spatial coverage, and pronounced scale variation, which complicate effective feature representation and contextual reasoning. Transformer-based detectors enable global dependency modeling but often introduce redundant feature interactions and insufficient feature regulation during multi-scale fusion, limiting discriminative representation, particularly for small and low-contrast targets. In this study, we propose a CNN-Transformer hybrid detection framework that improves global contextual modeling and stabilizes multi-scale feature interaction. Specifically, a partial-channel self-attention–based internal feature interaction module is introduced to selectively model long-range spatial dependencies within informative feature subspaces, improving contextual representation quality while reducing redundant interactions. In addition, an intensity-regulated fusion mechanism employing bounded nonlinear modulation is developed to stabilize cross-scale feature interaction and enhance discriminative object responses. Extensive experiments on three optical remote sensing benchmarks (HRSC2016, DOTA v1.5, and DIOR) demonstrate consistent detection performance improvements over a real-time transformer baseline, achieving mean average precision scores of 91.7%, 72.5%, and 86.6%, respectively, without increasing model complexity. These results demonstrate that selective feature interaction and bounded intensity regulation improve feature representation robustness and detection performance in complex remote sensing imagery.