Abstract:To address the cross-modal registration challenges between Synthetic Aperture Radar(SAR) and optical images, which arise from differences in imaging mechanisms, noise characteristics, and textures—particularly the difficulties in feature extraction from low-texture SAR images, significant domain disparities, and high computational costs of existing methods—this study proposes an improved algorithm named EnhancedXFeat based on the Accelerated Features(XFeat) network. The algorithm enhances the ability to preserve shallow-level details and distinguish deep-level semantics by doubling the number of channels in key modules of the backbone network; introduces a cross-modal dual-branch structure to independently process modality-specific features; and applies a dual attention mechanism twice. These designs effectively resolve the domain difference issue and improve the performance of feature extraction and fusion for low-texture images. Comparative experimental results on a self-constructed dataset demonstrate that EnhancedXFeat outperforms mainstream algorithms significantly: the average number of detected feature points reaches 37.0, which is much higher than 4.3 of Scale-Invariant Feature Transform(SIFT), 6.2 of Oriented FAST and Rotated BRIEF(ORB), and 17.9 of Optical-SAR Match Net(OSMNet); the matching accuracy achieves 88.1%, remarkably superior to 9.2% of SIFT,12.7% of ORB, 68.0% of OSMNet and 68.7% of VGG-16; meanwhile, the single-operation time is controlled to 0.801 2 s, achieving a favorable balance between accuracy and efficiency. The conclusion indicates that EnhancedXFeat effectively improves the accuracy and robustness of SAR-optical image registration through its targeted design, and its lightweight architecture provides an efficient and reliable solution for multi-source remote sensing image registration applications in resource-constrained environments.