© 2026 The authors. This article is published by IIETA and is licensed under the CC BY 4.0 license (http://creativecommons.org/licenses/by/4.0/).
OPEN ACCESS
Highway slope rockfall is a key factor inducing severe geological hazards and traffic accidents. However, rockfall events in nature exhibit a long-tail distribution, and rockfall samples are scarce in real monitoring images, which limits the detection accuracy and generalization ability of deep learning models in complex field scenes. To address these problems, this paper proposes an improved RF-DETR (Real-Time Detection Transformer) rockfall detection algorithm based on generative data transfer learning. First, the Stable Diffusion XL (SDXL) model combined with prompt engineering is introduced to construct a synthetic rockfall dataset covering complex environments such as rainy days, nighttime, occlusion, and multi-scale conditions, effectively alleviating the long-tail data scarcity problem. Second, in the RF-DETR network architecture, the CBAM (Convolutional Block Attention Module) dual attention mechanism is introduced to strengthen small-object rock fragment features, and the BiFPN (Bidirectional Feature Pyramid Network) is adopted to replace the original feature fusion module, improving the multi-scale representation capability; meanwhile, a denoising query (DN-Query) mechanism is introduced into the Transformer decoder to accelerate convergence in dense rockfall scenes. Finally, a two-stage transfer learning strategy is designed, which effectively aligns the feature distributions of the generative and real domains through synthetic-data pre-training and real-data fine-tuning. Experimental results show that on the self-constructed highway slope rockfall dataset, the improved RF-DETR algorithm achieves an mAP@0.5 of 84.2% and an inference speed of 38.7 FPS, comprehensively outperforming the latest mainstream detection algorithms such as YOLOv12-X, RT-DETR++, and Salience-DETR. This method provides a new theoretical perspective and engineering practice scheme for intelligent monitoring of geotechnical hazards.
slope rockfall, object detection, generative data, transfer learning, RF-DETR, geotechnical engineering monitoring
The geological structures along highways in southwestern and mountainous China are complex, and factors such as rock weathering, rainfall, and earthquakes can readily induce slope rockfall disasters [1]. Rockfall is characterized by sudden occurrence and unpredictable motion trajectories, seriously threatening the safe operation of traffic lifelines. Traditional slope monitoring methods (such as manual inspection, total stations, and fiber Bragg grating sensors) suffer from limitations of high cost, narrow coverage, and poor real-time performance [2] (see Figure 1). In recent years, with the rapid development of computer vision technology, image object detection algorithms based on deep learning have gradually become a frontier research hotspot in intelligent monitoring of geological hazards [3, 4].
Figure 1. Schematic diagram of the granite slope rockfall disaster scene on a highway
Currently, convolutional neural networks (CNNs), represented by the YOLO (You Only Look Once) series, have been widely applied in the field of rockfall detection [5]. For example, Zhong et al. [1] proposed the RAL-YOLO framework based on YOLOv8 for rockfall motion classification and localization. However, in real highway slope scenes, rockfall detection faces three severe challenges: first, “long-tail data scarcity”—rockfall events are low-probability sudden incidents, making it difficult to collect complete datasets covering various lighting, weather, and occlusion conditions; second, “severe background interference”—the natural textures of rock masses such as granite are highly similar to rockfall features, easily causing false detections; third, “multi-scale problems”—the scale span from huge rock blocks to fine rock fragments is extremely large, and the limitation of CNN local receptive fields leads to insufficient capability in capturing global contextual information.
To overcome the local limitation of CNNs, Transformer-based object detectors (the DETR series) have demonstrated great potential owing to their global attention mechanisms [6, 7]. In 2025, the Roboflow team proposed the RF-DETR algorithm [8], which combines the DINOv2 pre-trained backbone network with neural architecture search (NAS), achieving Pareto optimality between accuracy and speed in real-time detection tasks. However, directly applying RF-DETR to slope rockfall detection still cannot effectively solve the overfitting problem caused by data scarcity.
In view of this, this paper proposes an improved RF-DETR rockfall detection algorithm based on generative data transfer learning. The main innovations of this paper are summarized as follows:
(1) A paradigm for generating synthetic slope rockfall data based on the diffusion model (Stable Diffusion XL (SDXL)) is proposed. Through prompt engineering, a large-scale synthetic dataset covering extreme weather and multi-scale scenes is constructed, fundamentally alleviating the long-tail data scarcity problem.
(2) The RF-DETR architecture is deeply optimized: a CBAM (Convolutional Block Attention Module) dual attention mechanism is embedded at the output of the DINOv2 backbone to suppress rock-mass background interference; BiFPN (Bidirectional Feature Pyramid Network) is adopted to achieve efficient bidirectional cross-scale feature fusion; and a denoising query (DN-Query) mechanism is introduced into the decoder to improve matching stability.
(3) A two-stage domain-adaptation transfer learning strategy is designed, which effectively bridges the distribution difference (domain gap) between the generative and real domains through synthetic-domain pre-training and real-domain fine-tuning (see Figure 2).
Figure 2. Overall research framework based on generative data and the improved RF-DETR
2.1 Intelligent monitoring of slope rockfall and geological hazards
In the fields of geotechnical engineering and disaster prevention and mitigation, the introduction of machine vision technology has greatly promoted the intellectualization of monitoring means [9]. Early studies mostly relied on background subtraction and optical flow methods, which are susceptible to illumination changes. With the widespread adoption of deep learning, CNN-based object detection models have become mainstream [10]. Ma and Meng [11] explored the automatic identification of landslide cracks integrating multimodal sensors (such as LiDAR and vision). Alonso-Díaz et al. [12] reviewed the application of machine learning and InSAR in geohazard monitoring, pointing out that the combination of physics-informed neural networks (PINNs) and vision models is a future development direction. However, most existing rockfall detection models are trained under ideal illumination and limited samples, and their generalization ability degrades significantly when facing complex granite slope backgrounds and severe weather.
2.2 Synthetic data and transfer learning in object detection
To address the problem of data scarcity in specific domains, synthetic data generation provides an effective pathway [13]. The ODGEN model published by Zhu et al. [14] at NeurIPS demonstrated the potential of conditional diffusion models in generating high-quality bounding-box-conditioned images. Luo [15] further explored the evolution from GANs to diffusion models in remote sensing domain adaptation. However, a domain shift inevitably exists between synthetic data and real-world data. To this end, transfer learning has been widely adopted. By performing source-domain pre-training on synthetic datasets, followed by target-domain fine-tuning on a small number of real datasets, the feature space can be effectively aligned [16].
2.3 Evolution of DETR-series algorithms
Since the introduction of DETR, the end-to-end object detection paradigm has eliminated post-processing steps such as NMS [17]. Subsequently, RT-DETR, proposed by Zhao et al. [18] addressed the high computational complexity of Transformer-based detectors through decoupled intra-scale interaction and cross-scale fusion, achieving real-time object detection. RF-DETR, proposed in 2025, further introduced the self-supervised vision foundation model DINOv2 as the backbone network, significantly improving feature extraction capability [8]. In addition, DN-DETR proposed by Li et al. [19] effectively alleviates the instability of bipartite matching by introducing a denoising query mechanism, greatly accelerating model convergence. Combining the characteristics of the slope scene, this paper deeply modifies RF-DETR by integrating the above advanced mechanisms.
3.1 Collection of the real dataset
The real dataset used in this paper was collected from the high-definition video surveillance systems deployed along a highway in southwestern China. Through video frame extraction, 5,200 images containing granite rockfall events were screened out. The images cover sunny, rainy, dusk, and vegetation occlusion scenes of varying degrees. The LabelImg tool was used to manually annotate the rockfall targets in the images, with a unified category of “rockfall”.
3.2 Synthetic data generation based on SDXL
Following the diffusion-based generative data augmentation paradigm [20, 21], this paper adopts SDXL combined with ControlNet to construct a synthetic dataset. Through carefully designed prompts (prompt engineering), such as “Photorealistic highway mountain slope, rainy weather, scattered granite boulders on asphalt road”, 8,000 high-quality synthetic images were generated.
Figure 3. Comparison matrix of synthetic and real data samples
As shown in Figure 3, the synthetic data exhibits extremely high realism in illumination rendering, rock texture, and environmental interaction. Evaluated by the FID (Fréchet Inception Distance), the distribution distance between generated and real images is only 18.4, demonstrating the effectiveness of the generative data.
The original RF-DETR algorithm performs excellently on general datasets (such as COCO), but its perception capability for small-scale rock fragments is insufficient when facing slope scenes with complex backgrounds. This paper makes three core improvements to its architecture (see Figure 4).
Figure 4. Overall architecture of the improved RF-DETR network
4.1 Backbone network and CBAM dual attention mechanism
The model adopts DINOv2 (a Vision Transformer) [22] with partially frozen weights as the backbone network, exploiting the strong representation capability obtained through self-supervised learning to extract multi-scale features (C3, C4, C5). To suppress the interference of natural granite textures, a CBAM [23] is embedded at the output of the C5 layer. CBAM adaptively recalibrates the feature maps sequentially through the channel attention module and the spatial attention module (see Figure 5).
Figure 5. Structural details of the CBAM attention module
4.2 BiFPN Cross-Scale Feature Fusion
The original RF-DETR employs the CCFM (Cross-Scale Cross-Attention Feature Module) for feature fusion. This paper replaces it with the BiFPN [24]. BiFPN introduces learnable node weights, allowing the network to dynamically adjust the contribution of features at different resolutions along the bidirectional top-down and bottom-up paths, thereby significantly improving the capability to capture small-scale rock fragments.
4.3 DN-query decoder and loss function derivation
In the Transformer decoder part, this paper introduces the multi-head self-attention mechanism. Its core computation process is shown as follows:
$Attention(Q, K, V)=softmax\left(\frac{Q K^T}{\sqrt{d k}}\right) V$
To accelerate model convergence and address the instability of bipartite matching in dense rockfall scenes, this paper introduces the denoising query mechanism proposed by DN-DETR [19]. Ground-truth bounding boxes are perturbed with noise and fed into the decoder as denoising queries to reconstruct the true boxes.
In the label assignment stage, the Hungarian algorithm is used to find the optimal bipartite matching between the prediction set and the ground-truth set:
$\sigma^{\wedge}=\arg \sum_{i=1}^N L_{\text {match }}\left(y_i, \hat{y}_{\sigma(i)}\right)$
where, the matching cost function comprehensively considers the class prediction probability, the L1 distance of the bounding boxes, and the generalized intersection over union (GIoU):
$\begin{gathered}L_{\text {match }}=-\lambda_{\text {cls }} \cdot \hat{p}_{\sigma(i)}\left(c_i\right)+\lambda_{L 1} \cdot\left\|b_i-\hat{b}_{\sigma(i)}\right\|_1+ \lambda_{\text {giou }} \cdot L_{\text {giou }}\left(b_i, \hat{b}_{\sigma(i)}\right)\end{gathered}$
The computation formula of GIoU is as follows, where C is the smallest enclosing region containing the predicted box A and the ground-truth box B:
$L_{giou}=1-\left(IoU-\frac{|C \backslash(A \cup B)|}{|C|}\right)$
5.1 Experimental environment and two-stage transfer learning strategy
The experiments were conducted on the Ubuntu 22.04 operating system with the PyTorch 2.1 deep learning framework, using dual NVIDIA RTX 4090 GPUs (24GB VRAM). Model training adopts a two-stage transfer learning strategy:
(1) Source-domain pre-training: the model is pre-trained on the 8,000 synthetic images (100 epochs), during which the DINOv2 backbone parameters are frozen and only the neck and decoder are updated, enabling the model to initially learn the geometric shapes and semantic features of rockfall.
(2) Target-domain fine-tuning: the pre-trained weights are loaded and the model is fine-tuned on the 5,200 real images (120 epochs). All parameters are unfrozen, with the learning rate set to 1e-4, using the AdamW optimizer and a cosine annealing learning rate decay strategy (see Figure 6).
Figure 6. Model training loss and validation set mAP curves
5.2 Comprehensive performance comparison experiments
To verify the advancement of the proposed algorithm, it is compared with the latest mainstream object detection algorithms, including YOLOv8-X, YOLOv12-X (the latest version in 2025), RT-DETR++ (2024 version), and Salience-DETR. The evaluation metrics include mAP@0.5, mAP@0.5:0.95, F1-score, and inference speed (FPS). The experimental results are shown in Table 1 and Figure 7.
Table 1. Performance comparison of different object detection models on the slope rockfall dataset
|
Model |
Backbone |
mAP@0.5 (%) |
mAP@0.5:0.95 (%) |
F1-Score (%) |
FPS |
|
YOLOv8-X |
CSPDarknet |
74.3 |
52.1 |
71.2 |
52.1 |
|
YOLOv12-X |
A2-Net |
81.5 |
58.7 |
78.4 |
44.8 |
|
RT-DETR++ |
HGNetv2 |
82.4 |
59.3 |
79.1 |
35.2 |
|
Salience-DETR |
ResNet50 |
80.1 |
57.4 |
77.0 |
28.6 |
|
Ours |
DINOv2 |
84.2 |
61.8 |
81.5 |
38.7 |
Figure 7. Bar chart of comprehensive performance comparison of different models
Figure 8 intuitively demonstrates the detection performance of each model under a sunny rock-wall background and rainy-night complex scenes. The YOLO-series algorithms show obvious missed and false detections under granite texture interference; in contrast, with the feature recalibration of CBAM and the global perception capability of Transformers, the proposed method accurately frames all rockfall targets of various scales.
Figure 8. Visualization comparison of detection results of different models in complex scenes
It can be seen from the experimental results that the proposed method achieves 84.2% on the mAP@0.5 metric, an improvement of 9.9 percentage points over the original baseline YOLOv8-X, and outperforms the latest YOLOv12-X (81.5%) and RT-DETR++ (82.4%). Although the DINOv2 backbone and attention mechanisms are introduced, benefiting from the lightweight design of the RF-DETR architecture, the inference speed of the proposed method remains at 38.7 FPS, fully meeting the real-time processing requirements of highway video surveillance (typically 25-30 FPS) (see Figure 9).
Figure 9. Accuracy-speed trade-off (Pareto frontier) analysis of different models
To deeply investigate the contribution of each proposed innovative component to the overall performance, detailed ablation experiments were designed. Taking the original non-pre-trained RF-DETR as the baseline model, the synthetic data, transfer learning strategy, CBAM attention mechanism, and BiFPN feature fusion module were introduced step by step. The experimental results are shown in Figure 10.
Figure 10. Ablation results of different components on model performance
(1) Benefit of synthetic data: the baseline model trained only on real data achieves an mAP@0.5 of only 65.3%. After directly mixing in synthetic data for joint training, the mAP increases to 74.7% (+9.4%), proving that synthetic data greatly enriches the long-tail feature space.
(2) Effect of two-stage transfer learning: after adopting the two-stage strategy of “synthetic-data pre-training + real-data fine-tuning”, the mAP further increases to 78.4% (+3.7%). This significant improvement indicates that the strategy effectively mitigates the domain shift between synthetic images and real monitoring images.
(3) Contribution of architecture improvements: after introducing the CBAM dual attention mechanism, the anti-interference capability of the model against background textures is enhanced, with the mAP increasing to 80.9% (+2.5%); after further replacing the neck with BiFPN and introducing DN-Query, the final complete model achieves the optimal mAP of 84.2%.
Figure 11. Effect of different mixing ratios of synthetic data on performance
In addition, Figure 11 shows the effect of the proportion of synthetic data on model performance. The model achieves the best performance when the proportion of synthetic data is 70%; when the proportion is too high (>80%), the real-domain features are excessively diluted, leading to performance degradation.
To address the problems of long-tail data scarcity and severe background interference in highway slope rockfall monitoring, this paper proposes an improved RF-DETR detection algorithm based on generative data transfer learning. The main conclusions are as follows:
(1) The conditional generation technology based on the SDXL diffusion model, combined with prompt engineering, can construct rockfall datasets covering extreme scenes at low cost and high quality, providing a new paradigm for data augmentation in visual monitoring of geological hazards.
(2) The two-stage domain-adaptation transfer learning strategy effectively bridges the distribution difference between synthetic and real data, significantly improving the generalization ability of the model in real field scenes.
(3) Introducing the CBAM attention mechanism and the BiFPN feature fusion network into the RF-DETR architecture effectively enhances the perception capability for small-scale rock fragments and suppresses granite background interference. While achieving an mAP@0.5 of 84.2%, the improved algorithm maintains a real-time inference speed of 38.7 FPS, achieving an excellent balance between accuracy and speed, and possessing engineering application value for deployment on edge computing devices.
This paper was funded by the Jiangxi Provincial Department of Science and Technology 03 Special Project (Grant No.: S2022ZXXMC0059; 20224ABC03A03); The Project of Jiangxi Provincial Department of Education (Grant No.: GJJ2405305); and the Science and Technology Project of Jiangxi Provincial Department of Transportation (Grant No.: 2024ZG015; 2024YB009; 2024YB008; 2025YB032).
[1] Zhong, F., Cui, W., Li, T., Lu, F., Pang, R. (2025). Slope rockfall detection with impact localization and motion classification based on RAL-YOLO. Expert Systems with Applications, 285: 127797. https://doi.org/10.1016/j.eswa.2025.127797
[2] Nguyen, K.A., Jiang, Y.J., Huang, C.S., Kuo, M.H., Chen, W. (2024). Leveraging internet news-based data for rockfall hazard susceptibility assessment on highways. Future Internet, 16(8): 299. https://doi.org/10.3390/fi16080299
[3] Gao, Q., Li, D.H., Ding, Y. (2024). A study on the rock block tracking algorithm that utilizes the M-DBT framework. Hydro-Science and Engineering, 2024(3): 166-176. https://doi.org/10.12170/20230606004
[4] Zhan, J., Zong, D., Yang, Y., Zhu, W., Peng, J. (2026) Automatic identification of loess landslide cracks based on deep learning algorithms and performance comparison. Journal of Engineering Geology, 33(6): 2160-2173. https://doi.org/10.13544/j.cnki.jeg.2025-0330
[5] Hussain, M. (2024). YOLOv1 to v8: Unveiling each variant-A comprehensive review of YOLO. IEEE Access, 12: 42816-42833. https://doi.org/10.1109/ACCESS.2024.3378568
[6] Guo, M.H., Xu, T.X., Liu, J.J., et al. (2022). Attention mechanisms in computer vision: A survey. Computational Visual Media, 8(3): 331-368. https://doi.org/10.1007/s41095-022-0271-y
[7] Khan, S., Naseer, M., Hayat, M., Zamir, S.W., Khan, F.S., Shah, M. (2022). Transformers in vision: A survey. ACM Computing Surveys, 54(10s): 1-41. https://doi.org/10.1145/3505244
[8] Robinson, I., Robicheaux, P., Popov, M., Ramanan, D., Peri, N. (2026). RF-DETR: Neural architecture search for real-time detection transformers. In International Conference on Learning Representations (ICLR), https://doi.org/10.48550/arXiv.2511.09554
[9] Ghorbanzadeh, O., Blaschke, T., Gholamnia, K., Meena, S.R., Tiede, D., Aryal, J. (2019). Evaluation of different machine learning methods and deep-learning convolutional neural networks for landslide detection. Remote Sensing, 11(2): 196. https://doi.org/10.3390/rs11020196
[10] Peng, P., Gao, L., Li, J., Zhang, H. (2025). Optimized YOLOv8 framework for intelligent rockfall detection on mountain roads. Scientific Reports, 15: 14007. https://doi.org/10.1038/s41598-025-94910-5
[11] Ma, W., Meng, B. (2026). Progressive fusion of infrared and visible light for road rockfall detection in mining areas. Opto-Electronic Engineering, 53(2): 87-101. https://doi.org/10.12086/oee.2026.250226
[12] Alonso-Díaz, A., Fontes, M., Teixeira, A.C., Wdowinski, S., Sousa, J.J. (2026). Multi-temporal InSAR and machine learning for geohazard monitoring: A systematic review with emphasis on noise mitigation and model transferability. Remote Sensing, 18(9): 1356. https://doi.org/10.3390/rs18091356
[13] Yang, L., Zhang, Z., Song, Y., et al. (2024). Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 56(4): 1-39. https://doi.org/10.1145/3626235
[14] Zhu, J., Li, S., Liu, Y., et al. (2024). ODGEN: Domain-specific object detection data generation with diffusion models. Advances in Neural Information Processing Systems, 37: 63599-63633.
[15] Luo, Y. (2026). From GANs to diffusion models: A progressive approach to synthetic data domain adaptation for remote sensing change detection. Masters Thesis, UNSW Sydney. https://doi.org/10.26190/unsworks/32480
[16] Wu, Q. (2025). Domain-adapted diffusion for synthetic data augmentation in UAV small object detection. In 2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), Chengdu, China, pp. 1-6. https://doi.org/10.1109/ICCWAMTIP68645.2025.11352629
[17] Zou, Z., Chen, K., Shi, Z., Guo, Y., Ye, J. (2023). Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3): 257-276. https://doi.org/10.1109/JPROC.2023.3238524
[18] Zhao, Y., Lv, W., Xu, S., et al. (2024). DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16965-16978.
[19] Li, F., Zhang, H., Liu, S., Guo, J., Ni, L.M., Zhang, L. (2023). DN-DETR: Accelerate DETR training by introducing query DeNoising. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(4): 2239-2251. https://doi.org/10.1109/TPAMI.2023.3335410
[20] Cao, H., Tan, C., Gao, Z., et al. (2024). A survey on generative diffusion models. IEEE Transactions on Knowledge and Data Engineering, 36(7): 2814-2830. https://doi.org/10.1109/TKDE.2024.3361474
[21] Li, K., Luo, S., Tseng, K.K. (2025). RegionDiffusion: Generative data augmentation for object detection with diffusion models. Neurocomputing, 654: 131115. https://doi.org/10.1016/j.neucom.2025.131115
[22] Han, K., Wang, Y., Chen, H., et al. (2023). A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1): 87-110. https://doi.org/10.1109/TPAMI.2022.3152247
[23] Woo, S., Park, J., Lee, J.Y., Kweon, I.S. (2018). CBAM: Convolutional block attention module. In European Conference on Computer Vision, pp. 3-19. https://doi.org/10.1007/978-3-030-01234-2_1
[24] Tan, M., Pang, R., Le, Q.V. (2020). Efficientdet: Scalable and efficient object detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, pp. 10778-10787. https://doi.org/10.1109/CVPR42600.2020.01079