ScaleCeph: A Two-Stage Coarse-to-Fine ROI Cascade for Cross-Device Cephalometric Landmark Localization

Authors

  • Ubaid Ur Rehman University of Engineering and Technology, Taxila.
  • Syed M. Adnan University of Engineering and Technology, Taxila
  • Wakeel Ahmad University of Engineering and Technology, Taxila
  • Huzaifa Ehsan University of Engineering and Technology, Taxila
  • M. Nauman Tariq University of Engineering and Technology, Taxila

Keywords:

Automated Cephalometric Analysis, Cephalometric Landmark Localization, Cross-Device Robustness, Deep Learning, HRNet

Abstract

Cephalometric landmark localization is an important part of orthodontic diagnosis and treatment planning, but manual annotation is slow and has notable inter-observer variation. Automated methods have advanced rapidly, yet most are developed and tested on images from a single device and report a single aggregate error, leaving open how they behave when a clinic pools radiographs from several machines with different fields of view, contrast, and pixel spacing. We present ScaleCeph, a two-stage coarse-to-fine cascade that treats this cross-device variation as the central problem. A first HRNet-W32 localizes 19 landmarks on the whole radiograph; the bounding box of these predictions defines a craniofacial region of interest that is cropped and magnified to a canonical scale, absorbing the inter-device differences, and a second, independently trained network refines the landmarks on this magnified view. Coordinates are decoded with a robust windowed soft-argmax decoder, and ablations show that the design choices, prediction-driven crops, sharp σ = 2 heatmap targets with a weighted-BCE loss, and the windowed decoder, are each necessary rather than incidental. On the Aariz benchmark (1,000 cephalograms, seven devices), ScaleCeph attains a mean radial error (MRE) of 1.076 ± 1.18 mm (95% CI 0.99–1.16 mm) and 88.8% successful detection rate (SDR) within 2 mm (95% CI 87.7–90.0), against the single-stage baseline (1.252 ± 1.17 mm, 85.1%; paired Wilcoxon signed-rank p < 10⁻²⁰), and raises the macro-averaged per-device SDR@2 mm from 85.5% to 89.5%. On the best-resolved machines performance approaches inter-observer agreement, supporting prediction-driven scale normalization for robust cross-device localization.

Author Biographies

Ubaid Ur Rehman, University of Engineering and Technology, Taxila.

MS Computer Science, Dept. of Computer Sciene, UET Taxila.

Personal email: ubaidurrehman721@gmail.com

24-MS-CS-01@uettaxila.students.edu.pk

Syed M. Adnan, University of Engineering and Technology, Taxila

Designation Assistant Professor Department Dept. of Computer Science, UET Taxila. Highest Qualification Ph.D. Computer Engg. UET Taxila

Wakeel Ahmad, University of Engineering and Technology, Taxila

Designation Lecturer Department Dept. of Computer Science, UET Taxila Highest Qualification Ph.D. Computer Science

References

“Does orthodontic treatment affect patients’ quality of life? - PubMed.” Accessed: Aug. 10, 2026. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/18676797/

J. Monisha, U. Sangeetha, B. Nivethitha, and B. Madhan, “Agreement between cephalometric analyses in diagnosing the dento-skeletal characteristics of malocclusion,” J. Oral Biol. Craniofacial Res., vol. 15, no. 4, pp. 744–748, Jul. 2025, doi: 10.1016/j.jobcr.2025.04.012.

H. Zhang et al., “Deep Learning Techniques for Automatic Lateral X-ray Cephalometric Landmark Detection: Is the Problem Solved?,” Sep. 2024, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2409.15834

Q. Chang, Z. Wang, F. Wang, J. Dou, Y. Zhang, and Y. Bai, “Automatic analysis of lateral cephalograms based on high-resolution net,” Am. J. Orthod. Dentofac. Orthop., vol. 163, no. 4, pp. 501-508.e4, Apr. 2023, doi: 10.1016/J.AJODO.2022.02.020.

M. Han et al., “Automated Landmark Detection and Lip Thickness Classification Using a Convolutional Neural Network in Lateral Cephalometric Radiographs,” Diagnostics 2025, Vol. 15, Page 1468, vol. 15, no. 12, p. 1468, Jun. 2025, doi: 10.3390/DIAGNOSTICS15121468.

M. S. I. Sumon et al., “Self-CephaloNet: a two-stage novel framework using operational neural network for cephalometric analysis,” Neural Comput. Appl. 2025 3716, vol. 37, no. 16, pp. 9777–9805, Mar. 2025, doi: 10.1007/S00521-025-11097-6.

Y. Shimamura et al., “Accuracy of cephalometric landmark and cephalometric analysis from lateral facial photograph by using CNN-based algorithm,” Sci. Reports 2024 141, vol. 14, no. 1, pp. 31089-, Dec. 2024, doi: 10.1038/s41598-024-82230-z.

S. H. Han et al., “Accuracy of posteroanterior cephalogram landmarks and measurements identification using a cascaded convolutional neural network algorithm: A multicenter study,” Korean J. Orthod., vol. 54, no. 1, pp. 48–58, 2024, doi: 10.4041/KJOD23.075.

S. Kang, I. Kim, Y. J. Kim, N. Kim, S. H. Baek, and S. J. Sung, “Accuracy and clinical validity of automated cephalometric analysis using convolutional neural networks,” Orthod. Craniofac. Res., vol. 27, no. 1, pp. 64–77, Feb. 2024, doi: 10.1111/OCR.12683.

R. Khan, M. A. Khalid, K. Zulfiqar, U. Bashir, and M. M. Fraz, “Enhancing Cephalometric Landmark Detection with a Two-Stage Cascaded CNN on Multi-resolution Multi-modal Data,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 14860 LNCS, pp. 3–18, 2024, doi: 10.1007/978-3-031-66958-3_1/SAVE-RESEARCH.

M. A. Khalid, A. Khurshid, K. Zulfiqar, U. Bashir, and M. M. Fraz, “A two-stage regression framework for automated cephalometric landmark detection incorporating semantically fused anatomical features and multi-head refinement loss,” Expert Syst. Appl., vol. 255, p. 124840, Dec. 2024, doi: 10.1016/J.ESWA.2024.124840.

S. Rashmi, S. Srinath, R. Rakshitha, and B. V. Poornima, “Ensemble learning methods with single and multi-model deep learning approaches for cephalometric landmark annotation,” Discov. Artif. Intell. 2024 41, vol. 4, no. 1, pp. 93-, Nov. 2024, doi: 10.1007/S44163-024-00207-3.

H. Zhu, Q. Yao, L. Xiao, and S. K. Zhou, “You Only Learn Once: Universal Anatomical Landmark Detection,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 12905 LNCS, pp. 85–95, Mar. 2022, doi: 10.1007/978-3-030-87240-3_9.

Z. Zhu et al., “An ensemble-based deep learning method through multi-scale cross-attention training for cephalometric landmark localization on lateral X-ray images,” Mar. 2025, doi: 10.21203/RS.3.RS-6105085/V1.

Z. Jiao et al., “Deep learning for automatic detection of cephalometric landmarks on lateral cephalometric radiographs using the Mask Region-based Convolutional Neural Network: a pilot study,” Oral Surg. Oral Med. Oral Pathol. Oral Radiol., vol. 137, no. 5, pp. 554–562, May 2024, doi: 10.1016/j.oooo.2024.02.003.

S. A. Prativi, A. B. Suksmono, T. L. E. Rajab, D. Danudirdjo, and A. Laviana, “Automatic landmark detection in cephalometric lateral radiograph using deep learning one-stage detectors,” 2024 14th Int. Conf. Syst. Eng. Technol. ICSET 2024 - Proceeding, pp. 67–72, 2024, doi: 10.1109/ICSET63729.2024.10775273.

I. Tafala, F. E. Ben-Bouazza, A. Edder, O. Manchadi, M. Et-Taoussi, and B. Jioudi, “Cephalometric Landmarks Identification Through an Object Detection-based Deep Learning Model,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 2, pp. 859–867, 2024, doi: 10.14569/IJACSA.2024.0150286.

A. Jaheen et al., “CephRes-MHNet: A Multi-Head Residual Network for Accurate and Robust Cephalometric Landmark Detection,” Nov. 2025, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/abs/2511.10173v2

F. Laitenberger, H. T. Scheuer, H. A. Scheuer, E. Lilienthal, S. You, and R. E. Friedrich, “Cephalometric landmark detection using vision transformers with direct coordinate prediction,” J. Cranio-Maxillofacial Surg., vol. 53, no. 9, pp. 1518–1529, Sep. 2025, doi: 10.1016/J.JCMS.2025.05.021.

H. Wu et al., “Cephalometric Landmark Detection across Ages with Prototypical Network,” Jun. 2024, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2406.12577

M. A. Khalid et al., “A Benchmark Dataset for Automatic Cephalometric Landmark Detection and CVM Stage Classification,” Sci. Data 2025 121, vol. 12, no. 1, pp. 1336-, Jul. 2025, doi: 10.1038/s41597-025-05542-3.

C. W. Wang et al., “A benchmark for comparison of dental radiography analysis algorithms,” Med. Image Anal., vol. 31, pp. 63–76, Jul. 2016, doi: 10.1016/J.MEDIA.2016.02.004.

J. Wang et al., “Deep High-Resolution Representation Learning for Visual Recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 10, pp. 3349–3364, Aug. 2019, doi: 10.1109/TPAMI.2020.2983686.

Q. Wu, S. Y. Yeo, Y. Chen, and J. Liu, “Revisiting Cephalometric Landmark Detection from the view of Human Pose Estimation with Lightweight Super-Resolution Head,” Sep. 2023, Accessed: Aug. 10, 2026. [Online]. Available: https://arxiv.org/pdf/2309.17143

R. Chen, Y. Ma, N. Chen, D. Lee, and W. Wang, “Cephalometric Landmark Detection by AttentiveFeature Pyramid Fusion and Regression-Voting,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 11766 LNCS, pp. 873–881, Aug. 2019, doi: 10.1007/978-3-030-32248-9_97.

Downloads

Published

2026-08-20

How to Cite

Ubaid Ur Rehman, Syed M. Adnan, Wakeel Ahmad, Huzaifa Ehsan, & M. Nauman Tariq. (2026). ScaleCeph: A Two-Stage Coarse-to-Fine ROI Cascade for Cross-Device Cephalometric Landmark Localization. International Journal of Innovations in Science & Technology, 8(5), 2081–2097. Retrieved from https://journal.50sea.com/index.php/IJIST/article/view/1973

Most read articles by the same author(s)