INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTEGRATED CIRCUITS AND SYSTEMS The International Journal of Design, Analysis and Tools for Integrated Circuits and Systems (IJDATICS) was created by a netwo rk of researchers and engineers both from academia and industry. IJDATICS is an international journal intended for professionals and researchers in all fields of desig n, analysis and tools for integrated circuits and systems. The objective of the IJDATICS is to serve a better understanding between the community of researchers and practitioners both from academia and industry. Editor - In - Chief Ka Lok Man Xi'an Jiaotong - Liverpool University, China Associate Editor s Vijayakumar Nanjappan University College Cork, Ireland Jie Zhang Xi'an Jiaotong - Liverpool University , China Danny Hughes Katholieke Universiteit Leuven, Belgium Yuxuan Zhao Xi'an Jiaotong - Liverpool University, China Kamran Siddique University of Alaska Anchorage Hui - Huang Hsu Tamkang University, Taiwan M L Dennis Wong Heriot - Watt University, Scotland Tomas Krilavičius Vytautas Magnus University, Lithuania Young B. Park Dankook University, Korea Shuaibu Musa Adam Federal University Dutsin - Ma, Nigeria Shuhao Zhang Xi'an Jiaotong - Liverpool University, China Editorial Board Vladimir Hahanov Salah Merniz Kharkov National University of Radio Electronics, Ukraine Paolo Prinetto Politecnico di Torino, Italy Massimo Poncino Politecnico di Torino, Italy Alberto Macii Politecnico di Torino, Italy Joongho Choi University of Seoul, South Korea Wei Li Fudan University, China Michel Schellekens University College Cork, Ireland Emanuel Popovici University College Cork, Ireland Jong - Kug Seon LS Industrial Systems R&D Center, South Korea Umberto Rossi STMicroelectronics, Italy Franco Fummi University of Verona, Italy Graziano Pravadelli University of Verona, Italy Vladimir PavLov Intl. Software and Productivity Engineering Institute, USA Ajay Patel Intelligent Support Ltd, United Kingdom Thierry Vallee Georgia Southern University, USA Menouer Boubekeur University College Cork, Ireland Monica Donno Minteos, Italy Jun - Dong Cho Sung Kyun Kwan University, South Korea AHM Zahirul Alam International Islamic University Malaysia, Malaysia Gregory Provan University College Cork, Ireland Miroslav N. Velev Aries Design Automation, USA M. Nasir Uddin Lakehead University, Canada Dragan Bosnacki Eindhoven University of Technology, The Netherlands Dave Hickey University College Cork, Ireland Maria OKeeffe University College Cork, Ireland Milan Pastrnak Siemens IT Solutions and Services, Slovakia John Herbert University College Cork, Ireland Zhe - Ming Lu Sun Yat - Sen University, China Jeng - Shyang Pan National Kaohsiung University of Applied Sciences, Taiwan Chin - Chen Chang Feng Chia University, Taiwan Mong - Fong Horng Shu - Te University, Taiwan Liang Chen University of Northern British Columbia, Canada Chee - Peng Lim University of Science Malaysia, Malaysia Ngo Quoc Tao Vietnamese Academy of Science and Technology, Vietnam Mentouri University, Algeria Oscar Valero University of Balearic Islands, Spain Yang Yi Sun Yat - Sen University, China Damien Woods University of Seville, Spain Franck Vedrine CEA LIST, France Bruno Monsuez ENSTA, France Kang Yen Florida International University, USA Takenobu Matsuura Tokai University, Japan R. Timothy Edwards MultiGiG, Inc., USA Olga Tveretina Karlsruhe University, Germany Maria Helena Fino Universidade Nova De Lisboa, Portugal Adrian Patrick ORiordan University College Cork, Ireland Grzegorz Labiak University of Zielona Gora, Poland Jian Chang Texas Instruments Inc, USA Yeh - Ching Chung National Tsing - Hua University, Taiwan Anna Derezinska Warsaw University of Technology, Poland Kyoung - Rok Cho Chungbuk National University, South Korea Yong Zhang Shenzhen University, China R. Liutkevicius Vytautas Magnus University, Lithuania Yuanyuan Zeng University College Cork, Ireland D.P. Vasudevan University College Cork, Ireland Arkadiusz Bukowiec University of Zielona Gora, Poland Maziar Goudarzi University College Cork, Ireland Jin Song Dong National University of Singapore, Singapore Dhamin Al - Khalili Royal Military College of Canada, Canada Zainalabedin Navabi University of Tehran, Iran Lyudmila Zinchenko Bauman Moscow State Technical University, Russia Muhammad Almas Anjum National University of Sciences and Technology, Pakistan Deepak Laxmi Narasimha University of Malaya, Malaysia Danny Hughes Xi'an Jiaotong - Liverpool University, China Jun Wang Fujitsu Laboratories of America, Inc., USA A.P. Sathish Kumar PSG Institute of Advanced Studies, India N. Jaisankar VIT University. India Atif Mansoor National University of Sciences and Technology, Pakistan Steven Hollands Synopsys, Ireland Felipe Klein State University of Campinas, Brazil Enggee Lim Xi'an Jiaotong - Liverpool University, China Kevin Lee Murdoch University, Australia Prabhat Mahanti University of New Brunswick, Saint John, Canada Tammam Tillo Xi'an Jiaotong - Liverpool University, China Yanyan Wu Xi'an Jiaotong - Liverpool University, China Wen Chang Huang Kun Shan University, Taiwan Masahiro Sasaki The University of Tokyo, Japan Vineet Sahula Malaviya National Institute of Technology, India D. Boolchandani Malaviya National Institute of Technology, India Zhao Wang Xi'an Jiaotong - Liverpool University, China Shishir K. Shandilya NRI Institute of Information Science & Technology, India J.P.M. Voeten Eindhoven University of Technology, The Netherlands Wichian Sittiprapaporn Mahasarakham University, Thailand Aseem Gupta Freescale Semiconductor Inc., USA Kevin Marquet Verimag Laboratory, France Matthieu Moy Verimag Laboratory, France Ramy Iskander LIP6 Laboratory, France Suryaprasad Jayadevappa PES School of Engineering, India S. Hariharan B. S. Abdur Rahman University, India Chung - Ho Chen National Cheng - Kung University, Taiwan Kyung Ki Kim Daegu University, South Korea Shiho Kim Chungbuk National University, South Korea Hi Seok Kim Cheongju University, South Korea Siamak Mohammadi University of Tehran, Iran Brian Logan University of Nottingham, UK Ben Kwang - Mong Sim Gwangju Institute of Science & Technology, South Korea Asoke Nath St. Xavier's College, India Tharwon Arunuphaptrairong Chulalongkorn University, Thailand Shin - Ya Takahasi Fukuoka University, Japan Cheng C. Liu University of Wisconsin at Stout, USA Farhan Siddiqui Walden University, Minneapolis, USA Yui Fai Lam Hong Kong University of Science & Technology, Hong Kong Jinfeng Huang Philips & LiteOn Digital Solutions, The Netherlands Publisher Cooperation Name : Solari Co., Hong Kong Address : Unit 1 - 5, 20/F, Midas Plaza, 1 Tai Yau Street, San Po Kong, Kowloon, Hong Kong Phone : (852) 3966 - 2536 ISSN: 2071 - 2987 (online version), 2223 - 523X (print version) INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTEGRATED CIRCUITS AND SYSTEMS https://www.cicet.org/ijdatics / i Preface Welcome to the Volum e 1 5 Number 1 of the International Journal of Design, Analysis and Tools for Integrated Circuits and Systems (IJDATICS). This issue presents five high - quality academic papers, offering a well - rounded snapshot of current research in applied artificial intelligence, computer vision, natural language processing, and human - centered technology design. There are two key themes evident in these papers: • AI - Enhanced Detection and Classification Systems: Three papers highlight the growing role of AI in automating complex visual and textual recognition tasks across industrial, medical, and cybersecurity domains. • AI, VR, and Ethics in Applied Contexts : Two papers address the broader societal and pedagogical implications of emerging technologies. They underscore the need for human - centered design and responsible AI adoption beyond purely technical performance. We would also like to thank the IJDATICS editorial team, which is led by: Editor - I n - Chief Ka Lok Man Xi’an Jiaotong Liverpool University, China A ssociate Editors Vijayakumar Nanjappan University College Cork, Ireland Kamran Siddique University of Alaska Anchorage Jie Zhang Xi’an Jiaotong Liverpool University, China Yuxuan Zhao Xi'an Jiaotong - Liverpool University, China Shuaibu Musa Adam Federal University Dutsin - Ma, Nigeria Shuhao Zhang Xi'an Jiaotong - Liverpool University, China ii Table of Contents Vol. 15 , No. 1 , Au gust 20 2 6 Preface ................................................................................................. i Table of Contents ................................................................................... ii 1. Shu Jui Chang, Tim Watson, Iain Phillips and Andrew Peck, Symlets - Based Trainable Wavelet Downsampling for YOLOv12s in PCB Defect Inspection , Tamkang University , Taiwan 1 2. H sin - Yun Hsieh, Yi - Shiung Horng and Chii - Jen Chen , Automated Median Nerve Detection in Carpal Tunnel Ultrasound Using Improved YOLO - Based Model , Tamkang University , Taiwan 5 3. Lakshmi Indrasena Reddy Bandi, Po - Hung Yang , Network Traffic Prediction Using Temporal Correlation - Based LSTM Models , Tamkang University, Taiwan 9 4. J ean - Yves Le Corre and Mohanad Amin Salhab , Towards Instructional Design Frameworks for VR - Integrated Immersive Learning in Open - Source Ecosystems , Lincoln University College, Malaysia 15 5. A vantika Dube and Woon Kian Chong , Beyond the Algorithm: Perceived AI Authorship, Trust and Ethical Boundaries in AI Driven Advertising , SP Jain School of Global Management, Singapore 19 INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15 , NO. 1, AUGUST 202 6 1 Symlets - Based Trainable Wavelet Downsampling for YOLOv12s in PCB Defect Inspection Ching - Ming Chang and Chii - Jen Chen Abstract — Automated optical inspection (AOI) of printed circuit boards (PCBs) requires detectors that can accurately localize defects with diverse scales and irregular shapes while remaining lightweight enough for deployment on production lines. Recent YOLO - series detectors have achieved strong performance on PCB defect benchmarks; however, most improvements focus on attention mechanisms, feature fusion, or loss functions, whi le the downsampling operator in the neck is rarely revisited. In this paper, we propose a Sy mlets - based trainable wavelet downsampling module, termed SMD, and integrate it into YOLOv12s neck to better preserve high - frequency details in the feature maps SMD constructs 2 - D 4 N × 4 N kernels from analytic 1 - D Symlets filters and initializes the stride - 2 convolutions before the detection heads, allowing the network to learn task - adapted wavelet filters while preserving desirable multi - resolution properties. Experiments on the DsPCBSD+ dataset further compare two insertion locations(before the large - scale head and the middle - scale head), and results show that inserting SMD before the large - scale head provides the most favorable trade - off over the YOLOv12s baseline. Index Terms — Printed circuit board (PCB) defect detection, automated optical inspection (AOI), YOLOv12s, wavelet downsampling, Symlets , deep learning, object detection. I. INTRODUCTION Automated optical inspection ( AOI) of printed circuit boards ( PCBs) requires detectors that can accurately localize defects under complex textures and illumination variations This challenge is particularly evident in the DsPCBSD+ dataset, scratc hes and foreign_objects show high intra - class variability: scratches range from very thin and long traces to short marks and may overlap with other regions, while foreign_objects vary widely in size and shape. Recent advances in one - stage object detection, particularly the YOLO family, have made real - time industrial inspection increasingly practical by balancing accuracy and computational efficiency [ 1] Along with the progress of YOLO - style detectors, many PCB defect detection studies have improved performance by modifying attention mechanisms, feature - fusion necks, and localization losses. For instance, MAS - YOLO enhances YOLOv12 with attention and feature integration designs to improve PCB defect detection under challenging background interference [ 2] Meanwhile, public PCB defect datasets have been developed to support deep learning training and evaluation; DsPCBSD + is one representative dataset that categorizes PCB surface defects into multiple types to facilitate deep - learning - based inspection research [3] Despite these efforts, the downsampling operator in the neck - especially the stride - 2 convolution right before detection head -- has received comparatively less attention than attention mechanisms , fusion , and loss design. This is non - trivial for PCB AOI because downsampling operator determines how much high - frequency information (edges, thin traces, subtle defect boundaries) is preserved before multi - scale fusion and prediction. Motivated by wavelet theory, which provides a principled multi - resolution decomp osition for retaining both low - frequency structure and high - frequency details, we propose a Symlets - based trainable wavelet downsampling module, termed SMD , and integrate it into the YOLOv12s neck. Symlets wavelets offer compact support and higher - order vanishing moments, making them suitable for modeling both sharp defect edges and smoother contour variations. The main contributions are summarized as follows: 1. We propose SMD , a Symlets - based trainable wavelet downsampling module that replaces standard stride - 2 convolutions in the YOLOv12s neck. 2. We evaluate two insertion locations of SMD in the YOLOv12s neck (before the large - scale head and before the middle - scale head) to identify an effective accuracy – complexity trade - off. II. THE PROPOSED METHOD 2.1 Baseline: YOLOv12s Detector YOLOv12s follows the standard backbone – neck – head paradigm of YOLO - style one - stage detectors: the backbone extracts hierarchical features, the neck fuses multi - scale representations, and the detection heads predict bounding boxes and classes at multiple res olutions. In the baseline YOLOv12s, downsampling between neck stages is typically implemented by stride - 2 convolutions, which perform spatial decimation and learned filtering jointly. While effective and efficient, these layers do not explicitly encourage multi - resolution behavior that preserves high - frequency details important to thin, low - contrast PCB defects. 2.2 Symlet s - Based Trainable Wavelet Downsampling ( SMD ) The author's institution is as follows : Department of Computer Science and Information Engineering, Tamkang University, New Taipei City, Taiwan E mail: tianxinxin46@gmail.com , cjchen@mail.tku.edu.tw Corr es ponding author: Chii - Jen Chen INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15 , NO. 1, AUGUST 202 6 2 To better preserve high - frequency defect details during spatial downsampling, we design a Symlets - based trainable wavelet downsampling module, termed SMD For a given Symlets order N, we obtain analytic one - dimensional scaling and wavelet filters, denoted the 1 - D low - pass and high - pass filters as L N and H N , respectively. Each filter has length 4 N. Based on these 1 - D filters, four two - dimensional subband kernels are constructed via outer products: • K LL = L N ⊗ L N , K LH = L N ⊗ H N , • K HL = H N ⊗ L N , K HH = H N ⊗ H N yielding 4 N × 4 N kernels that form a wavelet - style multi - resolution decomposition when applied with stride 2. In SMD , the above Symlets kernels are used to initialize a depthwise stride - 2 convolution. For each input channel, the depthwise operation produces four subband responses (LL, LH, HL, HH), which are concatenated along the channel dimension. A subsequent 1 × 1 point wise convolution mixes the subband features and adjusts the output channel dimension to match the original YOLOv12s neck interface. All convolution weights remain trainable after initialization, enabling the module to retain the desirable multi - resolution behavior of Symlets transforms while adapting the filters end - to - end for PCB defect detection. 2.3 Depthwise and Pointwise Convolution Depthwise convolution applies one spatial kernel to each input channel independently. In SMD, the depthwise stride - 2 convolution is initialized by wavelet subband kernels (LL, LH, HL, HH) and operates channel - wise to preserve high - frequency information during downs ampling. Compared with a standard convolution, depthwise convolution significantly reduces the number of parameters from k^2· Cin· Cout to k^2· Cin, which is desirable for lightweight PCB inspection models ; P ointwise convolution is a 1× 1 convolution used to m ix information across channels. After the depthwise stage produces and concatenates multi - resolution subband responses, a pointwise convolution fuses these subband features and adjusts the channel dimension to match the original YOLOv12s neck interface. Th is design forms a depthwise – pointwise (separable) structure, whose parameter cost is k^2· Cin + Cin· Cout, providing an efficient way to integrate wavelet - style downsampling into YOLOv12s. 2.4 Depthwise – Pointwise Convolutions in Lightweight YOLO Variants Depthwise – pointwise (separable) convolutions have been widely adopted in lightweight YOLO variants to reduce parameters and computation while maintaining competitive accuracy. Early attempts such as YOLO - LITE [4] simplified the YOLO pipeline for non - GPU deployment, motivating subsequent designs that explicitly reorganize convolutional blocks toward shallower and narrower structures. Mixed YOLOv3 - LITE [5] further integrates lightweight backbone choices and restructured convolutional layers for efficiency. More recent s tudies extend this idea to YOLOv4/YOLOv5 - style frameworks by replacing standard convolutions in backbone/neck with depthwise – pointwise components, e.g., YoLite+ [6] and L - YOLOv4 [7] . These efforts support the practicality of separable designs inside YOLO; in contrast to purely efficiency - driven replacements, our module adopts a wavelet - initialized depthwise stage followed by pointwise fusion to better preserve informative details dur ing neck - stage downsampling. 2. 5 YOLOv12s Baseline Hyperparameter Details Tab le 1 summarizes the training settings used for the YOLOv12s baseline on DsPCBSD+, these hyperparameters are kept fixed across all baseline runs to provide a consistent reference for comparison with the proposed wavelet - based variant. Table 1 Training settings for the YOLOv12s baseline. Component Details Dataset DsPCBSD+ Input Resolution 640 ×640 pixels Epochs 100 Batch Size 16 Optimizer SGD warmup_epochs 2 lr0 0.01 momentum 0.9 2. 6 SMD - YOLOv12s Hyperparameter Details Table 2 lists the training settings for SMD - YOLOv12s. To ensure a fair comparison and isolate the effect of replacing the neck - stage stride - 2 downsampling layer with SMD, all hyperparameters are kept identical to the baseline except that the training sched ule is extended to 110 epochs. Table 2 Training settings for SMD - YOLOv12s Component Details Dataset DsPCBSD+ Input Resolution 640 ×640 pixels Epochs 110 Batch Size 16 Optimizer SGD warmup_epochs 2 lr0 0.01 momentum 0.9 III. EXPERIMENTAL RESULTS AND DISCUSSION 3. 1 DsPCBSD+ Dataset We conduct experiments on the DsPCBSD+ PCB surface defect dataset. The dataset contains 21,340 images in total, INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15 , NO. 1, AUGUST 202 6 3 where 19,698 images are used for training and 1,642 images are used for validation. All images are annotated with bounding boxes and defect categories in the YOLO detection format. DsPCBSD+ dataset covers nine defect categories For clear visualization, the nine defect categories are grouped into two figures according to their visual characteristics, including line - like/irregular - boundary patterns (as shown in Fig. 1 ) and blob - like/region - based patterns (as shown in Fig. 2 ). Notably, some categories (e.g., scratch - and foreign_object - related defects) exhibit high intra - class variability in scale and shape, which makes robust feature preservation during neck - stage downsampling particularly important for accurate PCB defect dete ction ( a ) conductor_scratch ( b )open ( c )short (d)spur ( e )mouse_bite Fig. 1 Representative defect examples in DsPCBSD+ with line - like or irregular boundary patterns: (a) conductor scratch, (b) open , (c) short , (d) spur , and (e) mouse_bite (a) spurious_copper (b) hole_breakout (c) conductor _foreign_object (d) base_material _foreign_object Fig. 2 Representative defect examples in DsPCBSD + with blob - like or region - based patterns: (a) spurious_copper, (b) hole breakout, (c) conductor_foreign_object, and (d) base_material_foreign_object. 3 2 Quantitative Results We evaluate the proposed SMD - enhanced YOLOv12s on the DsPCBSD+ dataset using Precision, mAP@0.5, and mAP@0.5:0.95. Table 3 reports the baseline YOLOv12s and two SMD insertion configurations: Location1 (before the large - scale head) and Location2 (before the middle - scale head). As shown in Table 3, Location1 improves Precision from 83.7% to 84.3% with only minor changes in mAP (86.3%→86.2% and 55.7%→55.5%). In contrast, Location2 does not improve Precision (83.1%) and yields slightly lower mAP than the baseli ne. Overall, t hese results indicate that SMD is insertion - dependent: replacing the downsampling operator before the large - scale head improves Precision, whereas inserting SMD before the middle - scale head does not provide additional benefits. This suggests that modifying the middle - scale downsampling path may interfere with subsequent feature fusion, leading to a less favorable trade - off. Table 3 The comparisons of YOLO12s baseline and SMD - YOLOv12s Model Precision mAP @0.5 mAP @0.5 :0.95 YOLOv 12 s (baseline) 83.7% 86.3% 55.7% YOLOv 12s + Location1 84.3% 86.2% 55. 5 % YOLOv 12s + Location2 8 3.1 % 86. 0 % 55. 5 % INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15 , NO. 1, AUGUST 202 6 4 3. 3 Discussion Table3 indicates that the effectiveness of SMD is insertion - dependent. Location1 (before the large - scale head) yields the best accuracy – complexity trade - off by improving Precision, while keeping mAP largely comparable to the baseline with a slight decrease. In co ntrast, inserting SMD before the middle - scale head (Location2) does not provide further gains and slightly degrades Precision and mAP. A plausible reason is that modifying the downsampling path at the middle - scale stage may interfere with the subsequent fe ature fusion process, whereas enhancing the large - scale branch better preserves discriminative structures for diverse PCB defect patterns. IV. CONCLUSIONS In this paper, we proposed SMD , a S y mlets - based trainable wavelet downsampling module, and integrated it into the YOLOv12s neck to revisit the design of stride - 2 downsampling operators for PCB defect detection. Experiments on the DsPCBSD+ dataset show that inserting SMD before the large - scale detection head (Location 1 ) achieves the best accuracy – complexity trade - off and yields consistent improvements over the YOLOv12s baseline , on the other hand, r eplacing the middle - scale downsampling layer (Location2) does not yield additional benefits, indicating that careful placement is critical when revisiting neck - stage downsampling operators. Future work will focus on improving robustness for highly variable defect categories and exploring more advanced integration strategies to further enhance multi - scale representation without introducing over - smoothing. REFERENCES [ 1] J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” arXiv preprint arXiv:1804.02767, 2018. [2] X. Yin, Z. Zhao, and L. Weng, “ .MAS - YOLO: A Lightweight Detection Algorithm for PCB Defect Detection Based on Improved YOLOv12,” Applied Sciences, vol. 15, no. 11, Art. no. 6238, 2025. [ 3 ] S. Lv, B. Ouyang, Z. Deng, T. Liang, S. Jiang, K. Zhang, et al., “A dataset for deep learning based detection of printed circuit board surface defect,” Scientific Data, vol. 11, no. 1, Art. no. 811, 2024. [4] R. Huang, J. Pedoeem, and C. Chen, “YOLO - LITE: A Real - Time Object Detection Algorithm Optimized for Non - GPU Computers,” in Proc. 2018 IEEE Int. Conf. on Big Data (Big Data), Seattle, WA, USA, Dec. 2018, pp. 2503 – 2510. [5] H. Zhao, Y. Zhou, L. Zhang, Y. Peng, X. Hu, H. Peng, and X. Cai, “Mixed YOLOv3 - LITE: A Lightweight Real - Time Object Detection Method,” Sensors, vol. 20, no. 7, Art. no. 1861, 2020, doi: 10.3390/s20071861. [6] Shuai , Y., Zhiyu, C., Shangdong, L., Mengxue, W., Feng, T., & Yimu, J. “YoLite+: a lightweight multi - object detection approach in traffic scenarios,” Procedia Computer Science, vol. 199, pp. 346 – 353, 2022, doi: 10.1016/j.procs.2022.01.042. [7] P. Ding, H. Qian, J. Bao, Y. Zhou, and S. Yan, “L - YOLOv4: lightweight YOLOv4 based on modified RFB - s and depthwise separable convolution for multi - target detection in complex scenes,” Journal of Real - Time Image Processing, vol. 20, Art. no. 71, 2023, d oi: 10.1007/s11554 - 023 - 01329 - 0. INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15, NO. 1, AUGUST 2026 5 Automated Median Nerve Detection in Carpal Tunnel Ultrasound Using Improved YOLO-Based Model a Hsin-Yun Hsieh, b,c Yi-Shiung Horng and a Chii-Jen Chen* Abstract This study proposes a deep learning framework that integrates YOLO11 with an improved segmentation model for the automated detection and segmentation of the median nerve cross-sectional area (CSA) in carpal tunnel ultrasound images. T o address the subjectivity and inefficiency of manual clinical analysis, the proposed approach aims to enhance the objectivity and consistency of the diagnostic process. Preliminary experimental results show that the method can be applied to the detection and segmentation of the median nerve, demonstrating its potential as a research direction for clinical application in the diagnosis of Carpal Tunnel Syndrome (CTS). Index Terms Carpal Tunnel Syndrome, Median nerve, Deep Learning, YOLO. I. INTRODUCTION Carpal Tunnel Syndrome (CTS) is a prevalent peripheral neuropathy caused by chronic compression of the median nerve (MN) within the carpal tunnel. Clinical symptoms, including numbness, pain, and loss of hand strength, significantly impair patients' daily activities and work productivity [1]. Epidemiological studies indicate that the global prevalence of CTS has reached 14.4% [2], with significantly higher incidence rates among workers in occupations involving repetitive hand movements or heavy gripping [3]. Because early-stage symptoms are often subtle, delayed diagnosis can lead to irreversible nerve damage, making timely and accurate screening critical for treatment outcomes. Recent advances have established ultrasound imaging as a primary clinical tool for evaluating MN cross-sectional area (CSA) due to its non-invasive, real-time, and cost- effective nature ! . However, manual interpretation requires extensive clinical experience and is inherently time-consuming and subjective, leading to inter- observer variability. Furthermore, the median nerve (MN) and adjacent structures, such as tendons and vessels, exhibit similar hypoechoic patterns in grayscale ultrasound, with boundaries frequently obscured by speckle noise or varying probe angles. Under these conditions, single-target segmentation models are prone to misidentifying vascular regions as nerve tissue, leading to significant measurement inaccuracies. Addressing these challenges through deep learning-based automated detection and segmentation is essential for enhancing diagnostic precision and minimizing human error in clinical decision-making. II. THE PROPOSED METHOD 2.1 Overview The automated image analysis pipeline proposed in this study is designed to address the clinical challenge of distinguishing the median nerve from surrounding tissues in ultrasound images. Initially, YOLO11 is employed for real- time detection of the median nerve and arteries. Due to the presence of significant background noise and non-target regions, the system utilizes the detected coordinates to automatically compute the union of the detected targets for ROI cropping. This preprocessing step condenses the original image into a focused Region of Interest (ROI), effectively reducing irrelevant interference and enabling the subsequent segmentation model to concentrate on fine- grained edge details of the target structures. In the segmentation stage, a YOLO-based detection model is first used to localize the region of interest (ROI), which is then cropped for further processing. Although YOLO is originally designed for object detection, its backbone and multi-scale feature representations can be extended to pixel-level prediction tasks, particularly through segmentation-oriented variants. This enables robust coarse localization of anatomical regions before fine segmentation. The cropped ROI is then processed by a dedicated segmentation network. In this study, a U-Net like architecture integrated with convolutional neural networks and Mamba-based state space modules (U-Mamba) is adopted for dense prediction. This design enhances both local feature extraction and long-range dependency modeling[5], which are important for biomedical image segmentation. By combining YOLO-based coarse localization with U-Mamba-based fine-grained segmentation in a two-stage pipeline, accurate delineation of structures such as the median nerve can be achieved. This framework may reduce inter-operator variability and assist in the assessment of Carpal Tunnel Syndrome. 2.2 Model Architecture The author's institution is as follows: a. Department of Computer Science and Information Engineering, Tamkang University, New Taipei City, Taiwan b. Department of Physical Medicine and Rehabilitation, Taipei Tzu Chi Hospital, Buddhist Tzu Chi Medical Foundation, Taipei, Taiwan c. Department of Medicine, Tzu Chi University, Hualien, Taiwan Email: elin3421@gmail.com, yshorng2015@gmail.com, cjchen@mail.tku.edu.tw Corresponding author: Chii-Jen Chen INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15, NO. 1, AUGUST 2026 6 The localization stage employs YOLO11 due to its high efficiency and favorable balance between speed and accuracy[6], making it suitable for real-time image analysis. YOLO11 generates accurate bounding boxes for the median nerve and arteries, enabling automatic ROI cropping. Its optimized architecture maintains reliable localization accuracy while delivering stable and efficient detection performance. The segmentation stage utilizes a deep learning based model that processes the cropped ROI to perform pixel-level prediction. This model incorporates multi-scale feature representation to capture rich contextual and boundary information, enabling it to delineate anatomical structures such as the median nerve within ultrasound images. By focusing on the ROI, the model effectively reduces interference from irrelevant background regions and improves segmentation stability under challenging imaging conditions. The two-stage framework, integrating detection and ROI-based segmentation, effectively minimizes interference from surrounding tissues[7]. 2.3 Segmentation Model Configuration The experimental settings and hyperparameter configurations of the segmentation model are summarized in Table 1. The dataset used in this study consists of private carpal tunnel ultrasound images clinically collected at Taipei Tzu Chi Hospital. The original images have a resolution of 532 434 pixels and are resized to a standard input size of 256 256 pixels to ensure computational efficiency and compatibility with the model architecture. The training process is conducted for 100 epochs with a batch size of 16. To achieve stable convergence toward a global optimum, the AdamW optimizer is adopted in combination with a cosine annealing learning rate schedule. Furthermore, to address the class imbalance commonly observed in medical imaging datasets, a hybrid loss function combining Weighted Cross-Entropy (WCE) and Weighted Dice Loss is employed. In this configuration, WCE emphasizes pixel-wise classification accuracy, while the weighted Dice loss focuses on improving overlap performance for small anatomical structures such as the median nerve. This dual-weighting strategy is designed to mitigate the impact of uneven sample distribution and ustness and consistency in segmenting clinically significant tissues. III. RESULTS AND DISCUSSIONS 3.1 Image Explanation Fig. 1 illustrates the ground truth segmentation masks for wrist ultrasound images, while Fig. 2 presents the ground truth YOLO bounding boxes, defining the spatial extent of the target objects. Fig. 3 shows the predicted segmentation masks generated by the trained model. In the predicted masks, each pixel is assigned a class label of 0, 1, or 2, where 0 (black) represents the background, 1 (blue) represents the median nerve, and 2 (red) represents the artery. Table 1. Experimental Components and Settings. Component Details Dataset A private carpal tunnel ultrasound dataset collected at Taipei Tzu Chi Hospital Image Size 532 × 434 pixels Model Input 256 × 256 pixels Epochs 100 Batch Size 16 Optimizer AdamW with a cosine annealing learning rate schedule Loss Functions Cross-Entropy Loss handles per-pixel classification, Dice Loss improves overlap accuracy, especially for small target regions Fig. 1. Ground truth segmentation of the median nerve and artery. Fig. 2. Ground truth YOLO bounding boxes. INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15, NO. 1, AUGUST 2026 7 Fig. 3. Final segmentation outputs of the proposed two-stage framework. 3.2 Quantitative Results To evaluate the effectiveness of the proposed two-stage framework, experiments were conducted on a private carpal tunnel ultrasound dataset using five-fold cross-validation. The framework consists of a detection stage based on YOLO11, followed by a segmentation stage using an improved YOLO-based model. Performance is evaluated separately for the detection and segmentation tasks. 3.2.1 Detection Performance Table 2 presents a comparison between YOLO11 and several widely used object detection models, including Faster R-CNN, YOLOv5, and YOLOv8. The evaluation metrics include mAP@0.5 and mAP@0.5:0.95. Among all compared models, YOLO11 achieves the best performance in terms of mAP@0.5 (80.50%), slightly outperforming YOLOv5 and YOLOv8. For mAP@0.5:0.95, YOLOv8 achieves the highest score (44.78%), while YOLO11 remains highly competitive with a score of 44.67%. In contrast, Faster R-CNN shows comparatively lower performance across both metrics. These results demonstrate that YOLO11 provides accurate and reliable bounding-box localization for the target structures (i.e., the median nerve and arteries), which is essential for subsequent ROI extraction and the downstream segmentation task. Table 2. Performance comparison of object detection models using five- fold cross-validation. Model mAP@0.5 mAP@0.5:0.95 Faster R-CNN 73.75% 39.59% YOLOv5 79.87% 42.56% YOLOv8 79.65% 44.78% YOLO11 80.50% 44.67% 3.2.2 Segmentation Performance Table 3 summarizes the average segmentation performance of the improved YOLO-based model across five-fold cross-validation. Evaluation metrics include the Dice coefficient and Intersection over Union (IoU) for three classes: background (C0), median nerve (C1), and artery (C2). The model achieves consistently high performance for the background class (C0), with Dice scores exceeding 0.98 across all folds. For the median nerve class (C1), the model demonstrates stable performance, with Dice scores of approximately 0.69. In contrast, the artery class (C2) exhibits significantly lower Dice scores (approximately 0.18), indicating the greater difficulty of accurately segmenting this small and less distinct anatomical structure in ultrasound images. Overall, the results show consistent performance across all folds, as reflected in the average metrics reported in Table 3. Table 3. Five-fold cross-validation results for median nerve and artery segmentation. Class Dice (avg) IoU (avg) Background (C0) 0.9921 0.9845 Median nerve (C1) 0.6930 0.5513 Artery (C2) 0.1832 0.1205 3.3 Model Results Analysis The experimental results demonstrate the effectiveness of the proposed two-stage framework, in which detection and segmentation are performed in a sequential manner. The accurate bounding boxes predicted by YOLO11 enable precise ROI extraction, effectively reducing background interference and allowing the improved YOLO-based segmentation model to focus on relevant anatomical structures. This sequential design improves both learning efficiency and segmentation stability. Class-wise analysis reveals distinct performance variations across different anatomical structures. The background class (C0) consistently achieves high accuracy due to its dominant presence and relatively homogeneous visual characteristics. The median nerve (C1) exhibits stable segmentation performance, indicating that the model is able to effectively capture its structural features. In contrast, the artery class (C2) remains challenging to segment due to its small size, low contrast in ultrasound images, and limited representation within the dataset. In addition, inaccuracies in detection may lead to suboptimal ROI extraction, further affecting segmentation performance for this class. To qualitatively evaluate the results, Fig. 4 presents representative examples, including YOLO-generated cropped images, corresponding ground truth masks, and predicted segmentation outputs from the proposed framework. The visual results indicate that ROI cropping effectively suppresses background noise and narrows the focus to target anatomical regions. While the model is INTERNATIONAL JOURNAL OF DESIGN, ANALYSIS AND TOOLS FOR INTERGRATED CIRCUITS AND SYSTEMS, VOL. 15, NO. 1, AUGUST 2026 8 generally capable of localizing the median nerve, the overlap with ground truth is not always consistent. The artery is particularly difficult to segment, especially in low-contrast or small-scale regions. Misclassifications are mainly observed at tissue boundaries or in areas with ambiguous texture. These qualitative observations are consistent with the quantitative results reported in Section 3.2.2, and highlight the current limitations of the proposed two-stage framework in handling fine-grained anatomical details in challenging ultrasound images. (a) (b) (c) Fig. 4. Visualization of segmentation results: (a) YOLO-cropped image, (b) ground truth mask, and (c) predicted mask. 3.4 Discussion The experimental results demonstrate the effectiveness of the proposed two-stage framework for wrist ultrasound analysis. The consistent Dice scores achieved for the median nerve indicate that accurate ROI extraction based on YOLO11 detection enables the segmentation model to focus on relevant anatomical structures, thereby contributing to stable performance. In contrast, segmentation of the artery remains challenging due to its small size, low contrast in ultrasound images, and limited representation within the dataset. In addition, errors in the detection stage may propagate to the segmentation stage through suboptimal ROI extraction, further affecting performance. These findings highlight the importance of improving both small object detection and fine-grained segmentation strategies in ultrasound image analysis. Overall, the proposed framework is effective for relatively prominent anatomical structures, while further refinement is required for accurate delineation of small and less distinct reg