Utilization of Image Analysis Techniques for Interpretation of Changes in Visual Patterns: An Overview

Utilization of Image Analysis Techniques for Interpretation of Changes in Visual Patterns: An Overview

Salwa S. Moustafa* | Fathi E. Abd El-Samie | Hossam M. Faheem | Nabil A. Ismail

Department of Computer Science & Engineering, Faculty of Electronic Engineering, Menoufia University, Menouf 32952, Egypt

Department of Electronics and Electrical Communications Engineering, Faculty of Electronic Engineering, Menoufia University, Menouf 32952, Egypt

Faculty of Computers & Informatics, Ain Shams University, Cairo 11566, Egypt

Corresponding Author Email: 
salwa1975eg@yahoo.com
Page: 
2047-2059
|
DOI: 
https://doi.org/10.18280/ts.430434
Received: 
19 February 2026
|
Revised: 
27 May 2026
|
Accepted: 
9 June 2026
|
Available online: 
31 August 2026
| Citation

© 2026 The authors. This article is published by IIETA and is licensed under the CC BY 4.0 license (http://creativecommons.org/licenses/by/4.0/).

OPEN ACCESS

Abstract: 

Image analysis for perfect object detection is considered one of the most important and challenging modern fields of technology in our time at all, which is related to the various tasks of computer vision (CV), image processing, and artificial intelligence (AI). In addition, to a continuous and noticeable development in the mechanisms and techniques of object detection and recognition in high-resolution visual imagery and in low-resolution infrared (IR) and X-ray imagery. The main goal of the task of image processing for object detection techniques is to accurately identify the position, size, and boundaries of one or more objects within a certain image or video. This task involves multiple steps for the subtasks of different techniques like image classification, image recognition, image localization, image segmentation, image enhancement, and so on. In the last few years, the rapid advances and development in deep learning (DL) techniques have greatly accelerated the momentum for better utilization of multiple subtasks of object detection technology. Therefore, the performance of object detection for visual image processors and trackers has greatly improved, achieving significant breakthroughs in object detection.

Keywords: 

image processing, object detection, machine learning, deep learning, artificial intelligence, computer vision, infrared and X-ray images, low-quality resolution

1. Introduction

A crucial computer vision (CV) job and one of the most important areas of artificial intelligence (AI) nowadays is object detection. As a result, a major factor in its rapid overall development is the improvement of object detection performance. Deep learning (DL) object detection and tracking, which are the core building blocks of many contemporary CV applications, are among the most recent technological developments in CV [1-10]. For instance, object detection enables anomaly detection, robot vision, intelligent healthcare monitoring, autonomous driving, and more. Every AI vision application often calls for a blend of various algorithms that create a pipeline with a number of processing steps.

The aim of the object detection task is to detect and locate one or more objects of interest in visual images or videos [11-20], as shown in Figure 1 and Figure 2. The task is to determine the location and boundaries of objects in the image and categorize those objects. Every object class has its own special features that help in classifying the class; therefore to deal with detecting instance objects of a certain class (such as humans, animals, buildings, cars, etc.) in proposal images and videos [21-38].  So, each class object detection task depends on specific features that are appropriate for the algorithm used [39-58]. In this paper, we focus on clarifying and explaining the nature of the overlap and intertwining between the different subtasks, such as image classification, image recognition, image localization, image segmentation, image enhancement, and more [6, 19-21]. This approach is clarified in Figure 3, and it is needed to improve the overall performance of the detection [59-71].

Figure 1. Sample for detection of one and multiple objects on a road

Figure 2. Block diagram showing the detection of a specified object in a visual image with determination of the feature points

Machine learning (ML) is a subfield of AI that primarily entails pattern learning from examples or sample data as the machine accesses the data and learns from it. On the other hand, DL is a particular type of ML that incorporates learning at several levels. In general, in our research paper, we reveal that the improvements in object detection techniques continue to surpass the performance gains of all their subtasks [39]. Object detection from normal and low-quality images is illustrated with the examples in Figure 1 and Figure 2.

Applications for object detection, or instances in which object detection is employed, are incredibly flexible; there are practically endless ways to let computers see like people in order to automate manual chores or develop brand-new, AI-powered products and services. It has been included in CV applications used for many different purposes [6-11, 20] for the most popular and recent CV applications [60].

Advantages and Disadvantages of Object Detection: Object detectors are very adaptable and can be trained for a variety of jobs and unique, special-purpose applications. Automating processes (counting, inspection, verification, etc.) throughout company value chains might benefit from the automatic identification of things, people, and scenes [56, 57]. However, the main drawbacks of object detectors are that they are computationally expensive and need a lot of computing power. Operating expenses can quickly rise, particularly when object detection algorithms are used at scale [37, 42], posing a threat to the economic feasibility of corporate use cases.

The field of image processing and object detection tasks for several degrees of resolution in the assumed image, whether it is visual, IR, or X-ray, is not easy to accomplish. Especially with detection in low-quality images like IR and X-ray. So, we need to go ahead with more advanced evolution in this field with the purpose of finding and locating instances of semantic objects of a certain class in both types of specified images: high-quality resolution of visual images with specific features, and low-quality resolution (IR and X-ray) images.

For this reason, the branch of image processing and analysis for different resolution patterns has become an extremely important technology because it has been widely used in all fields of people’s lives, for example, in fields of machine vision, intelligent transportation, security monitoring, autonomous driving, aerospace, military, medical diagnoses for healthcare treatment, and so on.

This research paper will provide an introduction to object detection technology and detector models of detectors to achieve fast and accurate object detection, whether we deal with high-quality resolution images or low-quality resolution images. It is considered a fundamental requirement for CV, ML, DL, and AI. Secondly, it provides an overview of the main steps of object detection algorithms and subtask techniques related to perfect detection [72]. Thirdly, the development of object detection and DL, and finally a brief history of low-quality images and especially the techniques of IR and X-ray images are presented.

1.1 Object detection – features – implementation

The object detection task steps in our research paper are given. We detect and locate the normal and visually specified object by determining the special feature points in the image as described in Figure 2.

We determine and visualize the matched feature points of the specified object in the proposed image of forest. As shown in the block diagram in the following Figure 2.

2. Main Steps of the Object Detection

The object detection task needs many subtasks combining image classification and object localization to identify, categorize, and classify objects in an image or video [59-61]. The following key steps constitute the basic object detection process:

•Input one or multiple images into an algorithm of object detection [5-8, 11].

•This algorithm analyzes the given image(s).

•The algorithm sets a bounding box around the object to identify its location and specifies a type and a class label for that object using image classification.

•After the step of determining the type and classification of the object, the algorithm shows where the object is in the image.

•Then, the algorithm outputs the location, size, type, and class label for each object in the image(s).

There are several different object detection algorithms, including R-CNN [18], Fast R-CNN [14], YOLO [41, 54, 55], and Single Shot MultiBox Detector (SSD) [13, 62]. These algorithms use a variety of techniques, including CNNs [5-8, 11], region-based fully convolutional networks, etc.

3. Relations and Connections Between the Terminology of Object Detection and the Multiple Subtasks in CV and Artificial Intelligence

To precisely detect, locate, and identify objects in an image, we need to understand the methods based on DL and ML to allow computers to recognize and discover objects in digitally specified images or video [63, 68]. The basic objective of object detection is to precisely identify one or more objects in a given image or video, as well as the position, size, and class label of each [69]. This method entails many key steps of subtasks, which are implicitly presented, like image classification, image recognition, image localization, image segmentation, image enhancement, and so on [19-21].

3.1 Object detection versus image classification

The process of labeling, categorizing, and classifying an image as containing a specific object based on established guidelines helps to clarify what it includes, clearly. An image of a zebra, for example, might be labeled as a “zebra”, and an image of a giraffe can be labeled as a “giraffe”, as can be seen in Figures 3 and 4.

Figure 3. Image classification vs. object detection vs. image segmentation

Figure 4. Description of image classification, image localization, image segmentation, and object detection into a different image shape

Convolutional neural networks (CNNs) and other DL models are frequently used in image classification to assess and classify images based on their visual features and characteristics [11, 20]. Therefore, the classification has the benefit of providing a better option for tags like "blurry" or "sunny" that do not actually have physical bounds [66].

However, for recognizable objects, object detection systems generally outperform standalone classification networks by incorporating spatial localization alongside categorization [52].

Simply, object detection is a little more sophisticated because it draws a bounding box around the object that has been detected. After classifying an image, the classifier works on the entire image to produce a tag. The classifier analyzes the entire image, and the tag appears when it is appropriate.

3.2 Object detection versus image localization & image recognition

Finding a specific object position and size within an image is known as object localization. To make it simpler for machines to recognize the object, a bounding box is drawn around it and a class label is given to it [10, 11, 19].

Object detection and image recognition are sometimes confounded; however, there are some important distinctions between both of them, as shown in Figure 4.

Without localizing or pinpointing the actual location of the objects, image recognition refers to the analysis of a full image and identification of the things therein. Contrarily, object detection entails both object recognition and object localization within an image [35]. Comparatively, object detection is a more challenging task than image recognition.

For instance, while image recognition can identify and classify a specific object within an image, it cannot pinpoint its exact location. Object detection, on the other hand, goes a step further by not only recognizing the object but also determining its precise position and outlining it with a bounding box [70].

3.3 Object detection versus image segmentation

Image segmentation is the process of identifying which pixels in an image belong to a specific object class. While semantic segmentation labels every pixel associated with a given class, it does not differentiate between individual instances of the same entity.

Instead, object detection will clearly indicate the location of each unique object with a box rather than segmenting the object.

Instance segmentation bridges semantic segmentation and object detection by first locating object instances and then segmenting each within its detected bounding box [21].

To effectively recognize objects and classify an image, it is essential to first understand its structural composition. As illustrated in Figure 3, image segmentation serves this purpose by partitioning the image into meaningful components.

The image can be divided or partitioned into a number of segments. Since some areas of the image will have no information, processing the full image at once is not a good idea. We can use the crucial portions of the image after segmenting it for processing.

Essentially, successful image segmentation partitions an image into distinct pixel sets based on shared attributes. Unlike object detection models, which merely enclose entities within rectangular or square bounding boxes that omit structural shape details, segmentation models generate pixel-wise masks for each object. This pixel-level precision provides a significantly more accurate representation of object geometry within the image [1, 19-21].

With this knowledge, we want to have a better understanding of image categorization, localization, object recognition, and segmentation, as illustrated in Figure 3 and Figure 4.

Briefly, image classification enables us to categorize the information included in an image. While object detection pinpoints the locations of many things in an image, image localization only specifies the location of a single object in an image [35]. A pixel-by-pixel mask of each object in the image is created using image segmentation at the end. Image segmentation will enable us to recognize the forms of various objects in the image.

3.4 Object detection versus image enhancement

Techniques for image enhancement are becoming more and more necessary to boost image recognition and object detection.

CNN-based object detection is a popular area of study in CV. In the real world, hotspots for object detection technologies include autonomous driving, pedestrian detection, and other uses [27, 29-34].

Academic object detection research has advanced significantly, especially since DL became popular. However, images taken in the real world frequently include a variety of quality issues, such as low light, low resolution, and color distortion, which greatly affect how well different detection algorithms work [46]. It is a common practice to firstly employ certain enhancement techniques to recover a high-quality image from the original image, followed by object detection on the recovered image [27, 29-34].

Figure 5. Description of infrared (IR) image enhancement, which is considered a preprocessing task before the task of object detection and recognition in different IR image shapes

As shown in Figure 5, when light levels are low, the lighting (illumination) aspect of the image has a big influence on object detection quality. As a pre-processing technique, low-light image enhancement techniques can increase image quality and yield better detection results [49-51].

However, the complexity of low-light situations may cause the present enhancement strategies to negatively affect some samples. Consequently, improving the total detection performance in low-light conditions is difficult. There are several factors, such as fog, rain, a sudden shift in illumination, or lack of illumination, that reduce the quality of captured images. Object identification and recognition for cars, people, fixed objects, and traffic lights may not succeed if the quality of the obtained images is compromised. Numerous image enhancement strategies have been suggested and assessed in order to increase the rate of object identification, recognition, and detection [43, 67].

Nowadays, several academic research works concentrate on image enhancement techniques to enable more advanced techniques for the object detection and recognition environment [27, 29-34].

3.4.1 Methods and techniques

There are many commonly used methods to achieve the performer’s goal of image enhancement and improvement, and we will review some of them in this section:

* Image enhancement based on histogram matching; this method will rely solely on the histogram matching (HM) tool, which will be utilized to improve the enhancement of IR images. It is noteworthy that visible images will have a better histogram distribution than other IR images, which are known to have band-limited histograms. In order to solve this, we might consider altering the histogram that depicts the IR image, which is dispersed throughout a specific range of the visible images. This would improve the IR image visual quality. The procedure described above, called histogram matching, involves altering the mean and specified variance of each IR image based on the counterparts in the particular visible images. The main idea of HM depends on stretching the low-contrast IR image histogram to match a high-contrast visible image, which is rich in details. We can use the mathematical model for the calculation steps and equations of HM summarized as follows:

Compute the mean of the original IR image from g(m,n)

$\hat{g}=\sum_{m=1}^M \sum_{n=1}^N g(m, n)$

where, M and N represent the dimensions of the IR images. 

Compute the mean of the reference visual image f(k, l)

$\hat{f}=\sum_{k=1}^K \sum_{l=1}^L f(k, l)$

where, K and L are the dimensions of the reference image.

Estimate the standard deviation of the IR image r1 ($\sigma_1$).

$\sigma_{1=} \sqrt{\frac{1}{M N} \sum_{m=1}^M \sum_{n=1}^N(g(m, n)-\hat{g})^2}$

Estimate of the STDEV of the reference image r2 ($\sigma_2$).

$\sigma_2=\sqrt{\frac{1}{K L} \sum_{k=1}^K \sum_{l=1}^L(f(k, l)-\hat{f})^2}$

Estimate the correction factor C by dividing the standard deviation of the reference image by the standard deviation of the IR image.

$C=\frac{\sigma_2}{\sigma_1}$

Compute the modified mean factor $f_c$.

$f_c=\hat{f}-C * \hat{g}$

Estimate the matched histogram $F_H$.

$F_H=f_c+f * C$

* Image enhancement based on fuzzy logic methods; many kinds of fuzzy image enhancement methods have been proposed in the fields of research, including: fuzzy contrast adjustment, subjective image enhancement, fuzzy image segmentation, fuzzy edge detection, and fuzzy image enhancement methods. While some of these methods directly improve the image from grayscale images, the majority rely on image binarization.

In the field of image enhancement, the concept of fuzziness has been widely utilized in recent years by other researchers to enhance the contrast of the image. To obtain an image with a higher contrast, the fuzziness concept has been applied as a final stage of the image processing operation. The general structure of the fuzzy enhancement technique consists of three basic stages: fuzzification, performing some operations on membership values (membership modification), and defuzzification. An image of size M×N and L gray levels can be considered fuzzy by assigning a membership function to the image. This work uses an intensification operator to increase the image contrast and to reduce the fuzziness.

The idea of fuzziness has been widely applied by other researchers in the field of image enhancement in recent years to improve the contrast of the image. The fuzziness notion has been used as a last step in the image processing process to produce a picture with a higher contrast. The fuzzy enhancement technique is generally divided into three stages: defuzzification, fuzzification, and membership modification, which involves performing certain operations on membership values. A membership function of an image can be used to classify an image of size M×N and gray levels L as fuzzy. The intensification operator is used in this work to minimize fuzziness and boost contrast in the image.

Enhancing the resolution of the image was a top priority over the last few decades, and numerous methods have been created to address this issue. These algorithms aim to improve the low contrast and low resolution of the images, while producing more detailed high-resolution images.

High-resolution images can be produced using super-resolution techniques in conjunction with several observations or only one observation. Many super-resolution approaches, like single-frame super-resolution, have been developed by utilizing dictionary techniques and the DL concept. The goal of many studies nowadays is to convert LR images to HR images by employing fuzzy models for image sharpening and techniques based on the interpolation concept.

The two types of models mentioned above are considered among the most important and commonly used technologies in this research field. The process will start by applying the mathematical description model of histogram matching to the registered images; then the fuzzy logic methods for image enhancement and fusion process will be used to produce and obtain better HR image.

4. The Development Stages of Object Detection and Deep Learning

The development of object detection over the last 20 years is broadly categorized into two historical periods [1-12]. Recent breakthroughs in DL have substantially elevated detection performance, expanding its applicability. In this review, we examine recent techniques individually to highlight the interplay between various components and their collective impact on performance.

Neural network-based approaches and non-neural methods are the two main groups into which object detection techniques usually fall [65]. For non-neural approaches, it is necessary to perform classification using a technique such as support vector machines (SVMs), after first defining features using one of the popular methods. In contrast, neural techniques, which frequently rely on CNNs [5, 6, 11], are able to perform end-to-end object detection without requiring explicit, handcrafted feature extraction [53].

4.1 The first stage (Before 2014)

The first stage encompasses traditional, non-neural object detection methods developed prior to 2014. These approaches typically involve a multi-step pipeline reflecting the early evolution of object detection [5, 18], outlined as follows:

•Viola–Jones object detection technique, the first common detector based on features.

•Scale-invariant feature transform (SIFT), the most popular traditional detector, pioneering research work in object detection [15].

•Histogram of oriented gradients (HOG), a famous and the most often used feature descriptor for object detection in CV and image processing [17, 45].

4.2 The second stage (After 2014)

The second stage encompasses DL object detection methods based on neural network architectures (post-2014). As documented in previous studies [1-4], these approaches represent the subsequent evolution of DL object detection, detailed as follows:

  • Two-Stage Proposal Object Detection

Two-stage methods prioritize detection accuracy. Prominent algorithms in this category belong to the R-CNN family, which includes milestone architectures such as R-CNN [13], SPPNet, Fast R-CNN, Faster R-CNN [14], Mask R-CNN, R-FCN, and Cascade R-CNN.

  • One-Stage Proposal Object Detection

One-stage methods prioritize inference speed. Prominent algorithms in this category include the single shot multi-box detector (SSD) [18], RetinaNet, and the You Only Look Once (YOLO) family [54, 55], which encompasses models such as YOLO, YOLOv2, YOLOv3, YOLOv4, and YOLOv5.

4.2.1 Evolution of object detection within different methods for models and algorithms

In this study, we investigate advanced techniques and algorithms to achieve highly efficient and robust object detection across diverse image modalities.

Object detection and recognition have changed dramatically over time, moving from conventional CV techniques to sophisticated DL models.

The fields of CV and DL-based object detection have experienced substantial progress, with CNNs and Vision Transformers (ViTs) currently standing as the two dominant paradigms [53, 64].

Leading this progression, the YOLO family of models continues to redefine the boundaries of real-time object detection. YOLO reformulates object detection as a unified regression problem, directly predicting bounding boxes and class probabilities from entire images in a single evaluation. This groundbreaking approach enables YOLO architectures to achieve significantly higher inference speeds than traditional two-stage detectors while maintaining high accuracy. Through successive architectural enhancements, each YOLO iteration has continuously advanced performance across multiple benchmarks [69]. Consequently, these models remain highly relevant, though selection depends on specific deployment constraints, available data, and computational budget.

In our paper, we highlight the developments in CNN [5, 6, 11] and YOLO models [54, 55], along with a practical application using YOLO.

5. Low-Quality Images and Their Origin

The main energy source for our planet is the sun, and solar energy is transmitted as electromagnetic radiation.

At the speed of light, electromagnetic energy moves through empty space as waves of electric and magnetic fields with different frequencies or wavelengths [22].

5.1 Images composed of electromagnetic radiation

In our daily life, electromagnetic radiation is a common occurrence. Electromagnetic (EM) spectrum waves span a broad spectrum, including visible light detectable by the human eye, radio waves powering communication receivers, microwave radiation used for heating food, X-rays enabling medical imaging, and ultraviolet radiation emitted by high-temperature sources.

Electromagnetic waves are the waves created by the interaction of vibrating electric and magnetic fields. Figures 6 and 8 illustrate electromagnetic waves as oscillating electric and magnetic fields.

In general, a charged particle produces an electric field. This electric field pushes other charged particles. While negative charges go more quickly in the opposite direction from the field, positive charges travel more quickly in the field direction.

A moving charged particle generates a magnetic field. This magnetic field pushes other particles in motion. As the force acting on these charges is always perpendicular to their motion, it affects only the direction of the velocity and not the speed. This results in the electromagnetic field being produced by a charged particle that is accelerating. Electromagnetic waves are simply made up of electric and magnetic fields traveling at the speed of light (c) through empty space.

Figure 6. Electromagnetic (EM) spectrum and comparison of wavelength, frequency, and energy for the EM spectrum

A charged particle is said to be accelerating when it oscillates about an equilibrium position. An electromagnetic wave with frequency f is produced by a charged particle whose oscillation frequency is f. The equation λ = c/f can be used to determine the wavelength.

Emitting waves is one instance of an energy transfer that occurs in space [22]. The range of electromagnetic wave frequencies, wavelengths, and photon energies that span from 1 Hz to 1025 Hz, or wavelengths as small as a few hundred meters to as large as an atomic nucleus, is known as the Electromagnetic (EM) spectrum. The range of all forms of electromagnetic radiation is, thus, a fundamental definition of the EM. In a vacuum, the speed of all electromagnetic waves is equal to that of light. But for different kinds of electromagnetic waves, the wavelengths, frequencies, and photon energy will vary, as Figure 6 illustrates.

Every electromagnetic wave, including radio waves, microwaves, visible light, and X-rays, oscillates in parallel electric and magnetic fields [22, 25]. Only their wavelengths are different. X-rays have wavelengths in the range of 0.01 to 10 nm, but radio waves can have wavelengths of meters or even kilometers.

Visible light makes up a very small fraction of the overall EM spectrum. Ultraviolet light, X-rays, and gamma rays are examples of electromagnetic waves with shorter wavelengths, higher frequencies, greater and higher intensity or brightness, and higher energy. Microwaves, radio waves, and television waves are examples of electromagnetic waves having longer wavelengths, lower frequencies, lower intensity or brightness, and lesser energy. So, the fact here is that the intensity and brightness of light vary with the values of the wavelengths, frequencies, and energy.

Shorter wavelengths have higher energy, which is why radio waves cannot harm living things, while ultraviolet rays, whose shorter wavelengths cause sunburn, and X-rays can cause cancer and radiation sickness, due to longer wavelengths but lower energies.

All forms of electromagnetic radiation are grouped and organized into the EM spectrum according to their wavelength. The relationship between the radiation frequency and energy and wavelength is inverse.

5.2 Images composed with X-ray and infrared

Both X-ray and IR radiation belong to the EM spectrum; however, they occupy distinct spectral regions situated on opposite sides of the visible light band, as illustrated in Figure 6.

The EM spectrum represents the continuous distribution and categorization of electromagnetic waves according to their wavelengths and frequencies. As illustrated in Figure 6, this spectrum encompasses all forms of radiation within the universe. Among these, gamma rays exhibit the highest frequency band, whereas radio waves occupy the lowest, with the visible light spectrum situated near the center.

Consequently, these variations in wavelength and frequency directly govern how electromagnetic radiation interacts with matter and how it is harnessed for image synthesis [22]. Thus, the foundational parameters of interest are radiation frequency (or wavelength) and the energy content per photon.

Infrared (IR) radiation interacts with matter similarly to visible light, exhibiting comparable absorption and transmission characteristics across various media. Unlike lower-frequency radio waves, IR photons carry sufficient energy to induce vibrational transitions in molecular bonds, making IR radiation a highly efficient mechanism for thermal heating. Commonly emitted by objects at ambient and elevated temperatures, IR radiation is frequently termed 'radiant heat' [32, 33]. Although thermal energy itself corresponds to microscopic molecular kinetic motion rather than electromagnetic radiation, IR radiation both originates from thermal emissions and induces molecular thermal excitation, establishing a direct physical linkage between the two phenomena [32, 33].

X-rays operate at high frequencies and energy levels, allowing photons to penetrate lower-density structures [25]. In clinical imaging, this behavior facilitates radiographic evaluation, as tissues attenuate X-ray beams proportionally to their physical density and opacity, producing high-contrast diagnostic imagery.

Fundamentally, X-rays and IR radiation share similar wave properties, differing primarily in their wavelength-dependent attenuation and material absorption profiles. However, a critical physical distinction lies in their individual photon energies. Unlike IR photons, which lack sufficient energy to disrupt molecular structures, single X-ray photons possess high energies capable of breaking chemical bonds [22, 25–34]. Consequently, while IR radiation merely induces thermal vibrational excitation, X-rays act as ionizing radiation. Upon absorption, high-energy X-ray photons can induce molecular dissociation, potential cell damage, or DNA lesions [22, 25]. Due to the associated risk of radiation-induced oncogenesis, clinical applications strictly control X-ray dosage and employ protective radiation shielding, such as leaded barriers, to minimize exposure.

However, the added risk with the responsible use of medical X-ray imaging is tiny, and with today’s modern technology, it is often zero. X-rays are weakly absorbed by air, and modern dental X-ray equipment is so sensitive that the X-rays used can be so weak that they are completely absorbed by a few feet of air.

5.2.1 X-ray images

X-rays serve a pivotal role across numerous scientific and technical disciplines. Beyond their transformative contributions to medical science and diagnostics, X-radiation has played a foundational role in quantum mechanics, crystallography, and astrophysics. In industrial and security domains, X-ray scanning systems are routinely deployed to detect structural flaws in manufactured components and screen for contraband or hazardous materials at transit hubs [25-29, 36].

Here are options for rewriting the historical account of Roentgen's discovery. The revisions correct several historical and physics inaccuracies in the original draft—such as the nature of theCrookes tube, the physics of cathode rays vs. gas discharge, and the mechanism behind the phosphor screen illumination—while raising the prose to IEEE/Springer academic standards.

In 1895, the German physicist Wilhelm Röntgen discovered X-rays during experiments with cathode rays in gas discharge tubes. Röntgen observed that a nearby barium platinocyanide screen glowed fluorescently when the discharge tube was energized. Although cathode rays were known to induce fluorescence, the effect persisted even after the tube was completely enclosed in heavy black cardboard to block all visible and ultraviolet light. When various opaque objects were placed between the tube and the screen, the fluorescence remained largely unattenuated. To confirm that a novel form of penetrating radiation was responsible, Röntgen placed his hand between the tube and the screen, projecting a distinct image of the underlying skeletal structure. This landmark experiment simultaneously revealed the existence of X-rays and demonstrated their initial diagnostic application [22, 23, 25].

Röntgen’s discovery represents one of the most transformative medical milestones in history, pioneering non-invasive diagnostic radiography. X-ray technology enabled clinicians to evaluate bone fractures and identify foreign objects within the human body without surgical intervention. Subsequent advancements in X-ray diagnostic modalities further expanded these capabilities to include the detailed visualization and delineation of blood vessels and internal biological organs.

Fundamentally, X-rays and visible light are both constituents of the electromagnetic spectrum, though X-rays possess significantly higher photon energies. To elucidate the physical differences between these two radiation domains [22, 25], their properties are evaluated through photon energy, wavelength, and frequency, which are interrelated by the following fundamental expressions:

Photon energy = Planck constant × Frequency          E = hv

Frequency = Speed of light / Wavelength                  v = c/λ

X-rays have more energy than visible ray photons, which means that their frequencies are large and their wavelengths are short [22].

Therefore, X-rays are invisible rays for us, just like radio, IR, and ultraviolet rays, but the difference between all these rays is their properties in terms of photon energy, frequency, and wavelength.

5.2.2 Infrared images

IR vision is a critical technology in several military [29] and civilian applications as shown in Figure 7. These applications include environmental monitoring, biological diagnostics, and thermal probing of active microelectronic devices [30-34].

Target acquisition, tracking, surveillance, and night vision are just a few of the military uses of IR technology. Its non-military applications encompass weather forecasting, remote thermal monitoring, structural energy efficiency analysis, spectroscopy, and short-range optical wireless communication. Furthermore, IR astronomy leverages sensor-equipped telescopes to penetrate optically obscured regions of space—such as dense molecular clouds—to detect low-temperature celestial bodies like exoplanets, as well as highly redshifted emissions from the early cosmos.

Figure 7. Detection in X-ray images that describe work samples for object detection of different weapons in X-ray images using the YOLO detector model

The body temperature distribution can be visualized using IR imaging. This modality holds significant potential for the non-invasive diagnosis and prognosis of various pathological conditions. Specialized IR cameras capture thermal radiation emitted from the body surface, generating high-resolution thermograms that map the spatial distribution of skin temperature and underlying vascular perfusion.

Heatwaves are used to describe IR waves, which are created by heated molecules and bodies. Because thermal motion within matter excites molecular vibrational and rotational modes, emitted IR radiation lies immediately adjacent to the low-frequency, long-wavelength red boundary of the visible spectrum [22, 23].

The IR images have low contrast between the background and the targets or objects and a small signal-to-noise ratio (SNR). This intrinsic characteristic significantly diminishes target detectability and complicates feature extraction in thermal imaging scenarios. Consequently, advanced pre-processing algorithms and clutter-suppression techniques are frequently required to enhance object-to-background discriminability prior to high-level CV tasks. 

Because IR images have poor contrast, image processing is required to improve them. Target objects within thermal scenes often lack sufficient pixel intensity for reliable identification, while background regions exhibit non-zero luminance due to ambient thermal radiation [30-34]. Consequently, specialized IR image enhancement techniques are indispensable for optimizing dynamic range and contrast prior to high-level CV tasks, as evidenced in Figure 5.

Owing to its significance, a variety of pre-processing and enhancement methods for IR images have surfaced in recent years. These enhancement approaches are primarily responsible for improving object detection and recognition from IR images by expanding dynamic range, suppressing background clutter, and sharpening structural edge gradients. By mitigating thermal noise and increasing feature salient contrast, such techniques serve as crucial prerequisites for downstream feature extraction, classification networks, and automated target recognition (ATR) architectures.

Here are options for rewriting the paragraph, strictly preserving your requested opening sentence while adding technical depth and elevating the biometrics and CV terminology to IEEE/Springer publication standards.

Figure 8. Discovery & usage of electromagnetic radiation

Furthermore, image enhancement techniques can be used to improve the performance of advanced security methods that rely on IR gait recognition. By mitigating thermal blur, low spatial resolution, and environmental thermal noise, pre-processing algorithms produce well-defined silhouetted gait sequences. This enhanced temporal and spatial clarity allows feature extraction networks to accurately isolate biomechanical motion patterns and stride dynamics, thereby significantly increasing biometric identification accuracy in low-visibility or nighttime surveillance operations.

IR images have low contrast, low resolution, and inherent non-uniformity, making object recognition from them a more difficult process. It is necessary to first pre-process and treat IR images for effective object detection in order to solve this issue. To improve the quality of IR images, several research efforts have been introduced [30-34]. Effective object detection typically requires methods such as histogram equalization, histogram matching, and IR image super-resolution. Among these, histogram-based processing has been widely adopted across numerous studies to achieve visual quality enhancement in thermal imagery [17, 32]. Moreover, DL-based IR image super-resolution has emerged as a highly promising research avenue. Historically, as illustrated in Figure 8, IR radiation was recognized as constituting a major portion of the electromagnetic spectrum since its initial discovery.

6. Discussions of Models and Methods for the Experimental Image Detection Results

Figure 1 shows the experimental results of the detection and locating task for one or multiple objects in visual images by determining the location, boundaries, and categorization of those objects.

The images in Figure 2 describe the task steps of the traditional object detection method by determining the specific and strongest feature points for objects in visual images [64] using the EstimateGeometricTransform function.

The detection steps for a specified animal object in the image of a forest are as follows:

1. Reading the images of an isolated animal object and the image of the forest.

2. Detecting feature points, which are matched in both images (detect SURF features), for the image of the specified animal object and the image of the forest.

3. Finding presumptive matching points of features (including outliers) by locating the animal object in the image of the forest and locating the animal object (without outliers).

4. Detecting the animal object in the image of the forest using the EstimateGeometricTransform function, which gives the transformation relating the matched points and then gets the bounding polygon of the specified image, as can be seen from the images in Figure 2.

The images in Figures 3, 4, and 5 refer to the different descriptions of the shapes for the output resulting images for each sub-task, which are considered key steps to reach functioning effectively for the basic task of object detection. These steps include sub-tasks such as image classification, image recognition, image localization, image segmentation (semantic or instance segmentation), and image enhancement. Together, these output descriptions illustrate the complete operational framework required for successful object detection.

The images in Figure 5 describe the different results of the IR image enhancement, which is considered a preprocessing task before the object detection and recognition tasks to verify its effectiveness. However, images taken in the real world frequently, and especially IR images, include a variety of quality issues such as low light, low resolution, and color distortion, which greatly affect how well different detection algorithms work. It is a common practice to first employ certain enhancement techniques to recover a high-quality image from the original image, followed by object detection from the recovered image.

As shown in Figure 5, when light levels are low, the lighting (illumination) aspect of the image has a big influence on object detection quality. As a pre-processing step, low-light image enhancement techniques can increase image quality and yield better detection results. Improving contrast and brightness during this stage allows detection algorithms to identify target features more reliably. These results have been obtained using MATLAB.

The images in Figure 7 describe sample results in our work for weapons detection from X-ray images. We considered one or more different selected threat objects (gun, knife, and razor) for X-ray scanner baggage images using the YOLOv3 detector model in Python and X-ray imagery datasets (like GDXray, SIXray).

The research was organized and arranged based on the progression of information related to the study and application of image processing for accurate object detection from a number of visual high-quality and low-quality images. The background of electromagnetic radiation was essential for understanding all its types of rays, especially the position of both IR and X-ray within the system, leading to their beneficial and harmful uses, how to handle low-quality images, how to detect objects, and how to address and resolve some problems encountered.

6.1 Discussion of negative impacts of new models

6.1.1 Processing low-resolution or blurry images

We wanted to shed light on these types of images, which are difficult to handle directly in the object detection process. For example, regarding image enhancement, as we mentioned in our discussion, this process can either improve the image or, conversely, have a negative impact, exacerbating image problems and blurring. Therefore, the decision to use the technology or not depends on several factors and proportions within the specific image conditions.

A key component of image processing systems is contrast enhancement. Enhancement is used to improve an image and facilitate visual interpretation, comprehension, and analysis. Otherwise, we leave the original image, as shown in the images in Figure 9.

Figure 9. Description of the negative effect for some samples of Infrared (IR) image enhancement, which is considered a failed preprocessing task to enforce enhancement strategies

Contrast enhancement modifies the original image gray levels based on the light and dark edges of objects.

Gray level enhancement techniques manipulate pixel intensities in the spatial domain, especially in low-contrast images, to improve visual interpretation (making the pure white and dark areas more visible), which improves image quality, contrast, and feature visibility.

Each pixel in a grayscale image is given a specified range of intensity values, or brightness. The brightness of each pixel in a specified image is represented as an integer. For 8-bit images, this usually ranges from 0 (black) to 255 (white), and for various shades of gray, it ranges from 1–254.

The application of an algorithm for image enhancement has been created and put into use based on the following fuzzy rules:

• A dark pixel intensity results in a darker output.

• The result will be gray if the pixel intensity is gray.

• The output is brighter if the pixel intensity is high.

The use of these rules in an image enhancement editor improvement and a fuzzy inference system is demonstrated in Figure 5 and Figure 9.

6.1.2 X-ray weapon detection in port operations

Also, we presented an example of one of the latest technologies used in object detection in general, and specifically in the field of weapons. However, it relies on numerous and sophisticated processes that require significant time for implementation, including learning, training, testing, the use of massive datasets, and the various models of the YOLO family. These methods must be studied and developed, and efforts must be made to reduce implementation time to obtain accurate results.

Currently, the methods used in the detection of weapons are similar to traditional methods by determining the specific and strongest feature points for threat objects in visual X-ray scanner baggage images, as shown in the steps of Figure 2. Or by using the detection of X-ray scanner baggage images camera.

X-ray scanners for baggage are made to see through the bag and show the contents within, not just the outside, by capturing them with two- or three-sided X-ray scanner baggage cameras and their detectors. The purpose of these devices is to detect hazards like guns, explosives, illegal drugs, and contraband by utilizing the principle of differential X-ray absorption (the materials absorb X-rays in different ways).  

These X-ray scanners allow image enlargement, take a reverse image and a three-dimensional image of each object individually, set specific colors for items in bags after determining their mass and density, create color-coded 2D or 3D view images of the goods packed inside, and sometimes use CT scanners in modern airports.

Modern scanners use a system called dual-energy X-ray scanner technology, which sends two different energy levels through the baggage, and then the computer can estimate both the location and the type of the object material inside the bags.

7. Conclusions

Object detection is a difficult task. Completing a variety of object identification tasks remains challenging due to variations in object appearance, form, and posture, alongside interfering factors such as lighting and shading. Additionally, consideration must be given to system costs and the balance between cost and accuracy. Designing algorithms that account for these complex parameters makes object detection a particularly demanding machine vision task [7-11]. Consequently, object detection is widely recognized as one of the most challenging areas in machine vision [1-4].

This paper has introduced the outlines of:

* Two main categories of object detection methods that operate across different image and video formats: traditional methods [5-18] and DL-based methods [1-4], applied to original specified images or videos containing visual features as well as low-quality images [25-34].

* Some common and necessary subtasks that contribute to effective object detection performance [14-16], for example, image classification, image recognition, image localization, image segmentation, image enhancement, and others.

* A popular area of study in CV is object detection using CNNs [25-34]. Image lighting has a significant impact on object detection and can substantially reduce detection accuracy in low-light situations. However, image quality can be enhanced, and superior detection outcomes can be achieved by applying low-light image enhancement as a pre-processing step.

* By measuring the emission and absorption of light and other radiation as it interacts with materials, spectroscopy can be used for determining wavelength or frequency [22-24]. A light beam is scattered as it travels through material media. Due to interactions with substance atoms and molecules based on their resonance frequencies, these atoms and molecules respond to light waves of similar frequencies. A line spectrum is produced when light rays interact with an excited atom, releasing a range of distinct frequencies. The emission lines within this line spectrum are non-continuous and arranged systematically, producing light with specific wavelengths. Conversely, an absorption spectrum results when light with continuous wavelengths passes through a low-density medium, where specific lines are missing from the continuous spectrum due to atomic and molecular absorption at characteristic frequencies matching those of the light waves.

* There are different modalities for images based on the electromagnetic spectral band used for imaging [22-26]. Popular imaging technologies include X-ray imaging, magnetic resonance (MR) imaging, computed tomography (CT) imaging, and IR imaging [30-34]. Although these technologies differ in their spectral bands and mechanisms, a common thread unites them; all can be utilized for biomedical applications.

* The characteristics of electromagnetic radiation include its brightness, wavelength, frequency, and period. Light wave frequency and energy have been shown to be inversely related [22-24]. When energy quantization was discovered around the turn of the 20th century, it was recognized that light can be characterized both as a wave and as a collection of particles called photons. Photons carry discrete energy units known as quanta, where the absorption of photons transfers this energy to atoms and molecules, while photons released by atoms and molecules can lose energy.

  References

[1] Wang, W., Lai, Q., Fu, H., Shen, J., Ling, H., Yang, R. (2021). Salient object detection in the deep learning era: An in-depth survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6): 3239-3259. https://doi.org/10.1109/TPAMI.2021.3051099

[2] Zhao, Z.Q., Zheng, P., Xu, S.T., Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11): 3212-3232. https://doi.org/10.1109/TNNLS.2018.2876865

[3] Liu, L., Ouyang, W., Wang, X., et al. (2020). Deep learning for generic object detection: A survey. International Journal of Computer Vision, 128(2): 261-318. https://doi.org/10.1007/s11263-019-01247-4

[4] Zou, Z., Chen, K., Shi, Z., Guo, Y., Ye, J. (2023). Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3): 257-276. https://doi.org/10.1109/JPROC.2023.3238524

[5] Zhiqiang, W., Jun, L. (2017). A review of object detection based on convolutional neural network. In 2017 36th Chinese Control Conference (CCC), Dalian, China, pp. 11104-11109. https://doi.org/10.23919/ChiCC.2017.8029130

[6] Huang, J., Rathod, V., Sun, C., et al. (2017). Speed/accuracy trade-offs for modern convolutional object detectors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, pp. 7310-7311. https://doi.org/10.1109/CVPR.2017.351

[7] Li, K., Cao, L. (2020). A review of object detection techniques. In 2020 5th International Conference on Electromechanical Control Technology and Transportation (ICECTT), Nanchang, China, pp. 385-390. https://doi.org/10.1109/ICECTT50890.2020.00091

[8] Yadav, N., Binay, U. (2017). Comparative study of object detection algorithms. International Research Journal of Engineering and Technology (IRJET), 4(11): 586-591.

[9] Huang, G., Laradji, I., Vazquez, D., Lacoste-Julien, S., Rodriguez, P. (2022). A survey of self-supervised and few-shot object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4): 4071-4089. https://doi.org/10.1109/TPAMI.2022.3199617

[10] Gupta, A., Puri, R., Verma, M., Gunjyal, S., Kumar, A. (2019). Performance comparison of object detection algorithms with different feature extractors. In 2019 6th International Conference on Signal Processing and Integrated Networks (SPIN), Noida, India, pp. 472-477. https://doi.org/10.1109/SPIN.2019.8711763

[11] Agarwal, S., Terrail, J.O.D., Jurie, F. (2018). Recent advances in object detection in the age of deep convolutional neural networks. arXiv preprint arXiv:1809.03193. https://doi.org/10.48550/arXiv.1809.03193

[12] Borji, A., Cheng, M.M., Hou, Q., Jiang, H., Li, J. (2019). Salient object detection: A survey. Computational Visual Media, 5(2): 117-150. https://doi.org/10.1007/s41095-019-0149-9

[13] Dai, J., Li, Y., He, K., Sun, J. (2016). R-FCN: Object detection via region-based fully convolutional networks. Advances in Neural Information Processing Systems, 29.

[14] Girshick, R. (2015). Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, pp. 1440-1448. https://doi.org/10.1109/ICCV.2015.169

[15] Rublee, E., Rabaud, V., Konolige, K., Bradski, G. (2011). ORB: An efficient alternative to SIFT or SURF. In 2011 International Conference on Computer Vision, Barcelona, Spain, pp. 2564-2571. https://doi.org/10.1109/ICCV.2011.6126544

[16] Viola, P., Jones, M. (2001). Rapid object detection using a boosted cascade of simple features. In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR, Kauai, HI, USA, pp. I-I. https://doi.org/10.1109/CVPR.2001.990517

[17] Dalal, N., Triggs, B. (2005). Histograms of oriented gradients for human detection. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), 1: 886-893. https://doi.org/10.1109/CVPR.2005.177

[18] Liu, W., Anguelov, D., Erhan, D., et al. (2016). SSD: Single Shot MultiBox Detector. European Conference on Computer Vision, 9905: 21-37. https://doi.org/10.1007/978-3-319-46448-0_2

[19] Russakovsky, O., Deng, J., Su, H., et al. (2015). Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3): 211-252. https://doi.org/10.1007/s11263-015-0816-y

[20] Krizhevsky, A., Sutskever, I., Hinton, G.E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25.

[21] Girshick, R., Donahue, J., Darrell, T., Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, pp. 580-587. https://doi.org/10.1109/CVPR.2014.81

[22] Butcher, G. (2016). Tour of the Electromagnetic Spectrum. Government Printing Office.

[23] Jayakrishnan, V.M., Liya, M.L. (2020). A survey of electromagnetic waves based metamaterials and applications in various domains. In 2020 Third International Conference on Smart Systems and Inventive Technology (ICSSIT), Tirunelveli, India, pp. 662-666. https://doi.org/10.1109/ICSSIT48917.2020.9214163

[24] Sabah, C., Uckun, S. (2007). Electromagnetic wave propagation through frequency-dispersive and lossy double-negative slab. Opto-Electronics Review, 15(3): 133-143. https://doi.org/10.2478/s11772-007-0011-y

[25] Als-Nielsen, J., McMorrow, D. (2011). Elements of Modern X-Ray Physics. John Wiley & Sons.

[26] Franzel, T., Schmidt, U., Roth, S. (2012). Object detection in multi-view X-ray images. Joint DAGM (German Association for Pattern Recognition) and OAGM Symposium, 7476: 144-154. https://doi.org/10.1007/978-3-642-32717-9_15

[27] Chen, Z., Zheng, Y., Abidi, B.R., Page, D.L., Abidi, M.A. (2005). A combinational approach to the fusion, de-noising and enhancement of dual-energy x-ray luggage images. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05)-Workshops, San Diego, CA, USA, pp. 2-2. https://doi.org/10.1109/CVPR.2005.386

[28] Haff, R.P., Toyofuku, N. (2008). X-ray detection of defects and contaminants in the food industry. Sensing and Instrumentation for Food Quality and Safety, 2(4): 262-273. https://doi.org/10.1007/s11694-008-9059-8

[29] Mery, D., Mondragon, G., Riffo, V., Zuccar, I. (2013). Detection of regular objects in baggage using multiple X-ray views. Insight-Non-Destructive Testing and Condition Monitoring, 55(1): 16-20. https://doi.org/10.1784/insi.2012.55.1.16

[30] Maini, R., Aggarwal, H. (2010). A comprehensive review of image enhancement techniques. arXiv preprint arXiv:1003.4053. https://doi.org/10.48550/arXiv.1003.4053

[31] Gupta, A.K., Chauhan, S.S., Shrivastava, M. (2016). Low contrast image enhancement technique by using fuzzy method. International Journal of Engineering Research and General Science, 4(2): 518-526.

[32] Angaitkar, P.G., Saxena, P.K. (2012). Enhancement of infrared image: A review. Journal of Signal and Image Processing, 2(2): 1186-1189.

[33] Algarni, A.D. (2020). Efficient object detection and classification of heat emitting objects from infrared images based on deep learning. Multimedia Tools and Applications, 79(19): 13403-13426. https://doi.org/10.1007/s11042-020-08616-z

[34] Lu, Q., Conners, R. (2006). Using image processing methods to improve the explosive detection accuracy. IEEE Transactions on Systems, Man, and Cybernetics - Part C: Applications and Reviews, 36(6): 750-760. https://doi.org/10.1109/TSMCC.2005.855532

[35] Sharma, D., Jaffery, Z.A. (2023). Multiple object localization and tracking based on boosted efficient binary local image descriptor. MethodsX, 11: 102354. https://doi.org/10.1016/j.mex.2023.102354

[36] Mery, D., Pieringer, C. (2021). Computer Vision for X-Ray Testing (Vol. 805). Switzerland: Springer International Publishing. https://doi.org/10.1007/978-3-030-56769-9

[37] Aziz, L., Salam, M.S.B.H., Sheikh, U.U., Ayub, S. (2020). Exploring deep learning-based architecture, strategies, applications and current trends in generic object detection: A comprehensive review. IEEE Access, 8: 170461-170495. https://doi.org/10.1109/ACCESS.2020.3021508

[38] Khairdoost, N. (2022). Driver behavior analysis based on real on-road driving data in the design of advanced driving assistance systems. Doctoral Dissertation, The University of Western Ontario (Canada).

[39] Moustafa, S.S., Berbar, A., Ismail, A. (2004). Efficient memory performance for multi-issue processors. In Proceeding of International Conference on Electrical, Electronic and Computer Engineering, Cairo, Egypt, pp. 265-268.

[40] Li, T., Ma, X., Chen, Z., Wang, R., Xiong, X. (2000). Fuzzy feature matching between radar image and optical image. In Process Control and Inspection for Industry, 4222: 224-229. https://doi.org/10.1117/12.403879

[41] Şahin, Ö. (2023). Improving the performance of yolo-based detection algorithms for small object detection in UAV-taken images. Master's Thesis, Bilkent Universitesi (Turkey).

[42] Boesch, G. (2024). Object detection: The definitive guide. viso.ai. https://viso.ai/deep-learning/object-detection/.

[43] Sadic, N., Hassan, E., El-Rabaie, S., El-dolil, S., Dessoky, M.I., El-Samie, F. (2021). Enhancement technique of infrared images. Menoufia Journal of Electronic Engineering Research, 30(1): 58-64.

[44] Chen, J., Qin, D., Hou, D., Zhang, J., Deng, M., Sun, G. (2022). Multiscale object contrastive learning-derived few-shot object detection in VHR imagery. IEEE Transactions on Geoscience and Remote Sensing, 60: 1-15. https://doi.org/10.1109/TGRS.2022.3229041

[45] Joseph, S., Olugbara, O.O. (2021). Detecting salient image objects using color histogram clustering for region granularity. Journal of Imaging, 7(9): 187. https://doi.org/10.3390/jimaging7090187

[46] Li, Y., Hou, C., Tian, F., et al. (2007). Enhancement of infrared image based on the retinex theory. In 2007 29th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, Lyon, France, pp. 3315-3318. https://doi.org/10.1109/IEMBS.2007.4353039

[47] Amjoud, A.B., Amrouch, M. (2023). Object detection using deep learning, CNNs and vision transformers: A review. IEEE Access, 11: 35479-35516. https://doi.org/10.1109/ACCESS.2023.3266093

[48] Holland, A.D., Minoglou, K. (2024). X-ray, optical, and infrared detectors for astronomy XI. In: Proceedings of SPIE, 13103: 1310301-1.

[49] Jana, M., Roy, P. (2023). Low light scene enhancement for object detection and tracking using multiscale decomposition and fusion. In 2023 3rd International Conference on Range Technology (ICORT), Chandipur, Balasore, India, pp. 1-6. https://doi.org/10.1109/ICORT56052.2023.10249164

[50] Sharmila, V., Kannadhasan, S., Kannan, A.R., Sivakumar, P., Vennila, V. (Eds.). (2024). Challenges in information, communication and computing technology. Proceedings of the 2nd International Conference on Challenges in Information, Communication, and Computing Technology (ICCICCT 2024), Namakkal, Tamil Nadu, India. 

[51] Chung, Y.L., Yang, J.J. (2021). Application of a mask R-CNN-based deep learning model combined with the retinex image enhancement algorithm for detecting rockfall and potholes on hill roads. In 2021 IEEE 11th International Conference on Consumer Electronics, ICCE-Berlin, Berlin, Germany, pp. 1-6. https://doi.org/10.1109/ICCE-Berlin53567.2021.9720001

[52] Paramasivam, R., Kumar, P., Lai, W.C., Bidare Divakarachari, P. (2024). Deep ensemble model-based moving object detection and classification using SAR images. Frontiers in Earth Science, 11: 1288003. https://doi.org/10.3389/feart.2023.1288003

[53] Khan, A., Fouda, M.M., Do, D.T., Almaleh, A., Alqahtani, A.M., Rahman, A.U. (2024). Underwater target detection using deep learning: methodologies, challenges, applications, and future evolution. IEEE Access, 12: 12618-12635. https://doi.org/10.1109/ACCESS.2024.3353688

[54] Docto, J.P., Labininay, A.I., Villaverde, J.F. (2022). Third eye hand glove object detection for visually impaired using You Only Look Once (YOLO) v4-tiny algorithm. In 2022 IEEE International Conference on Artificial Intelligence in Engineering and Technology (IICAIET), Kota Kinabalu, Malaysia, pp. 1-6. https://doi.org/10.1109/IICAIET55139.2022.9936740

[55] Kodipaka, V., Marques, L., Cortesão, R., Araújo, H. (2023). APH-YOLOv7t: A YOLO attention prediction head for search and rescue with drones. In Iberian Robotics Conference, 978: 256-268. https://doi.org/10.1007/978-3-031-59167-9_22

[56] Yu, H., Liu, J., Liu, L., Ju, Z., Liu, Y., Zhou, D. (Eds.) (2019). Intelligent Robotics and Applications. Springer Science and Business Media LLC. https://link.springer.com/book/10.1007/978-3-030-27526-6.

[57] Sharma, L. (Ed.). (2020). Towards Smart World: Homes to Cities Using Internet of Things. CRC Press.

[58] Zhao, R., Tang, S., Supeni, E.E.B., Rahim, S.B.A., Fan, L. (2024). A review of object detection in traffic scenes based on deep learning. Applied Mathematics and Nonlinear Sciences, 9(1): 0322. https://doi.org/10.2478/amns-2024-0322

[59] Lin, Z., Wang, L., Yang, J., et al. (Eds.) (2019). Pattern Recognition and Computer Vision. Springer Science and Business Media LLC. https://link.springer.com/book/10.1007/978-3-030-31654-9.

[60] Tang, W., Chen, J., Ning, Y., Xu, K. (2023). Object detection algorithms with deep learning: An overview. In 2023 3rd International Conference on Electronic Information Engineering and Computer Science (EIECS), Changchun, China, pp. 844-858. https://doi.org/10.1109/EIECS59936.2023.10435579

[61] CSA & CUTE. (2015). Advances in Computer Science and Ubiquitous Computing. Springer Nature. https://link.springer.com/book/10.1007/978-981-10-0281-6.

[62] Arora, A., Grover, A., Chugh, R., Reka, S.S. (2019). Real time multi object detection for blind using Single Shot MultiBox Detector. Wireless Personal Communications, 107(1): 651. https://doi.org/10.1007/s11277-019-06294-1

[63] Choudhary, P., Satpathy, S., Dagur, A., Shukla, D.K. (Eds.). (2025). Recent trends in intelligent computing and communication: Volume 1. CRC Press.

[64] Chen, E.H., Röthig, P., Zeisler, J., Burschka, D. (2019). Investigating low level features in CNN for traffic sign detection and recognition. In 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, New Zealand, pp. 325-332. https://doi.org/10.1109/ITSC.2019.8917340

[65] Fayaz, F.A., Nasti, S.M., Khan, U.I., Dar, A.H. (2025). Edge-driven real-time UAV thermal imaging for survivor identification. International Conference on Innovative Computing and Communication, 1432: 521-534. https://doi.org/10.1007/978-981-96-6715-4_36

[66] Senjyu, T., So-In, C., Joshi, A. (2024). Smart trends in computing and communications. Proceedings of SmartCom, 650: 2.

[67] Lan, Q. (2025). Knowledge distillation for efficient object detection: Toward scalable and deployable vision models. Doctoral Dissertation, The University of Alabama at Birmingham.

[68] Al-Sharafi, M.A., Al-Emran, M., Mahmoud, M.A., Arpaci, I. (Eds.). (2025). Current and Future Trends on AI Applications. Springer.

[69] Deepak, G.D., Bhat, S.K. (2025). Maximizing YOLOv2 efficiency: A study on multiclass detection of indoor objects. Results in Engineering, 26: 105405. https://doi.org/10.1016/j.rineng.2025.105405

[70] Arrahmah, A.I., Rahmania, R., Saputra, D.E. (2022). Comparison between convolutional neural network and K-nearest neighbours object detection for autonomous drone. Bulletin of Electrical Engineering and Informatics, 11(4): 2303-2312. https://doi.org/10.11591/eei.v11i4.3784

[71] Wang, T., Wei, R., Wang, L., et al. (2021). Detection of transmission towers and insulators in remote sensing images with deep learning. In 2021 China Automation Congress (CAC), pp. 3298-3303. https://doi.org/10.1109/CAC53003.2021.9728166

[72] Jiang, J. (2025). Multimedia forensics: Identification and verification of source camera, vehicle speed estimation, and deepfakes detection. Doctoral Dissertation, Old Dominion University.