© 2026 The authors. This article is published by IIETA and is licensed under the CC BY 4.0 license (http://creativecommons.org/licenses/by/4.0/).
OPEN ACCESS
The abnormalities in the human heart can be screened by the Electrocardiogram (ECG) signals. In this article, the abnormalities in ECG signals, such as Atrial Fibrillation (AF), are detected from Non-Atrial Fibrillation (NAF) signals using the proposed transformer-based deep learning algorithm. The main contribution of the proposed ECG signal classification system is a preprocessing module that uses a Butterworth filter (BWF) for detecting and removing aliasing and noise in the acquired ECG signals in both training and testing stages. Then, the low and high frequency sub-band (HFS) are obtained using Non Subsampled Contourlet Transform (NSCT) module through pyramidal and directional filters. The statistical feature set is computed from the NSCT decomposed low and HFS. Finally, the derived statistical features are classified using the proposed Transformer Driven Convolutional Neural Networks (TDCNet) classification architecture. This proposed TDCNet based ECG signal classification system has been evaluated on ECG signal datasets Massachusetts Institute of Technology – Beth Israel Hospital (MIT-BIH) and China Physiological Signal Challenge (CPSC). The experimental results of this proposed ECG signal classification system have been compared with recent methods in terms of functional metrics: Signal detection Sensitivity (SSe), Signal detection Specificity (SSp), Signal detection Accuracy (Sac), Signal detection Jaccard Index (SJI), and Signal detection Precision (SPr). The proposed ECG signal classification system obtains 98.8% SSe, 98.5% SSp, 99.3% Sac, 99.1% SJI and 98.9% SPr on the MIT-BIH dataset. The proposed ECG signal classification system obtains 99.3% SSe, 99.1% SSp, 98.9% Sac, 99.3% SJI and 99.1% SPr on the CPSC dataset. The experimental results obtained in this research article are validated by the k-fold cross validation algorithm.
Electrocardiogram, noises, features, sub-bands, classifications
Nowadays, cardiovascular diseases are identified as the life killing diseases that can affect all types of age groups irrespective of their gender types. The cardiovascular related abnormalities in the human body can be screened by detecting the electrical activity of the human heart [1, 2]. This electrical activity can be measured using the Electrocardiogram (ECG). This ECG signal can be captured by placing the 12-lead ECG electrode system on the human body at various skin parts. These electrodes are placed over the body on conductive gel. The dry electrodes, which are connected to the ECG equipment, were placed on various skin parts of the human body, and they captured the ECG waveforms [3-5]. The ECG waves are categorized into three different types as P-waveform, T-wave, and QRS waveform. Among these ECG waveforms, the atrial contraction can be measured or referred to by the P-wave, and it indicates the size and activity of the atrium region of the heart. The activity of ventricular repolarization has been measured by the T-wave, and it is smaller than the length of the QRS waveform. The ventricular contraction can be measured by the QRS waveform [6-8].
The irregular and fast movement of pulses from the heart leads to the formation of a heart abnormality, which is called Atrial Fibrillation (AF). The upper chamber of the human heart generates fast pulses due to this AF, and the heart beats are high, with symptoms of shortness of breath, dizziness, and fatigue. Other correlated symptoms are high blood pressure and stroke [9]. The heart may fail due to this abnormal AF. The timely detection of this AF abnormality can save a human’s life. The detection of this AF can be performed by a manual process through an expert cardiologist [10]. In large population countries, it is difficult to screen for AF for all heart patients on time due to the shortage of expert cardiologists. Hence, the main motivation of this research work is to develop a fully automated system to detect AF abnormality using artificial intelligence algorithms. In this article, the deep learning algorithm extracts the fundamental concept from the transformer classification model to detect the ECG abnormalities.
The research article has the following contributions.
Figure 1(a) shows the Non-Atrial Fibrillation Electrocardiogram (NAF-ECG) waveform and Figure 1(b) shows the Abnormal AF-ECG waveform.
Figure 1. (a) Non-Atrial Fibrillation Electrocardiogram (NAF-ECG) waveform, (b) Abnormal Atrial Fibrillation Electrocardiogram (AF-ECG) waveform
There are numerous Convolutional Neural Network (CNN) model based AF detection methods and used by many researchers in the past two decades. These conventional CNN models mainly focused on computing the spatial features from the source ECG signal and failed to extract the long-range temporal features, which are important for the AF signal detection process. Moreover, the conventional CNN models required a larger number of trainable parameters, which increased the computational complexity of the AF detection system and also led to an increase in the signal training time period. Hence, these models are not suitable for real time healthcare monitoring system. To overcome these limitations, this research work proposes a novel TDCNet model for the detection and classification of AF signals. The proposed TDCNet integrates the functionality of the proposed parallel formed CNN and Transformer model within a unified framework. Specifically, the Transformer model splits the entire frequency feature matrix into parallel features, and then the proposed parallel formed CNN module efficiently extracts discriminative local morphological features with a global discriminative feature set from ECG signals for improving the AF Signal Detection Accuracy (Sac). This hybrid design significantly enhances the model's capability to recognize subtle AF patterns while maintaining computational efficiency.
The comprehensive concise comparison table depicted in Table 1 shows the conventional classification methods, their limitations, and also states that how the proposed TDCNet model addresses these limitations for the ECG signal classification process.
Table 1. Comprehensive concise comparison table
|
Conventional Methods |
Limitations |
How the Proposed TDCNet Addresses the Limitations |
|
Conventional CNN |
Extracts only local features from the signals |
It efficiently integrates both global and local signal features |
|
LSTM model |
Consumes higher training time period |
The parallel formed model reduces the training time period |
|
Vision Transformer Model |
Higher computational complexity |
The computational complexity can be reduced by extracting global and local features from the signals |
|
Machine learning algorithm |
Functions only on hand crafted feature maps |
The signal classification accuracy has been improved using extracted discriminative features |
This research article is divided into various sections. Section 2 represents the conventional AF and NAF detection approaches using automated algorithms, and Section 3 proposes a novel TDCNet classification algorithm for the effective classification of the AF and NAF signals. The experimental details along with the simulation results are given in Section 4. The limitations and the future direction of this research work are given in Section 5 as a conclusion.
In this section, the various conventional ECG signal analysis methods are presented with respect to pre-processing approach, feature computation approach, and classification models.
2.1 Pre-processing approaches
Azmy [11] used a discrete transformation process for the decomposition task, and the fractional Fourier transform was used to eliminate the aliasing error during the transformation process of the ECG signals. The hyperbolic functions of this method improved the ECG signal classification rate. Then, the ECG signal classifications were performed using the multifractal detrended fluctuation analysis, and the experimental analyses were compared with the DL and SVM techniques. This work obtained 97.1% Signal detection Sensitivity (SSe), 97.3% Signal detection Specificity (SSp), 97.9% Sac, 97.6% Signal detection Jaccard Index (SJI) and 97.3% Signal detection Precision (SPr) on MIT-BIH dataset and this work obtained 98.2% SSe, 98.5% SSp, 98.1% Sac, 98.4% SJI and 97.3% SPr on China Physiological Signal Challenge (CPSC) dataset. Lim et al. [12] proposed a QRS centric methodology for AF detection and classification based on a deep learning algorithm. The heart rate variations were identified in the ECG signal wave format and segmented using an adaptive algorithm. The robust pattern features were computed from the segmented heart rate variation format, and they were further classified by the robust machine learning algorithm. This work obtained 96.8% SSe, 96.9% SSp, 97.1% Sac, 97.2% SJI, and 97.1% SPr on the MIT-BIH dataset, and this work obtained 97.9% SSe, 98.2% SSp, 97.6% Sac, 97.2% SJI, and 97.1% SPr on the CPSC dataset. Xie et al. [13] developed an automated AF and NAF classification system using artificial intelligence approaches. The ECG signals, which were obtained through the single lead ECG electrodes, were quantized using 16-bit quantization format. The quantized ECG signal data was analyzed and classified using the ResNet deep learning algorithm. The internal variations of the feature metrics were used in this work to differentiate the AF signals from the NAF signals. This work obtained 96.4% SSe, 96.6% SSp, 96.5% Sac, 96.6% SJI, and 96.7% SPr on the MIT-BIH dataset, and this work obtained 97.3% SSe, 97.8% SSp, 97.3% Sac, 96.7% SJI, and 96.8% SPr on the CPSC dataset.
2.2 Feature computation and classification approaches
Wang et al. [14] differentiated the AF signal from the NAF signals using the signal feature fusion approach. This method used a spatial-frequency feature fusion algorithm that fused the derived pattern features from the obtained feature matrix. Finally, the fused features from this module were classified using a CNN algorithm, which provided the AF classification results. This work obtained 96.1% SSe, 96.2% SSp, 96.4% Sac, 96.3% SJI, and 96.3% SPr on the MIT-BIH dataset, and this work obtained 97.1% SSe, 97.3% SSp, 96.2% Sac, 96.3% SJI, and 96.2% SPr on the CPSC dataset. Gupta et al. [15] used the Local Mean Decomposition (LMD) algorithm to decompose the entire ECG signal into a number of local subbands. Then the correlation metric features were computed and classified using the ensemble classification algorithm. The hyperparameter optimization of this ensemble classification algorithm improved the final ECG signal classification rate. This work obtained 95.8% SSe, 95.3% SSp, 95.7% Sac, 95.8% SJI, and 95.7% SPr on the MIT-BIH dataset, and this work obtained 96.8% SSe, 96.2% SSp, 95.8% Sac, 96.1% SJI, and 95.7% SPr on the CPSC dataset. Pandey et al. [16] developed an AF and NAF identification algorithm from the ECG signals, and the pattern features were computed from the decomposed sub-band coefficients. The hybrid algorithm, which was developed in this work, classified the ECG signal into either AF or NAF using the computed decomposed coefficients. This work obtained 95.3% SSe, 95.1% SSp, 95.2% Sac, 95.1% SJI, and 95.3% SPr on the MIT-BIH dataset, and this work obtained 96.3% SSe, 96.1% SSp, 95.3% Sac, 95.4% SJI, and 95.3% SPr on the CPSC dataset.
In this article, the abnormalities in ECG signals are detected using the proposed transformer diffused deep learning algorithm. This proposed ECG signal classification system contains a preprocessing module that uses BWF for detecting and removing aliasing and noise in the acquired ECG signals in both training and testing stages. Then, the low and HFSs are obtained using the NSCT module through pyramidal and directional filters. The statistical feature set is computed from the NSCT decomposed low and HFSs. Finally, the derived statistical features are classified using the proposed TDCNet classification architecture. Figure 2(a) shows the AF and NAF ECG signals trained by the proposed TDCNet classification algorithm, and Figure 2(b) shows the AF and NAF ECG signals tested and validated by the proposed TDCNet and K-fold algorithms.
Figure 2. (a) Atrial Fibrillation (AF) and Non-Atrial Fibrillation Electrocardiogram (NAF-ECG) signals training by the proposed Transformer Driven Convolutional Neural Networks (TDCNet) classification algorithm, (b) AF and NAF ECG signals testing and validating by the proposed TDCNet and K-fold algorithms
3.1 Filtration of electrocardiogram signals
In this research article, BWF has been used to remove the noise from the source ECG signals. The BWF is otherwise called a flat magnitude filter, which has a flat frequency response for filtering the ECG signals. It filters the ECG signal without any ripples and signal distortion during the filtering process. The signal integrity has been improved by passing the ECG signal through this BWF module in the proposed ECG signal classification system. It preserves the low frequency components of the ECG signals and removes the high frequency noise component from the signal. The transition from passband to stopband of BWF is entirely based on the order of the BWF, which is noted as 'n' and the cut-off frequency '$w_c$'.
The frequency response of the BWF for the noise removal process in the ECG signal classification system is given in Eq. (1) below.
$H(w)=\frac{1}{\sqrt{1+\left(\frac{w}{w_c}\right)^{2 n}}}$ (1)
Figure 3. (a) Source Electrocardiogram (ECG) signal with noise component to be tested for Non-Atrial Fibrillation (NAF) before the Butterworth filter (BWF) is applied as a preprocessing method, (b) Non-Atrial Fibrillation Electrocardiogram (NAF-ECG) signal after applying the BWF as a preprocessing method
Figure 4. (a) Source Electrocardiogram (ECG) signal with noise component to be tested for Atrial Fibrillation (AF) before the Butterworth filter (BWF) is applied as a preprocessing method, (b) AF-ECG signal after applying the BWF as a preprocessing method
Figure 3(a) shows the source ECG signal with noise component to be tested for NAF, and Figure 3(b) shows the BWF filtered NAF-ECG signal. Figure 4(a) shows the source ECG signal with a noise component to be tested for AF, and Figure 4(b) shows the BWF filtered AF-ECG signal. In the case of NAF signals, the density of signal amplitudes and frequency variations with the noise components are high, which can be clearly viewed in Figure 3(a). The presence of the noise components in this NAF signal has been removed by BWF, and the output of the BWF module is clearly shown in Figure 3(b). From Figure 3(b), the noise components between the signal samples are significantly removed, which creates a higher impact in the signal classification process.
In the case of AF signals, the density of signal amplitudes and frequency variations with the noise components is low, which can be clearly viewed in Figure 4(a). The presence of the noise components in this AF signal has been removed by BWF, and the output of the BWF module is clearly shown in Figure 4(b). From Figure 4(b), the noise components between the signal samples are significantly removed, which creates a higher impact in the signal classification process.
3.2 Non subsampled contourlet transform of electrocardiogram signals
In order to compute and derive the contextual information from the ECG signals, the decomposition is an important module that decomposes or downsamples the source filtered ECG signals into a number of downsampled subbands. The obtained downsampled subbands represent the detailed contextual information that is important for computing the features for the differentiation of AF and NAF signals. In this research article, the NSCT transform has been used for performing the decomposition process on the ECG signals [17]. The filtered 1D ECG signal is converted into a 2D ECG signal and then applied to the NSCT transform for the decomposition process. The NSCT transformation module contains two Low Frequency Subband (LFS) filters and one HFS filter along with the Non Subsampled Pyramidal Filter (NSPF) and Non Subsampled Directional Filter (NSDF), as depicted in Figure 5. The filtered ECG signals through the BWF are passed through the HFS and LFS filters simultaneously, which separates the low and high frequency components from the filtered ECG signals. The HFS filter output is directly applied to NSDF to obtain the HFS. The NSDF separates the HFS components without aliasing and mostly derives the edge variations and textures in the signals. It performs shift invariance on the signal to provide a multi-directional decomposition process with respect to various scales and directions. NSPF performs the shift invariance based decomposition process that captures the contextual information from the signal. The obtained LFSs (LFS 1 and LFS 2) and HFS are combined into a matrix, which is called the Unique Decomposition Matrix (UDM) with respect to M rows and N columns.
Figure 5. Electrocardiogram (ECG) signal decomposition process using Non Subsampled Contourlet Transform (NSCT)
Figure 6(a) shows the decomposed HFS of the NAF signal, Figure 6(b) shows the decomposed LFS 1 of the NAF signal, and Figure 6(c) shows the decomposed LFS 2 of the NAF signal.
The NSCT decomposes the ECG signal into low frequency subband and high frequency subband as shown in Figure 6 and Figure 7. The low-frequency components carry out the global characteristic features, while the high-frequency components carry out the fine variation features of morphology and discrimination. Both subbands are used to determine the global discriminative features, which are further trained by the proposed TDCNet model to enhance the AF and NAF signal classification performance accuracy.
Figure 6. Decomposition results of Non-Atrial Fibrillation Electrocardiogram (NAF-ECG) signal: (a) decomposed high frequency sub-band, (b) decomposed low frequency sub-band 1, (c) decomposed low frequency sub-band 2
Figure 7. Decomposition results of Atrial Fibrillation Electrocardiogram (AF-ECG) signal: (a) decomposed high frequency sub-band, (b) decomposed low frequency sub-band 1, (c) decomposed low frequency sub-band 2
Figure 7(a) shows the decomposed HFS of the AF signal, Figure 7(b) shows the decomposed LFS 1 of the AF signal, and Figure 7(c) shows the decomposed LFS 2 of the AF signal.
The novelty of this research work is to effectively integrate these two modules, BWF and NSCT, with the proposed TDCNet for obtaining both global and local discriminative features for the effective classification of AF signals. The BWF is particularly used in this proposed system to suppress the high frequency artifacts. In this proposed system, the high frequency artifacts have more of an impact on the ECG signals. Moreover, the ECG signal components P wave, QRS, and the time periods of the RR interval are much preserved by the BWF, which is applied to the direct raw ECG signals. These signal components are more important for AF/NAF signal classification process. After noise suppression, the shift-invariant property of NSCT preserves the subtle morphological changes seen in AF without inducing any artifact of translation.
3.3 Feature determination from Unique Decomposition Matrix
The following statistical features (Eqs. (2)–(5)) are computed from the decomposed subband UDM from the NSCT module.
Signal Detection Skewness $(S D S)=\frac{1}{M * N} \frac{\sum_{i=1}^M\left(U D M_i-\mu\right)^3}{S D^2}$ (2)
Signal Detection Kurtosis $(S D K)=\frac{1}{M * N} \frac{\sum_{i=1}^M\left(U D M_i-\mu\right)^4}{S D^2}$ (3)
Signal Detection Correlation $(S D C)=\frac{\sum\left({U D M_i}^{-\mu}\right) *\left(U D M_j-\mu\right)}{\sqrt{\left(U D M_i-\mu\right)^2 *\left(U D M_j-\mu\right)^2}}$ (4)
Signal Detection Covariance $(S D C o)=\frac{\sum\left(U D M_i-\mu\right) *\left(U D M_j-\mu\right)}{M-1}$ (5)
where, UDM is the sub-band matrix from the NSCT module with respect to row M and Column N, and the mean and standard deviations are denoted as $\mu$ and SD, respectively.
The irregularities in ECG signals are determined using the derived statistical features. The SDS of the NAF signal is symmetric, distributions which are having the value close to zero, and the SDS of the AF signal is asymmetric due to the absence of signal components. The SDK of the NAF signal is typically a low value due to its lower amplitude of the signal, and the SDK of the AF signal is typically a high value due to its higher amplitude of the signal. The SDC of the NAF signal is typically a low value due to its regular interval, and the SDC of the AF signal is typically a high value due to its irregular interval sequences. The SDCo of the NAF signal typically has a lower value due to signal component regularity, and the SDCo of the AF signal typically has a higher value due to signal component irregularity. These values are helpful to differentiate the AF and NAF signals by the proposed TDCNet model.
All the derived statistical features from this feature computation module are grouped into a Feature Matrix (FM), and they are given to the proposed classification module for differentiating the AF and NAF ECG signals.
3.4 Proposed transformer driven convolutional neural networks classification process
The ECG signal classification process is an important module of the entire proposed ECG signal classification system. The classifiers used in the proposed system classify the input ECG signal into either AF or NAF based on the learning capacity. Many researchers from the past two decades have used deep learning CNN architectures for the ECG signal classification process. Though these methodologies stated in these works attained sustainable experimental results for the AF/NAF signal classifications, their computational time is not optimal due to their complex internal layering design, which is identified as the main limitation of these traditional ECG signal classification methods. In this research work, such limitation has been resolved by proposing a novel transformer model based CNN architecture for the classification of the ECG signals into AF/NAF cases. In the case of a traditional CNN architecture, the computed spatial feature values are fed into the CNN in sequence order, which consumes more computational time. This limitation has been mitigated by proposing a Transformer model, and this model reduces the significant time period of the ECG signal classification system. Even though it has certain advantages in terms of computational time, its internal architectural model is complex due to its encoding and decoding procedure and also exhibits lower experimental performance metrics. Hence, the proposed TDCNet Architecture eliminates such limitations by feeding the feature values into the CNN in parallel form. This proposed architecture performs a significant reduction in computational time as well as an improvement in experimental performance metrics.
Figure 8 is the proposed TDCNet Architecture for the classification of the AF/NAF ECG signals. The proposed TDCNet Architecture for the classification of AF/NAF ECG signals splits the derived Feature Matrix with m*n size into a number of s*r size FM. All the FM are having same size with respect to rows and columns. Zero padding is performed if the last FM contains null feature values, and hence the size of each module output is obtained from each internal layer of this proposed architecture. In this article, the entire Feature Matrix is split into four FMs, as specified as FM1, FM2, FM3, and FMn. All these four FMs have been processed at the same time, and hence the AF/NAF detection time is significantly reduced by the proposed architecture. The values in each FM are applied to the Consistent Adder Module (CAM), which contains arithmetic adder units. FM1 is arithmetically added to FM2 and fed into the first parallel CNN layers, FM2 is arithmetically added to FM3 and fed into the second parallel CNN layers, and FM3 is arithmetically added to FMn and fed into the third parallel CNN layers of the TDCNet architecture.
Figure 8. Proposed Transformer Driven CNN (TDCNet) architecture for the classification of the f Atrial Fibrillation/Non-Atrial Fibrillation Electrocardiogram (AF/NAF ECG) signals
The first parallel layer contains two Convolutional layers, two Rectified Linear Unit (ReLU) layers, and a spatial pooling layer as illustrated in the following Eqs. (7)–(9).
Parallel Layer $-1=\{$ConL11, ReLU, ConL12, ReLU, P1$\}$ (6)
Convolutional layers $=\{$ConL11, ConL12$\}$ (7)
ConL11 $=32$ filtering layers with $3 * 3$ stride function (8)
ConL12 $=512$ filtering layers with $5 * 5$ stride function (9)
It is used to extract the non-linear features from the input feature sequences. The variations between each feature value and its surrounding feature values are computed using this convolutional layer. It uses the convolution process by producing the feature map through convolution with the learnable kernel. The feature map from each convolutional layer contains complex feature patterns. It also preserves the spatial hierarchy between the feature values in the generated feature map. The size of the generated feature is depending on the kernel size, such as 3*3 and 5*5, and its stride value.
The spatial dimension of the output sequences of each convolutional layer is high, and hence it is not able to be processed further by the next layer in the proposed architecture. Hence, the spatial dimensionality has been reduced using a spatial pooling function, which is applied on the final convolutional layer output in each parallel layer of the proposed architecture. It splits the output convolved sequences into equal 4*4 regions and selects the maximum feature value in the 4*4 region. The most prominent feature value is chosen by this spatial pooling layer with reduced spatial dimensionality.
In this research work, the ReLU has been used as the Activation function in the proposed TDCNet Architecture. This ReLU has been placed at the output of reach Convolutional layer to eliminate negative values in the response of the convolved sequences in each layer. It is used to prevent gradient shrinking during the training of the deep learning algorithm. Due to this activation function, the training of the proposed system becomes faster. Due to its simple comparison process for vanishing gradient solution, it attains more computational efficiency. By avoiding negative (non-zero) values at each neuron in the internal layer, the sparsing property of the proposed TDCNet architecture is improved due to this activation function. The following Eq. (10) represents the ReLU activation function.
$f\left(y_i\right)=\max \left(0, y_i\right)$ (10)
where, $y_i$ is the response of the convolved sequences in each convolution layer of the proposed architecture.
It is more important to learn the input parameter than the number of layers designed in the proposed TDCNet Architecture. This is achieved through the non-linearity provided by the activation function.
The second parallel layer contains two Convolutional layers, two ReLU layers, and a spatial pooling layer as illustrated in the following equations.
Parallel Layer $-2=\{$ConL21, ReLU, ConL22, ReLU,P2$\}$ (11)
Convolutional layers $=\{$ConL21, ConL22$\}$ (12)
ConL11 = 64 filtering layers with 3 * 3 stride function (13)
ConL12 $=512$ filtering layers with $5 * 5$ stride function (14)
The third parallel layer contains two Convolutional layers, two ReLU layers, and a spatial pooling layer, as illustrated in the following Eqs. (15)–(18).
Parallel Layer - 3 = {ConL31, ReLU, ConL32, ReLU, P3} (15)
Convolutional layers $=\{$ConL31, ConL32$\}$ (16)
ConL11 $=128$ filtering layers with $3 * 3$ stride function (17)
ConL12 $=512$ filtering layers with $5 * 5$ stride function (18)
The spatial pooling sequences from Parallel Layer-1, Parallel Layer-2, and Parallel Layer-3 have been arithmetically added by the CAM as depicted in Figure 4. The outputs of Parallel Layer-1 and Parallel Layer-2 are arithmetically added. At the same time period, the outputs of Parallel Layer-2 and Parallel Layer-3 are arithmetically added by the CAM. Now, the feature fusion module is used to fuse the output added feature values from the CAM unit, and the fused feature values are fed into Fully Connected Neural Network (FCNN) layers. This layer is particularly designed with three internal layers as depicted in the following Eqs. (19)–(22).
$F C N N$ layers $($TDCNet Architecture$)=\{F C N N$ layer -1, FCNN layer -2, FCNN layer -3$\}$ (19)
FCNN layer $-1=4096$ biased neurons with weighting function (20)
FCNN layer $-2=4096$ biased neurons with weighting function (21)
FCNN layer $-3=512$ biased neurons with weighting function (22)
All the neurons in the FCNN layer-3 are summed up, and the output value is transferred to the SoftMax layer to produce the significant AF/NAF classification results.
The SoftMax is another activation function that is used to perform the probability distribution on the output class of the proposed TDCNet Architecture. It is used to normalize the output process of the final FCNN layer as depicted in the following Eq. (23).
$f\left(x_i\right)=\frac{e^{x_i}}{\sum e^{x_{i k}}}$ (23)
The most important advantage of the proposed TDCNet model for AF signal classification is the parallel processing capability, which eliminates the sequential computation bottleneck that was encountered in recurrent networks such as Long Short-Term Memory (LSTM) and conventional CNN models. The parallel discriminative feature maps learning accelerates the convergence process, which further improves computational efficiency and facilitates the extraction of comprehensive temporal patterns from large-scale ECG recordings on both ECG datasets. The fusion of locally discriminative feature maps generated by the parallel architecture enables the proposed TDCNet model to obtain higher AF and NAF signal classification accuracy compared with conventional CNN models, LSTM, and vision transformer models.
Due to the parallel feature learning methodology of the proposed TDCNet model, more nummber of internal discriminative features are computed through the higher number of internal convolutional layers. The computational ECG signal classification accuracy is dependent on the highly derived internal fused feature maps. Due to the generation of higher derived discriminative feature maps through the subsequent and parallel configured convolutional layers, the ECG signal classification accuracy is increased. At the same time, the computational time period is significantly reduced due to the parallel feature learning methodology of the proposed TDCNet model. Hence, the performance of the proposed system is improved in terms of various analysis parameters.
Table 2 depicts the architectural specifications of the proposed TDCNet for the classification of the AF/NAF signals.
Table 2. Architectural specifications of the proposed TDCNet for the classification of AF/NAF signals
|
Architectural Layers |
Internal Layers |
Remarkable Specifications |
|
Convolutional layers |
ConL11 |
32 filtering layers with 3*3 stride function |
|
ConL12 |
512 filtering layers with 5*5 stride function |
|
|
ConL21 |
64 filtering layers with 3*3 stride function |
|
|
ConL22 |
512 filtering layers with 5*5 stride function |
|
|
ConL31 |
128 filtering layers with 3*3 stride function |
|
|
ConL32 |
512 filtering layers with 5*5 stride function |
|
|
Pooling layers |
P1 |
Spatial pooling with 4*4 sub window matrix |
|
P2 |
Spatial pooling with 4*4 sub window matrix |
|
|
P3 |
Spatial pooling with 4*4 sub window matrix |
|
|
FCNN layers |
FCNN layer1 |
4096 biased neurons with weighting function |
|
FCNN layer2 |
4096 biased neurons with weighting function |
|
|
FCNN layer3 |
512 biased neurons with weighting function |
Table 3 shows the hyperparameters that control the functionality of the proposed TDCNet for the classification of the AF/NAF signals. The proposed classification model architecture accuracy has been improving by varying the value of the hyperparameters. It balances the trade-off between the proposed internal architecture design and its computational time period.
Table 3. Specified experimental setup for hyperparameters of the proposed Transformer Driven Convolutional Neural Networks (TDCNet) architecture (ablation study)
|
Hyperparameter Name |
Specified Value |
|
Learning rate |
0.01 |
|
Convolutional layers count |
6 |
|
Filter size |
3*3 and 5*5 |
|
Stride |
2*2 |
|
Pooling type |
Spatial poolig function |
|
Batch size |
100 |
|
Epochs |
50 |
|
Dropout |
0.1 |
|
Optimizer |
RMSProp |
This research work for AF and NAF detection system uses two independent ECG signal datasets, MIT-BIH [18] and CPSC [19]. These ECG signal datasets are in the open access category, and hence a license is not required for accessing the ECG signals from these datasets in any research work. All the ECG signals and their capturing methods are verified by the expert team, and the patient who was involved during the creation of this dataset is not available.
The MIT-BIH dataset is created by Boston’s Beth Israel Hospital, and it is presently available under the department of Beth Israel Deaconess Medical Center. The 25 ECG recordings are obtained from each patient, and all these ECG signals are sampled at 512 samples per second with a 16-bit resolution format. This dataset contains 1400 AF signals and 1200 NAF signals. In this dataset, a 60:40 ratio has been maintained for the training and testing split, whereas 60% of ECG signals are used for training and the remaining 40% of ECG signals are used for testing the proposed ECG signal classification system in this research article. Hence, 840 AF signals are used for training, and the remaining 560 AF signals are used for testing. Similarly, 720 NAF signals are used for training, and 480 NAF signals are used for testing the proposed ECG signal classification system.
The CPSC dataset is created by the CPSC program in China, and it is presently available under the Department of Physiological Medical Center. The 15 ECG recordings are obtained from each patient, and all these ECG signals are sampled at 512 samples per second with a 12-bit resolution format. This dataset contains 1450 AF signals and 1750 NAF signals. In this dataset, a 60:40 ratio has been maintained for the training and testing split, whereas 60% of ECG signals are used for training and the remaining 40% of ECG signals are used for testing the proposed ECG signal classification system in this research article. Hence, 870 AF signals are used for training, and the remaining 580 AF signals are used for testing. Similarly, 1050 NAF signals are used for training, and 700 NAF signals are used for testing the proposed ECG signal classification system.
A 60:40 training-testing split was used for the evaluation of the proposed TDCNet model on the ECG dataset from the MIT-BIH and CPSC. This partition was chosen to ensure a good balance between the proposed model learning and fair assessment. These two independent datasets are large enough to contain a significant number of ECG recordings that have a high degree of inter-patient variability and diversity of rhythm. This guarantees sufficient data for the network to learn discriminative representations associated with AF and NAF signals. For classification problems, such as AF detection, a larger set of tests is desirable due to the fact that there are significant morphological variations within ECG signals from different patients. The 40% testing subset allows for thorough testing of the model under various physiological conditions and prevents overestimation of the performance. Additionally, the proposed split guarantees that enough samples are left for each rhythm category in both training and test sets, which helps prevent class imbalance in the training and evaluation sets and preserves class distribution.
The proposed ECG signal classification system has been performance measured with respect to the following Eqs. (24)–(28).
Signal detection Sensitivity $(S S e)=\frac{T P}{T P+F N}$ (24)
Signal detection Specificity $(S S p)=\frac{T N}{T N+F P}$ (25)
Signal detection Accuracy $(S A c)=\frac{T P+T N}{T P+T N+F P+F N}$ (26)
Signal detection Jaccard Index $(S J I)=\frac{T P}{T P+F P+F N}$ (27)
Signal detection Precision $(S P r)=\frac{T P}{T P+F P}$ (28)
where, TP and TN correlate the truly identified AF ECG signals and NAF signals. FP and FN correlate the falsely identified AF ECG signals and NAF signals.
All these ECG signal detection and classification system performances are evaluated on both open access ECG signal datasets individually, and the experimental results are given in the following section.
Table 4 shows the observation of ECG signal classification system performances on ECG signal datasets. The proposed ECG signal classification system obtains 98.8% SSe, 98.5% SSp, 99.3% Sac, 99.1% SJI and 98.9% SPr on the MIT-BIH dataset. The proposed ECG signal classification system obtains 99.3% SSe, 99.1% SSp, 98.9% Sac, 99.3% SJI and 99.1% SPr on the CPSC dataset.
Table 4. Observation of ECG signal classification system performances on ECG signal datasets
|
Measurement Parameters |
ECG Signal Datasets |
|
|
MIT-BIH |
CPSC |
|
|
SSe |
98.8 |
99.3 |
|
SSp |
98.5 |
99.1 |
|
Sac |
99.3 |
98.9 |
|
SJI |
99.1 |
99.3 |
|
SPr |
98.9 |
99.1 |
Figure 9. (a) Experimental results comparisons of the proposed Transformer Driven Convolutional Neural Networks (TDCNet) with other recent Electrocardiogram (ECG) signal classification systems on the MIT-BIH dataset, (b) Experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the China Physiological Signal Challenge (CPSC) dataset
In this research article, the proposed TDCNet classifier for ECG signal classification has been compared with other deep learning classification algorithms with respect to various parameters. The proposed TDCNet classifier obtained higher values of experimental results than the conventional Deep Neural Networks (DNN), ResNet50, Ensemble, and LSTM classifiers. The reason behind these higher experimental results by the proposed TDCNet classifier is that its feature computation process by the internal layers and its feature fusion process. This increases the performance of the proposed ECG classification system with respect to various experimental parameters on both ECG signal datasets. Figure 9(a) shows the experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the MIT-BIH dataset, and Figure 9(b) shows the experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the CPSC dataset.
Table 5 shows the experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the MIT-BIH dataset. From Table 4, the proposed TDCNet classification architecture proposed in this work attained higher experimental performance on the MIT-BIH dataset with respect to various experimental parameters.
Table 5. Experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the MIT-BIH dataset
|
Approaches |
Performance Measurement Parameters (%) |
||||
|
SSe |
SSp |
Sac |
SJI |
SPr |
|
|
Proposed TDCNet Classifier |
98.8 |
98.5 |
99.3 |
99.1 |
98.9 |
|
Azmy [11] |
97.1 |
97.3 |
97.9 |
97.6 |
97.3 |
|
Lim et al. [12] |
96.8 |
96.9 |
97.1 |
97.2 |
97.1 |
|
Xie et al. [13] |
96.4 |
96.6 |
96.5 |
96.6 |
96.7 |
|
Wang et al. [14] |
96.1 |
96.2 |
96.4 |
96.3 |
96.3 |
|
Gupta et al. [15] |
95.8 |
95.3 |
95.7 |
95.8 |
95.7 |
|
Pandey et al. [16] |
95.3 |
95.1 |
95.2 |
95.1 |
95.3 |
Table 6. Experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the CPSC dataset
|
Approaches |
Performance Measurement Parameters (%) |
||||
|
SSe |
SSp |
Sac |
SJI |
SPr |
|
|
Proposed TDCNet Classifier |
99.3 |
99.1 |
98.9 |
99.3 |
99.1 |
|
Azmy [11] |
98.2 |
98.5 |
98.1 |
98.4 |
97.3 |
|
Lim et al. [12] |
97.9 |
98.2 |
97.6 |
97.2 |
97.1 |
|
Xie et al. [13] |
97.3 |
97.8 |
97.3 |
96.7 |
96.8 |
|
Wang et al. [14] |
97.1 |
97.3 |
96.2 |
96.3 |
96.2 |
|
Gupta et al. [15] |
96.8 |
96.2 |
95.8 |
96.1 |
95.7 |
|
Pandey et al. [16] |
96.3 |
96.1 |
95.3 |
95.4 |
95.3 |
Table 6 shows the experimental results comparisons of the proposed TDCNet with other recent ECG signal classification systems on the CPSC dataset. From Table 6, the proposed TDCNet classification architecture proposed in this work attained higher experimental performance on the CPSC dataset with respect to various experimental parameters.
Table 7 shows the comparative analysis of computational time period (in milliseconds) on different ECG signal datasets. From Table 7, the computational time period is significantly lower in the proposed TDCNet model.
Table 8 shows the comparative analysis of memory requirements (in MB) on different ECG signal datasets. From Table 8, the memory requirement is significantly lower in the proposed TDCNet model.
Table 7. Comparative analysis of computational time period on different ECG signal datasets
|
Methods |
Computational Time Period (ms) |
|||||
|
MIT-BIH Dataset |
CPSC Dataset |
|||||
|
Proposed TDCNet model |
0.43 |
0.48 |
||||
|
Concentional CNN |
0.59 |
0.62 |
||||
|
LSTM model |
0.77 |
0.84 |
||||
|
Vision Transformer model |
0.82 |
0.94 |
||||
|
ResNet model |
0.98 |
1.11 |
||||
|
Pandey et al. [16] |
96.3 |
96.1 |
95.3 |
95.4 |
95.3 |
|
Table 8. Comparative analysis of memory requirements on different ECG signal datasets
|
Methods |
Memory Requirement (MB) |
|
|
MIT-BIH Dataset |
CPSC Dataset |
|
|
Proposed TDCNet model |
41 |
42 |
|
Conventional CNN |
52 |
58 |
|
LSTM model |
61 |
67 |
|
Vision Transformer model |
65 |
71 |
|
ResNet model |
98 |
87 |
In this research work, the measured experimental results of the proposed ECG signal classification system have been validated through the k-fold cross validation algorithm [20], where k denotes the number of folds to be tested on datasets. Here, 4 fold validations have been used for evaluating the performance of the proposed TDCNet classifier based ECG signal classification system with respect to various epoch counts and various ECG signal datasets.
Totally, 560 AF signals from the MIT-BIH dataset and 580 AF signals from the CPSC dataset are tested in this research work. Hence, each fold in the 4-fold validation contains 140 AF signals. One fold is tested by the proposed system, and the remaining folds are trained by the proposed system at 95% Confidential Interval (CI) and 0.14 standard deviation.
Table 9 illustrates the cross validation performances with respect to epochs on the MIT-BIH dataset. The Sac has been significantly increased with respect to the incremental value of epoch count in each cross validation count on the MIT-BIH dataset. The 99.3% Sac has been obtained at an epoch count of 50 on cross validation count 1. The 99.2% Sac has been obtained at an epoch count of 50 on cross validation count 2. The 99.3% Sac has been obtained at epoch count 50 on cross validation count 3. The 99.4% Sac has been obtained at an epoch count of 50 on cross validation count 4. Hence, the average 99.3% Sac has been obtained through this 4-fold cross validation algorithm on the MIT-BIH dataset at the maximum epoch count of 50.
Table 10 illustrates the cross validation performances with respect to epochs on the CPSC dataset. The Sac has been significantly increased with respect to the incremental value of epoch count in each cross validation count on the CPSC dataset. The 98.9% Sac has been obtained at epoch count 50 on cross validation count 1. The 98.8% Sac has been obtained at epoch count 50 on cross validation count 2. The 98.8% Sac has been obtained at epoch count 50 on cross validation count 3. The 98.9% Sac has been obtained at epoch count 50 on cross validation count 4. Hence, the average 98.9% Sac has been obtained through this 4-fold cross validation algorithm on the CPSC dataset at the maximum epoch count of 50.
Table 9. Cross validation performances with respect to epochs on Massachusetts Institute of Technology-Beth Israel Hospital (MIT-BIH) dataset
|
Cross Validation Count |
Epochs Count |
SAc |
|
1 |
10 |
97.4 |
|
20 |
97.9 |
|
|
30 |
98.6 |
|
|
40 |
98.9 |
|
|
50 |
99.3 |
|
|
2 |
10 |
97.3 |
|
20 |
97.8 |
|
|
30 |
98.4 |
|
|
40 |
98.8 |
|
|
50 |
99.2 |
|
|
3 |
10 |
98.1 |
|
20 |
98.4 |
|
|
30 |
98.8 |
|
|
40 |
99.1 |
|
|
50 |
99.3 |
|
|
4 |
10 |
96.9 |
|
20 |
97.3 |
|
|
30 |
97.9 |
|
|
40 |
98.7 |
|
|
50 |
99.4 |
Table 10. Cross validation performances with respect to epochs on China Physiological Signal Challenge (CPSC) dataset
|
Cross Validation Count |
Epochs Count |
SAc |
|
1 |
10 |
96.1 |
|
20 |
97.3 |
|
|
30 |
97.9 |
|
|
40 |
98.2 |
|
|
50 |
98.9 |
|
|
2 |
10 |
96.5 |
|
20 |
96.9 |
|
|
30 |
97.1 |
|
|
40 |
97.9 |
|
|
50 |
98.8 |
|
|
3 |
10 |
96.4 |
|
20 |
97.3 |
|
|
30 |
97.9 |
|
|
40 |
98.2 |
|
|
50 |
98.8 |
|
|
4 |
10 |
96.5 |
|
20 |
96.9 |
|
|
30 |
97.3 |
|
|
40 |
98.3 |
|
|
50 |
98.9 |
By comparing the average value of Sac of the k-fold validation algorithm with the experimental results of the proposed method, this proposed method has cross validated correctly by the k-fold algorithm in this research article.
The potential bias can be mitigated using k-fold validation as stated in the points.
A train-test split could be biased towards certain data distributions, resulting in an overoptimistic or pessimistic performance estimate. K-fold cross validation stated in this work, can be operated with several partitions of the training and testing datasets, so that the experimental results obtained in this research work are not affected by the specific splitting strategy, which mitigates the bias.
The proposed model's generalizability can be enhanced as stated below.
The proposed model needs to be reliable when applied to new patient data (ECG recordings or signals) in medical applications. The k-fold cross validation method helps to assess the stability of the proposed model by repeatedly training and testing the model using different sets of ECG recordings. Thus, the experimental results obtained by the proposed model are more realistic for generalizability.
In this research work, the AF and NAF ECG signals are differentiated using the proposed transformer diffused deep learning methodology. This proposed ECG signal classification algorithm has been significantly tested on MIT-BIH and CPSC datasets. The proposed ECG signal classification system obtains 98.8% SSe, 98.5% SSp, 99.3% Sac, 99.1% SJI and 98.9% SPr on the MIT-BIH dataset. The proposed ECG signal classification system obtains 99.3% SSe, 99.1% SSp, 98.9% Sac, 99.3% SJI and 99.1% SPr on the CPSC dataset. All these experimental values are compared by the recent ECG signal detection and classification methods. The measured experimental results of the proposed ECG signal classification system have been validated through the k-fold cross validation algorithm. The Sac has been significantly increased with respect to the incremental value of epoch count in each cross validation count on MIT-BIH and CPSC datasets. Though the proposed TDCNet based ECG signal classification system stated in this research article attained higher experimental results, it has certain limitations, and they are listed below.
All the experimental results obtained by the proposed methods are based on the open access ECG signal datasets, and no clinical ECG signal datasets are involved in this study to measure the robustness of the developed system, as identified as the main limitation of this work.
The AF ECG signals are not severity diagnosed by the proposed method, which is another limitation of this work.
These limitations of this research work will be eliminated in future work.
The ECG signals from the datasets that are used in this research work for AF and NAF signal classification process are free from dynamic noise environments and artifacts. Therefore, these datasets do not represent the real-world clinical noise conditions, which are identified as the main limitation of this work. Another challenge is the limited interpretability of the proposed TDCNet model with other clinical environment datasets. By incorporating the explainable Grad-CAM technique in the future, these limitations will be resolved to improve the real time clinical system.
The authors would like to thank their friends and colleagues for their constant help and support throughout the study and in obtaining the results.
[1] Gupta, V. (2023) Application of chaos theory for arrhythmia detection in pathological databases. International Journal of Medical Engineering and Informatics, 15(2): 191-202. https://doi.org/10.1504/IJMEI.2023.129353
[2] Serbes, G., Aydin, N. (2011) Modified dual tree complex wavelet transform for processing quadrature signals. Biomed Signal Process Control, 6(3): 301-306. https://doi.org/10.1016/j.bspc.2010.09.007
[3] Butt, F.S., Wagner, M.F., Schäfer, J., Ullate, D.G. (2022) Toward automated feature extraction for deep learning classification of electrocardiogram signals. IEEE Access, 10: 118601-118616. https://doi.org/10.1109/ACCESS.2022.3220670
[4] Abubaker, M.B., Babayigit, B. (2023) Detection of cardiovascular diseases in ECG images using machine learning and deep learning methods. IEEE Transactions on Artificial Intelligence, 4(2): 373-382. https://doi.org/10.1109/TAI.2022.3159505
[5] Prabhakararao, E., Dandapat, S. (2023) Congestive heart failure detection from ECG signals using deep residual neural network. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 53(5): 3008-3018. https://doi.org/10.1109/TSMC.2022.3221843
[6] Yang, S.X., Lian, C., Zeng, Z.G., Xu, B.R., Zang, J.B., Zhang, Z.D. (2023) A multi-view multi-scale neural network for multi-label ECG classification. IEEE Transactions on Emerging Topics in Computational Intelligence, 7(3): 648-660. https://doi.org/10.1109/TETCI.2023.3235374
[7] Le, D., Truong, S., Brijesh, P., Adjeroh, D.A., Le, N. (2023) sCL-ST: Supervised contrastive learning with semantic transformations for multiple lead ECG arrhythmia classification. IEEE Journal of Biomedical and Health Informatics, 27(6): 2818-2828. https://doi.org/10.1109/JBHI.2023.3246241
[8] Andreotti, F., Carr, O., Pimentel, M.A.F., Mahdi, A., De Vos, M. (2017) Comparing feature-based classifiers and convolutional neural networks to detect arrhythmia from short segments of ECG. Computational Cardiology, 44: 1-4. https://doi.org/10.22489/CinC.2017.360-239
[9] Datta, S., Puri, C., Mukherjee, A., et al. (2017). Identifying normal, AF and other abnormal ECG rhythms using a cascaded binary classifier. Computational Cardiology, 44: 1-4. https://doi.org/10.22489/CinC.2017.173-154
[10] García, M., Ródenas, J., Alcaraz, R., Rieta, J.J. (2017). Atrial fibrillation screening through combined timing features of short single-lead electrocardiograms. Computational Cardiology, 44: 1-4. https://doi.org/10.22489/CinC.2017.340-047
[11] Azmy, M.M. (2025). Detection of electrocardiogram atrial fibrillation using modified multifractal detrended fluctuation analysis based on discrete transforms and fractional Fourier transform. Discover Applied Sciences, 7: 1366. https://doi.org/10.1007/s42452-025-07086-y
[12] Lim, J., Han, D., Chon, K.H. (2025). QRS-centric beat-wise atrial fibrillation detection in ECG signals using deep neural networks. Computers in Biology and Medicine, 192(Part B): 110282. https://doi.org/10.1016/j.compbiomed.2025.110282
[13] Xie, J.X., Stavrakis, S., Yao, B. (2024). Automated identification of atrial fibrillation from single-lead ECGs using multi-branching ResNet. Frontiers in Physiology, 15: 1362185. https://doi.org/10.3389/fphys.2024.1362185
[14] Wang, B.C., Chen, G.R., Rong, L., Liu, Y.C., Yu, A.N., He, X.H. (2023). Arrhythmia disease diagnosis based on ECG time-frequency domain fusion and convolutional neural network. IEEE Journal of Translational Engineering in Health and Medicine, 11: 116-125. https://doi.org/ 10.1109/JTEHM.2022.3232791
[15] Gupta, K., Bajaj, V., Ansari, I.A. (2023). Atrial fibrillation detection using electrocardiogram signal input to LMD and ensemble classifier. IEEE Sensors Letters, 7(6): 1-4. https://doi.org/10.1109/LSENS.2023.3281129
[16] Pandey, S.K., Kumar, G., Shukla, S., Kumar, A., Singh, K.U., Mahato, S. (2022). Automatic detection of atrial fibrillation from ECG signal using hybrid deep learning techniques. Journal of Sensors, 1(1): 6732150. https://doi.org/10.1155/2022/6732150
[17] Khan, M.Z., Diwakar, M., Srivastava, P., et al. (2026). Multimodality medical image fusion using directional total variation based linear spectral clustering in NSCT domain. Scientific Reports, 16: 5367. https://doi.org/10.1038/s41598-025-26916-y
[18] Goldberger, A.L., Amaral, L.A.N., Glass, L., et al. (2000). PhysioBank, physiotoolkit, and physionet: Components of a new research resource for complex physiologic signals. Circulation, 101(23): e215-e220. https://doi.org/10.1161/01.cir.101.23.e215
[19] Liu, F.F., Liu, C.Y., Zhao, L.N., et al. (2018). An open access database for evaluating the algorithms of electrocardiogram rhythm and morphology abnormality detection. Journal of Medical Imaging and Health Informatics, 8(7): 1368-1373. https://doi.org/10.1166/jmihi.2018.2442
[20] Abedin, T., Xu, H.M., Uddin, S. (2026). The impact of K selection in K-fold cross-validation on bias and variance in supervised learning models. Scientific Reports, 16: 6084. https://doi.org/10.1038/s41598-026-37247-x