Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
InfoTech Spectrum: Iraqi Journal of Data Science

Search published articles

Search by title, keywords, author name, or all fields at once.

13 result(s) for “Computer science” Keywords

Research article 2025 Vol. 2 · No. 2

Loan Repayment Default Prediction Using Supervised Machine Learning Techniques on Financial Data

Ali RAAD · Muna Abdulmunem Othman · Ahmad Ghandour

With the enhancement of technology facilitating the expansion of businesses and thoughts, more and more people are applying for loans for personal or business use. However, banks have limited assets, which limit the amount of loans that can be granted. Identifying the right persons to grant loans to can be a time-consuming process. Banks seek to grant loans to individuals who can repay the loan on time, enabling the bank to obtain maximum profits. This work aims to solve the loan default problem with minimum costs to banks. This work consists of five main stages: pre-processing, feature extraction, machine learning techniques, evaluation models, and performance analysis to select the best machine learning models. Then, two datasets with different features are used. The first dataset has five features, and the second contains eighteen features. We are splitting the datasets into various training percentages (40%, 50%, 60% and 70%). The rest of the dataset is used for testing using only the Weka application. KNN is applied with different cross-validations, such as 15, 10, and 5, and different numbers of nearest neighbours (1, 5, 10, and 15). For the first dataset, the highest accuracy is 97.47% with two cross-validation values, 15 and 10, in the 10 nearest neighbours. The KNN was also implemented on the second dataset to compute the highest accuracy, 88.21% in three cross-validation values (15, 10, and 5) with the 15 nearest neighbours. Then, logistic regression is applied to compare the results of the correct classification value computed at the highest value of 96.93% with the (70% training set for the first dataset. The highest accuracy was obtained at 88.32% after splitting the second dataset (40%) for training and the rest for testing.

Research article 2025 Vol. 2 · No. 2

An Enhanced Intrusion Detection System for Wireless Sensor Networks Using Cuckoo-Optimized Neural Networks

Munther Twaij · Amir Lakizadeh

Wireless Sensor Networks are critical from the security point of view because of their distributed nature and resource constraints. Artificial Intelligence techniques have shown promising results in intrusion detection, but their performance optimization is paramount. This paper proposes a new approach based on combining Multi-Layer Perceptron neural networks with the Cuckoo Optimization Algorithm for efficient intrusion detection in WSN. Our methodology involves three main steps: (1) data preprocessing using the k-nearest neighbor for missing value imputation and normalization, (2) reduction of dimensionality through Principal Component Analysis, reducing the features from 41 to 38 dimensions, and (3) neural network optimization using COA for weight and bias parameter tuning. Our approach has yielded an accuracy of 99.1% in intrusion detection using the NSL-KDD dataset, which shows an improvement of about 3% compared to traditional methods. The proposed system performs better in terms of detection accuracy, reduction of false alarm rate, and computational efficiency.

Research article 2025 Vol. 2 · No. 2

An Advanced Framework for Intrusion Detection in Network Security Utilizing Machine Learning Algorithms: Challenges, Solutions, and Future Direction

Hussein Alrammahi · Mohammed Thakir Mahmood

Intrusion Detection Systems (IDS) are elementary building blocks of network security that can be used to detect unauthorized access and malicious activity. But traditional IDS approaches often suffer from problems such as high false positives, inability to adapt quickly to new threats, and scalability. This paper presents an advanced intrusion detection model that uses machine learning algorithms like Random Forest, Support Vector Machine (SVM), and Neural Networks to enhance detection. Using the KDD Cup 1999 data, the framework was highly preprocessed, feature engineered, and hyperparameters adjusted to achieve optimal performance. The Neural Network model outperformed other algorithms at 92.5% accuracy, 93.8% recall, and 92.4% F1-score, proving its ability to identify complex attack patterns with minimal false positives effectively. Additionally, the proposed framework reflected significant improvement over existing IDS solutions that always achieve accuracies of 80–85%. Intrusion Detection Systems (IDS) are important components of security, assuming the task of monitoring, detecting, and responding to unauthorized activities in network frameworks. This work's most notable contributions are its integration of sophisticated machine learning methods, systematic assessment of detection performance on a wide range of attack types, and comparison with well-established IDS benchmarks. In spite of facing issues like the complexity of the dataset and computational requirements, findings point to the efficacy of machine learning-based IDS in countering modern-day cybersecurity threats. Real-time data fusion and improving model interpretability for real-world implementation are areas that need to be addressed in the future.

Research article 2025 Vol. 2 · No. 1

Impact of Colour Space Transformation on Smoke Detection Accuracy using RESNET50

Mohannad Taha · Mohamed Safaa

Detecting smoke that precedes fire is a vital matter since it will detect fire incidents in a very early stage since these incidents have very high catastrophic effects on people's lives as well as industrial matters. In order to produce a more reliable detection system, in this article, we dove deeper to examine the effect of colour conversion of the captured footage to enhance the detection percentage using a pre-trained CNN model (ResNet50) that was altered to do a binary classification and was trained on a dataset that consists of smoke and non-smoke scenario images. We examined the system using the footage's original status (RGB) and also tested four colour spaces (HSV, YCbCr, LAB, and grayscale). The testing results showed that HSV had the highest accuracy of 92.1% and the lowest errors during training and testing. Regarding accuracy, the order after HSV was RGB, YCbCr, LAB, and finally, grayscale. Grayscale was the lowest in the testing results, with 85.4%. These results indicate that colour spaces do affect the detection quality and using them would improve the quality of smoke detection systems.

Research article 2025 Vol. 2 · No. 1

Medical Image Compression Utilizing The Serial Differences and Coding Techniques

Ghalib Ahmed Salman · Ahmed Ahmed · HAREER MOAIAD HUSSEN

Different medical devices for imaging used by centers and clinics produce an increasing number of sequential medical images. ‎Different imaging techniques such as Computed Tomography (CT), Magnetic Resonance Imaging (MRI) and Fluoroscopy ‎produce a set of series for the same patient. Within these images, most of the image parts are fixed against noticeable changes ‎in the remaining part. This consumes non-ignorable storage space. This paper proposes a near-lossless compression ‎technique that considers the fixed image parts to focus on changing parts for a higher compression ratio. In some applications, lossless compression techniques are highly preferable against preferring lossy techniques in some applications. In other applications, near-lossless compression techniques are preferable to lossless and lossy compression techniques, where lossy ones may ‎lose significant details, and the lossless ones produce less compression ratios than near-lossless ones. Previous works dealt with Fluoroscopy images as individual images or using ‎video compression techniques. This work tends to handle the whole series of ‎images as an integrated object. This paper considers subtracting successive ‎images to detect ROI areas producing zero overall values over similar ‎areas and non-zero ones within ROI ones. The double coding technique and near-lossless concept of compression increase the compression ratio. ‎Conducted experiments showed encouraging results benchmarking the other published ‎works in medical image compression.‎

Research article 2025 Vol. 2 · No. 1

Evaluating the Effectiveness of AI Tools in Mathematical Modelling of Various Life Phenomena: A Proposed Approach

Hadeel N. Abosaooda · Syaiba Balqish Ariffin · Osamah Mohammed Alyasiri · Ameen A. Noor

Advances in artificial intelligence (AI) are transforming the landscape of mathematical modelling in areas including physics, biology, and chemistry. Research suggests that ChatGPT, Gemini, and other AI tools can change the way researchers use simulation and modeling for complex phenomena by helping to produce models faster with less computational complexity and real-time insights. Here, we introduce a novel framework for building mathematical models of life sciences using AI tools for applications in disease dynamics and ecological systems. The approach integrates AI tools into the process for a hybrid model that combines initial model formulations based on AI-assisted discussions and refinements based on expert validation of AI-generated output. To give an example, if we are interested in modelling disease outbreaks, AI platforms such as ChatGPT or Gemini can instantly build a simple susceptible-infectious-recovered (SIR) model. This also helps with high dataset processing and making parameter suggestions based on real-time data, which in turn helps in the dynamic adaptation of models to changing data (e.g. transmission rates or intervention strategies). Likewise, in ecological modelling, AI tools can aid in the generation of predator-prey models that consider these complex interactions, such as habitat fragmentation or reserved zones and then suggest parameter sensitivities based on observed trends. These abilities make the future of AI-based mathematical modelling especially exciting, as they will further decrease the time that is traditionally spent by researchers on manually defining models and allow them to focus on result interpretation and strategic decision-making. With the rapidly changing advances in AI tools, incorporating some new capabilities and developments in the mathematical modelling procedure may allow for unprecedented improvements in predictive performance, model flexibility and interdisciplinary investigations. Further research and real-world efforts with this approach are needed to determine if AI tools can improve the cost-effectiveness and affordability of mathematical modelling in many fields of science.

Research article 2025 Vol. 2 · No. 1

Using Density Criterion and Increasing Modularity to Detect Communities in Complex Networks

Iman Hasan Abed · Sondos Bahadori

The selection of the initial centers of the communities is also significant in iteration-based methods for finding the communities in the networks. This is the reason why, if the initial centers of the communities are not chosen correctly, the errors and the time required for the application of the algorithm in the detection of the communities will be higher. Hence, selecting more significant nodes as starting points of communities can be the appropriate solution. Various techniques can be employed to achieve the selection of more significant nodes. In this thesis, the algorithm under discussion employs density and modularity criteria in the identification of communities in complex networks. This algorithm initially defines the number of nodes or the distinctive members of the community, in which these nodes have higher density levels and all the other nodes in their neighborhood have lower density levels. Next, the local communities are defined as the nodes that are in some way connected to the core nodes. Finally, the final communities are defined with the assistance of the merging algorithm, which is based on increasing modularity. In this algorithm, increasing modularity is used as a criterion for joining local communities together. Modularity is a criterion that indicates how the graph is like a modular or an organized community. When modularity becomes higher, local communities merge to form the final community. This means that it is possible to apply the presented algorithm and to use both density and modularity criteria to detect communities in complex networks. When the core nodes and local communities are first detected and then merged based on the increasing value of modularity, the resultant communities are more accurate. The results of the conducted experiments prove that the method applied in the Karate Club network clustering is equal to 0. 6913 for the NMI criterion and a value of 0. 733 for the accuracy criterion.

Research article 2024 Vol. 2 · No. 2

Evaluating AI Language Models in News Retrieval: A Comparative Study Of ChatGPT-Plus and DeepSeek (R1)

Omar Al-Janabi · Osamah Mohammed Alyasiri · Elaf Ayyed Jebur · Shahad Mohgoob Nafl

The increasing complexity of how humans interact with and process information has demonstrated significant advancements in Natural Language Processing (NLP), transitioning from task-specific architectures to generalized frameworks applicable across multiple tasks. Despite their success, challenges persist in specialized domains such as translation, where instruction tuning may prioritize fluency over accuracy. Against this backdrop, the present study conducts a comparative evaluation of ChatGPT-Plus and DeepSeek (R1) on a high-fidelity bilingual retrieval-and-translation task. A single standardize prompt directs each model to access the Arabic-language news section of the College of Medicine, University of Baghdad, retrieve the three most recent articles, and translate them into English. ChatGPT-Plus fulfilled the prompt successfully, extracting authentic Arabic content and delivering fluent, semantically accurate English translations. DeepSeek (R1), by contrast, failed to retrieve the requested articles and instead produced only generic procedural advice – evidence of its lack of real-time web access and a retrieval-augmented generation (RAG) mechanism.

Research article 2024 Vol. 1 · No. 1

Evaluating The Impact of Feature Extraction Techniques on Arabic Reviews Classification

Hawraa Alshammary · Mohammed Fadhil Ibrahim · Hafsa Ataallah Hussein

With the advent of AI text-based tools and applications, the need to introduce and investigate word-processing tools has also been raised. NLP tools and techniques have developed rapidly for some languages, such as English. However, other languages, such as Arabic, still need to introduce more methods and techniques to provide more explanations. In this study, we present a sample to classify customer reviews which are written in Arabic. The data set (HARD) is used to be certified as a dataset for work. This study adopted four classifications in machine learning and deep learning (CNN, RNN, NB, LR). In addition, the texts were cleaned using data cleaning techniques, and the stemming technique was used, and three types of them were implemented (Khoja Stemmer, Snowball Stemmer, Thashaphyne Stemmer). Moreover, two methods of feature extraction were used (TF-IDF, N-gram). The results of the model provided several explanations. The best performance resulted from the use of (CNN+ Snowball Stemmer +N-gram) with accuracy (%93.5). The results of the model stated that some workbooks are sensitive to the use of different tools, and some accuracy performance can also be affected if there are different methods for extracting the features used. Either feature extraction has an impact on accuracy performance. The model also proved that colloquial Arabic could cause some limitations because different dialects can give different meanings across different regions or countries. The results of the study open the door to exploring other tools and methods to enrich natural Arabic language processing and contribute to the development of new applications that support Arabic content.

Research article 2024 Vol. 1 · No. 1

A Hybrid Technique Based on RF-PCA and ANN for Detecting DDoS Attacks IoT

Hayder Jalo · Mohsen Heydarian

The increasing reliance on smart products has increased vulnerabilities in Internet of Things (IoT) traffic, which poses significant security risks. These vulnerabilities allowed some hackers to exploit them, which led to system performance degradation. Attacks can lead to these vulnerabilities to various undesirable outcomes, including data leakage, economic losses, data breaches, operational disruptions, and damage to the company's reputation. To address these security challenges, network intrusion detection alarms play a crucial role in assessing system security. In recent years, the proliferation of intelligent and soft computing-based algorithmic and structural frameworks has been evident. However, previous studies have faced challenges related to comprehensiveness, zero-day attacks, realism, and data interpretation. In light of these concerns, this study proposes to design a neural network for proactive detection of attacks. Moreover, we propose to use a hybrid system called RF-PCA to facilitate dimensionality reduction and help classifiers. Notably, this is the first application of a BOT-IoT data set in such an approach. The study also includes a discussion of relevant IoT terms in the context of our work. The proposed method uses high-level data features to represent and draw conclusive conclusions. To evaluate its effectiveness, an experiment was conducted using Python as the programming environment, achieving a remarkable detection rate of 99.73%.

Research article 2024 Vol. 1 · No. 1

Rotation Invariant Technique for Sign Language Recognition

Mohamed T. Dardoh Al-Obaidi · Ali M. Sahan · Ali S. Al-Itbi

Sign language recognition is an assistive technology that has garnered significant attention from researchers, particularly with respect to its potential benefits for individuals with hearing impairments. This paper proposes an effective technique for sign language recognition based on the Contourlet Transform (CT) and deep learning. The CT is employed in the pre-processing stage to reduce complexity and processing time, while deep learning is utilized to extract and classify sign language features. The proposed method was evaluated using two sign language databases: a direct feed database and an American sign language database. The experimental analysis demonstrated that the proposed method gives good results in processing time by more than 70% while maintaining high accuracy