Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
InfoTech Spectrum: Iraqi Journal of Data Science

Search published articles

Search by title, keywords, author name, or all fields at once.

6 result(s) for “Machine learning” Keywords

Research article 2026 Vol. 3 · No. 1

A Deep Learning Framework for Extracting and Summarizing Text from Images

Abbas EL DOR · Osama Emad Abdulhussein

In the digital era, substantial amounts of textual information are embedded in images, especially across news outlets, social platforms, and scanned documents. This presents a significant technical challenge: efficiently extracting and summarizing text from images in an automated way that preserves context and meaning. Traditional text summarization techniques are not directly applicable to image-based content because they depend on pre-structured input text. In this paper, we propose a framework that integrates Optical Character Recognition (OCR) and advanced Natural Language Processing (NLP) models to address this challenge. The proposed method implements OCR to extract raw text from images, followed by deep learning-based summarization using models such as LSTM, Bi-LSTM, BERT and T5. These models are trained on large-scale news datasets to enhance their ability to generate coherent summaries from unstructured text. To ensure accessibility and practical usability, our framework is deployed via an interactive web-based interface that allows end-users to upload images and receive concise summaries in real time. Experimental evaluation demonstrates the efficacy of the proposed approach, particularly with transformer-based models, in delivering high-quality summarization from visual text sources

Research article 2025 Vol. 2 · No. 2

Loan Repayment Default Prediction Using Supervised Machine Learning Techniques on Financial Data

Ali RAAD · Muna Abdulmunem Othman · Ahmad Ghandour

With the enhancement of technology facilitating the expansion of businesses and thoughts, more and more people are applying for loans for personal or business use. However, banks have limited assets, which limit the amount of loans that can be granted. Identifying the right persons to grant loans to can be a time-consuming process. Banks seek to grant loans to individuals who can repay the loan on time, enabling the bank to obtain maximum profits. This work aims to solve the loan default problem with minimum costs to banks. This work consists of five main stages: pre-processing, feature extraction, machine learning techniques, evaluation models, and performance analysis to select the best machine learning models. Then, two datasets with different features are used. The first dataset has five features, and the second contains eighteen features. We are splitting the datasets into various training percentages (40%, 50%, 60% and 70%). The rest of the dataset is used for testing using only the Weka application. KNN is applied with different cross-validations, such as 15, 10, and 5, and different numbers of nearest neighbours (1, 5, 10, and 15). For the first dataset, the highest accuracy is 97.47% with two cross-validation values, 15 and 10, in the 10 nearest neighbours. The KNN was also implemented on the second dataset to compute the highest accuracy, 88.21% in three cross-validation values (15, 10, and 5) with the 15 nearest neighbours. Then, logistic regression is applied to compare the results of the correct classification value computed at the highest value of 96.93% with the (70% training set for the first dataset. The highest accuracy was obtained at 88.32% after splitting the second dataset (40%) for training and the rest for testing.

Research article 2025 Vol. 2 · No. 2

An Enhanced Intrusion Detection System for Wireless Sensor Networks Using Cuckoo-Optimized Neural Networks

Munther Twaij · Amir Lakizadeh

Wireless Sensor Networks are critical from the security point of view because of their distributed nature and resource constraints. Artificial Intelligence techniques have shown promising results in intrusion detection, but their performance optimization is paramount. This paper proposes a new approach based on combining Multi-Layer Perceptron neural networks with the Cuckoo Optimization Algorithm for efficient intrusion detection in WSN. Our methodology involves three main steps: (1) data preprocessing using the k-nearest neighbor for missing value imputation and normalization, (2) reduction of dimensionality through Principal Component Analysis, reducing the features from 41 to 38 dimensions, and (3) neural network optimization using COA for weight and bias parameter tuning. Our approach has yielded an accuracy of 99.1% in intrusion detection using the NSL-KDD dataset, which shows an improvement of about 3% compared to traditional methods. The proposed system performs better in terms of detection accuracy, reduction of false alarm rate, and computational efficiency.

Research article 2025 Vol. 2 · No. 2

An Advanced Framework for Intrusion Detection in Network Security Utilizing Machine Learning Algorithms: Challenges, Solutions, and Future Direction

Hussein Alrammahi · Mohammed Thakir Mahmood

Intrusion Detection Systems (IDS) are elementary building blocks of network security that can be used to detect unauthorized access and malicious activity. But traditional IDS approaches often suffer from problems such as high false positives, inability to adapt quickly to new threats, and scalability. This paper presents an advanced intrusion detection model that uses machine learning algorithms like Random Forest, Support Vector Machine (SVM), and Neural Networks to enhance detection. Using the KDD Cup 1999 data, the framework was highly preprocessed, feature engineered, and hyperparameters adjusted to achieve optimal performance. The Neural Network model outperformed other algorithms at 92.5% accuracy, 93.8% recall, and 92.4% F1-score, proving its ability to identify complex attack patterns with minimal false positives effectively. Additionally, the proposed framework reflected significant improvement over existing IDS solutions that always achieve accuracies of 80–85%. Intrusion Detection Systems (IDS) are important components of security, assuming the task of monitoring, detecting, and responding to unauthorized activities in network frameworks. This work's most notable contributions are its integration of sophisticated machine learning methods, systematic assessment of detection performance on a wide range of attack types, and comparison with well-established IDS benchmarks. In spite of facing issues like the complexity of the dataset and computational requirements, findings point to the efficacy of machine learning-based IDS in countering modern-day cybersecurity threats. Real-time data fusion and improving model interpretability for real-world implementation are areas that need to be addressed in the future.

Research article 2024 Vol. 2 · No. 2

Evaluating AI Language Models in News Retrieval: A Comparative Study Of ChatGPT-Plus and DeepSeek (R1)

Omar Al-Janabi · Osamah Mohammed Alyasiri · Elaf Ayyed Jebur · Shahad Mohgoob Nafl

The increasing complexity of how humans interact with and process information has demonstrated significant advancements in Natural Language Processing (NLP), transitioning from task-specific architectures to generalized frameworks applicable across multiple tasks. Despite their success, challenges persist in specialized domains such as translation, where instruction tuning may prioritize fluency over accuracy. Against this backdrop, the present study conducts a comparative evaluation of ChatGPT-Plus and DeepSeek (R1) on a high-fidelity bilingual retrieval-and-translation task. A single standardize prompt directs each model to access the Arabic-language news section of the College of Medicine, University of Baghdad, retrieve the three most recent articles, and translate them into English. ChatGPT-Plus fulfilled the prompt successfully, extracting authentic Arabic content and delivering fluent, semantically accurate English translations. DeepSeek (R1), by contrast, failed to retrieve the requested articles and instead produced only generic procedural advice – evidence of its lack of real-time web access and a retrieval-augmented generation (RAG) mechanism.

Research article 2024 Vol. 1 · No. 1

Enhancing Malware Detection Through Machine Learning Techniques

Zeina S. Jassim · Mohamad M. Kassir

Malware detection is important to computer network security since it is the principal attack vector against modern enterprises. As a result, firms must remove viruses from computer systems. Using artificial intelligence, namely machine learning techniques, to function in real-time with an IT system is the ideal solution to this problem. This issue has yet to be fixed, but it is still significant because a lack of processing power and memory constrains these features. The most popular method for evaluating systems and intrusion detection models is using the Application Program Interface (API) calls via the KDD-CUP99 data set to give this solution. KDD-CUP99 has more than three hundred thousand samples, each with 54 features. However, the data set attributes were designed and chosen to provide us with a high malware detection rate. The quality of this data was lowered to produce results. To get the desired results, the attributes of this data were reduced. Data transformation and purification are used in this process. Inaccurate, unnecessary, duplicated, or missing information is eliminated by data cleansing. Data cleaning eliminates inaccurate, excessive, redundant, or lacking information. By comparing this study to earlier research that employed lengthy sequences of software interface (API) calls with the same machine-learning classifiers, data transformation includes discretization, which transforms the continuous process of discretizing continuous data into discrete forms is a type of data transformation. Using more advanced algorithms to do the task at hand with the best precision and the least expense increases accuracy and performance. The data set was divided into two categories using a Support Vector Machine (SVM), Decision Tree (DT), and Iterative Dichotomiser 3 (ID3). The findings revealed that little previous research uses a five-class classification strategy for malware detection. The accuracy of several works is comparable to the accuracy acquired in the proposed work.