Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
InfoTech Spectrum: Iraqi Journal of Data Science

Search published articles

Search by title, keywords, author name, or all fields at once.

5 result(s) for “Data mining” Keywords

Research article 2025 Vol. 3 · No. 1

Proposed Model for Credit Card Fraud Detection Model Using Machine Learning Technique

Harith Safwan Ezzulddin

The online payment system is at high risk due to the increasing rates of credit card theft. The primary objective is to identify cases of credit card theft by analysing the purchase history of cardholders and categorising them accordingly. These include an increase in the slope of logistics, steep slope, and scattered woodlands. The proposed model utilises tools such as logistic regression and random forest as machine learning techniques. Additionally, a set of preprocessing techniques is employed, including data balancing using SMOTE. After being trained on a large dataset of credit card transactions, the model is used to detect trends and anomalies that may indicate fraudulent activity, taking into account factors such as transaction amount, location, and time of day. We have used artificial minority oversampling to put the data set into proper perspective. The two algorithms were applied, yielding 97.34% accuracy for Logistic Regression and 99.99% accuracy for Random Forest. The accuracy metric is used for performance evaluation. The results indicate a promising performance that can enhance credit card security, potentially helping to reduce financial losses to victims of fraud.

Research article 2025 Vol. 2 · No. 2

An Advanced Framework for Intrusion Detection in Network Security Utilizing Machine Learning Algorithms: Challenges, Solutions, and Future Direction

Hussein Alrammahi · Mohammed Thakir Mahmood

Intrusion Detection Systems (IDS) are elementary building blocks of network security that can be used to detect unauthorized access and malicious activity. But traditional IDS approaches often suffer from problems such as high false positives, inability to adapt quickly to new threats, and scalability. This paper presents an advanced intrusion detection model that uses machine learning algorithms like Random Forest, Support Vector Machine (SVM), and Neural Networks to enhance detection. Using the KDD Cup 1999 data, the framework was highly preprocessed, feature engineered, and hyperparameters adjusted to achieve optimal performance. The Neural Network model outperformed other algorithms at 92.5% accuracy, 93.8% recall, and 92.4% F1-score, proving its ability to identify complex attack patterns with minimal false positives effectively. Additionally, the proposed framework reflected significant improvement over existing IDS solutions that always achieve accuracies of 80–85%. Intrusion Detection Systems (IDS) are important components of security, assuming the task of monitoring, detecting, and responding to unauthorized activities in network frameworks. This work's most notable contributions are its integration of sophisticated machine learning methods, systematic assessment of detection performance on a wide range of attack types, and comparison with well-established IDS benchmarks. In spite of facing issues like the complexity of the dataset and computational requirements, findings point to the efficacy of machine learning-based IDS in countering modern-day cybersecurity threats. Real-time data fusion and improving model interpretability for real-world implementation are areas that need to be addressed in the future.

Research article 2025 Vol. 2 · No. 1

Medical Image Compression Utilizing The Serial Differences and Coding Techniques

Ghalib Ahmed Salman · Ahmed Ahmed · HAREER MOAIAD HUSSEN

Different medical devices for imaging used by centers and clinics produce an increasing number of sequential medical images. ‎Different imaging techniques such as Computed Tomography (CT), Magnetic Resonance Imaging (MRI) and Fluoroscopy ‎produce a set of series for the same patient. Within these images, most of the image parts are fixed against noticeable changes ‎in the remaining part. This consumes non-ignorable storage space. This paper proposes a near-lossless compression ‎technique that considers the fixed image parts to focus on changing parts for a higher compression ratio. In some applications, lossless compression techniques are highly preferable against preferring lossy techniques in some applications. In other applications, near-lossless compression techniques are preferable to lossless and lossy compression techniques, where lossy ones may ‎lose significant details, and the lossless ones produce less compression ratios than near-lossless ones. Previous works dealt with Fluoroscopy images as individual images or using ‎video compression techniques. This work tends to handle the whole series of ‎images as an integrated object. This paper considers subtracting successive ‎images to detect ROI areas producing zero overall values over similar ‎areas and non-zero ones within ROI ones. The double coding technique and near-lossless concept of compression increase the compression ratio. ‎Conducted experiments showed encouraging results benchmarking the other published ‎works in medical image compression.‎

Research article 2025 Vol. 2 · No. 1

Using Density Criterion and Increasing Modularity to Detect Communities in Complex Networks

Iman Hasan Abed · Sondos Bahadori

The selection of the initial centers of the communities is also significant in iteration-based methods for finding the communities in the networks. This is the reason why, if the initial centers of the communities are not chosen correctly, the errors and the time required for the application of the algorithm in the detection of the communities will be higher. Hence, selecting more significant nodes as starting points of communities can be the appropriate solution. Various techniques can be employed to achieve the selection of more significant nodes. In this thesis, the algorithm under discussion employs density and modularity criteria in the identification of communities in complex networks. This algorithm initially defines the number of nodes or the distinctive members of the community, in which these nodes have higher density levels and all the other nodes in their neighborhood have lower density levels. Next, the local communities are defined as the nodes that are in some way connected to the core nodes. Finally, the final communities are defined with the assistance of the merging algorithm, which is based on increasing modularity. In this algorithm, increasing modularity is used as a criterion for joining local communities together. Modularity is a criterion that indicates how the graph is like a modular or an organized community. When modularity becomes higher, local communities merge to form the final community. This means that it is possible to apply the presented algorithm and to use both density and modularity criteria to detect communities in complex networks. When the core nodes and local communities are first detected and then merged based on the increasing value of modularity, the resultant communities are more accurate. The results of the conducted experiments prove that the method applied in the Karate Club network clustering is equal to 0. 6913 for the NMI criterion and a value of 0. 733 for the accuracy criterion.

Research article 2024 Vol. 1 · No. 1

Enhancing Malware Detection Through Machine Learning Techniques

Zeina S. Jassim · Mohamad M. Kassir

Malware detection is important to computer network security since it is the principal attack vector against modern enterprises. As a result, firms must remove viruses from computer systems. Using artificial intelligence, namely machine learning techniques, to function in real-time with an IT system is the ideal solution to this problem. This issue has yet to be fixed, but it is still significant because a lack of processing power and memory constrains these features. The most popular method for evaluating systems and intrusion detection models is using the Application Program Interface (API) calls via the KDD-CUP99 data set to give this solution. KDD-CUP99 has more than three hundred thousand samples, each with 54 features. However, the data set attributes were designed and chosen to provide us with a high malware detection rate. The quality of this data was lowered to produce results. To get the desired results, the attributes of this data were reduced. Data transformation and purification are used in this process. Inaccurate, unnecessary, duplicated, or missing information is eliminated by data cleansing. Data cleaning eliminates inaccurate, excessive, redundant, or lacking information. By comparing this study to earlier research that employed lengthy sequences of software interface (API) calls with the same machine-learning classifiers, data transformation includes discretization, which transforms the continuous process of discretizing continuous data into discrete forms is a type of data transformation. Using more advanced algorithms to do the task at hand with the best precision and the least expense increases accuracy and performance. The data set was divided into two categories using a Support Vector Machine (SVM), Decision Tree (DT), and Iterative Dichotomiser 3 (ID3). The findings revealed that little previous research uses a five-class classification strategy for malware detection. The accuracy of several works is comparable to the accuracy acquired in the proposed work.