Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
InfoTech Spectrum: Iraqi Journal of Data Science

Search published articles

Search by title, keywords, author name, or all fields at once.

7 result(s) for “Natural language processing” Keywords

Research article 2026 Vol. 3 · No. 1

ChatGPT: Precision Answer Comparison and Evaluation Model

Aso Mohammed Aladdin · Rebwar Khalid Muhammed · Hemin Sardar Abdulla · Tarik Ahmad Rashid

Artificial Intelligence (AI) has made advancements, among other things, OpenAI created the sophisticated model ChatGPT. Conversational, ChatGPT supports natural interactions, providing human-like responses to queries across myriad topics. But it is not infallible, and the degree of accuracy also depends on the complexity of the queries, the context, and how often the prompts are repeated. This work thus proposes a new model, the Precision Answer Comparison and Evaluation Model (PACEM), to systematically address these types of questions and assess ChatGPT's performance. PACEM assesses the correctness and coherence of ChatGPT's answers across numerous fields, including literature, history, law, ethics, and sports. By providing these analyses and comparisons, PACEM goes on record with a detailed understanding of what ChatGPT does well and poorly as a source of reliable information. On top of that, it includes an assessment of response time, considering ChatGPT's speed in producing answers in relation to real or expected ones. The findings show that ChatGPT's answers are usually substantially accurate and often of superior quality compared to those written by the user and other alternatives. Response time generally increases with the complexity or length of the answer. Finally, the study reviews notable takeaways from PACEM's deployment and offers suggestions for future research to address the evolving challenges in AI-driven response assessment.

Research article 2026 Vol. 3 · No. 1

A Deep Learning Framework for Extracting and Summarizing Text from Images

Abbas EL DOR · Osama Emad Abdulhussein

In the digital era, substantial amounts of textual information are embedded in images, especially across news outlets, social platforms, and scanned documents. This presents a significant technical challenge: efficiently extracting and summarizing text from images in an automated way that preserves context and meaning. Traditional text summarization techniques are not directly applicable to image-based content because they depend on pre-structured input text. In this paper, we propose a framework that integrates Optical Character Recognition (OCR) and advanced Natural Language Processing (NLP) models to address this challenge. The proposed method implements OCR to extract raw text from images, followed by deep learning-based summarization using models such as LSTM, Bi-LSTM, BERT and T5. These models are trained on large-scale news datasets to enhance their ability to generate coherent summaries from unstructured text. To ensure accessibility and practical usability, our framework is deployed via an interactive web-based interface that allows end-users to upload images and receive concise summaries in real time. Experimental evaluation demonstrates the efficacy of the proposed approach, particularly with transformer-based models, in delivering high-quality summarization from visual text sources

Research article 2025 Vol. 3 · No. 1

Proposed Model for Credit Card Fraud Detection Model Using Machine Learning Technique

Harith Safwan Ezzulddin

The online payment system is at high risk due to the increasing rates of credit card theft. The primary objective is to identify cases of credit card theft by analysing the purchase history of cardholders and categorising them accordingly. These include an increase in the slope of logistics, steep slope, and scattered woodlands. The proposed model utilises tools such as logistic regression and random forest as machine learning techniques. Additionally, a set of preprocessing techniques is employed, including data balancing using SMOTE. After being trained on a large dataset of credit card transactions, the model is used to detect trends and anomalies that may indicate fraudulent activity, taking into account factors such as transaction amount, location, and time of day. We have used artificial minority oversampling to put the data set into proper perspective. The two algorithms were applied, yielding 97.34% accuracy for Logistic Regression and 99.99% accuracy for Random Forest. The accuracy metric is used for performance evaluation. The results indicate a promising performance that can enhance credit card security, potentially helping to reduce financial losses to victims of fraud.

Research article 2025 Vol. 2 · No. 1

Medical Image Compression Utilizing The Serial Differences and Coding Techniques

Ghalib Ahmed Salman · Ahmed Ahmed · HAREER MOAIAD HUSSEN

Different medical devices for imaging used by centers and clinics produce an increasing number of sequential medical images. ‎Different imaging techniques such as Computed Tomography (CT), Magnetic Resonance Imaging (MRI) and Fluoroscopy ‎produce a set of series for the same patient. Within these images, most of the image parts are fixed against noticeable changes ‎in the remaining part. This consumes non-ignorable storage space. This paper proposes a near-lossless compression ‎technique that considers the fixed image parts to focus on changing parts for a higher compression ratio. In some applications, lossless compression techniques are highly preferable against preferring lossy techniques in some applications. In other applications, near-lossless compression techniques are preferable to lossless and lossy compression techniques, where lossy ones may ‎lose significant details, and the lossless ones produce less compression ratios than near-lossless ones. Previous works dealt with Fluoroscopy images as individual images or using ‎video compression techniques. This work tends to handle the whole series of ‎images as an integrated object. This paper considers subtracting successive ‎images to detect ROI areas producing zero overall values over similar ‎areas and non-zero ones within ROI ones. The double coding technique and near-lossless concept of compression increase the compression ratio. ‎Conducted experiments showed encouraging results benchmarking the other published ‎works in medical image compression.‎

Research article 2024 Vol. 2 · No. 2

Evaluating AI Language Models in News Retrieval: A Comparative Study Of ChatGPT-Plus and DeepSeek (R1)

Omar Al-Janabi · Osamah Mohammed Alyasiri · Elaf Ayyed Jebur · Shahad Mohgoob Nafl

The increasing complexity of how humans interact with and process information has demonstrated significant advancements in Natural Language Processing (NLP), transitioning from task-specific architectures to generalized frameworks applicable across multiple tasks. Despite their success, challenges persist in specialized domains such as translation, where instruction tuning may prioritize fluency over accuracy. Against this backdrop, the present study conducts a comparative evaluation of ChatGPT-Plus and DeepSeek (R1) on a high-fidelity bilingual retrieval-and-translation task. A single standardize prompt directs each model to access the Arabic-language news section of the College of Medicine, University of Baghdad, retrieve the three most recent articles, and translate them into English. ChatGPT-Plus fulfilled the prompt successfully, extracting authentic Arabic content and delivering fluent, semantically accurate English translations. DeepSeek (R1), by contrast, failed to retrieve the requested articles and instead produced only generic procedural advice – evidence of its lack of real-time web access and a retrieval-augmented generation (RAG) mechanism.

Research article 2024 Vol. 1 · No. 1

Evaluating The Impact of Feature Extraction Techniques on Arabic Reviews Classification

Hawraa Alshammary · Mohammed Fadhil Ibrahim · Hafsa Ataallah Hussein

With the advent of AI text-based tools and applications, the need to introduce and investigate word-processing tools has also been raised. NLP tools and techniques have developed rapidly for some languages, such as English. However, other languages, such as Arabic, still need to introduce more methods and techniques to provide more explanations. In this study, we present a sample to classify customer reviews which are written in Arabic. The data set (HARD) is used to be certified as a dataset for work. This study adopted four classifications in machine learning and deep learning (CNN, RNN, NB, LR). In addition, the texts were cleaned using data cleaning techniques, and the stemming technique was used, and three types of them were implemented (Khoja Stemmer, Snowball Stemmer, Thashaphyne Stemmer). Moreover, two methods of feature extraction were used (TF-IDF, N-gram). The results of the model provided several explanations. The best performance resulted from the use of (CNN+ Snowball Stemmer +N-gram) with accuracy (%93.5). The results of the model stated that some workbooks are sensitive to the use of different tools, and some accuracy performance can also be affected if there are different methods for extracting the features used. Either feature extraction has an impact on accuracy performance. The model also proved that colloquial Arabic could cause some limitations because different dialects can give different meanings across different regions or countries. The results of the study open the door to exploring other tools and methods to enrich natural Arabic language processing and contribute to the development of new applications that support Arabic content.

Research article 2024 Vol. 1 · No. 1

Rotation Invariant Technique for Sign Language Recognition

Mohamed T. Dardoh Al-Obaidi · Ali M. Sahan · Ali S. Al-Itbi

Sign language recognition is an assistive technology that has garnered significant attention from researchers, particularly with respect to its potential benefits for individuals with hearing impairments. This paper proposes an effective technique for sign language recognition based on the Contourlet Transform (CT) and deep learning. The CT is employed in the pre-processing stage to reduce complexity and processing time, while deep learning is utilized to extract and classify sign language features. The proposed method was evaluated using two sign language databases: a direct feed database and an American sign language database. The experimental analysis demonstrated that the proposed method gives good results in processing time by more than 70% while maintaining high accuracy