Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
InfoTech Spectrum: Iraqi Journal of Data Science

Search published articles

Search by title, keywords, author name, or all fields at once.

3 result(s) for “Feature (linguistics)” Keywords

Research article 2026 Vol. 3 · No. 1

A Deep Learning Framework for Extracting and Summarizing Text from Images

Abbas EL DOR · Osama Emad Abdulhussein

In the digital era, substantial amounts of textual information are embedded in images, especially across news outlets, social platforms, and scanned documents. This presents a significant technical challenge: efficiently extracting and summarizing text from images in an automated way that preserves context and meaning. Traditional text summarization techniques are not directly applicable to image-based content because they depend on pre-structured input text. In this paper, we propose a framework that integrates Optical Character Recognition (OCR) and advanced Natural Language Processing (NLP) models to address this challenge. The proposed method implements OCR to extract raw text from images, followed by deep learning-based summarization using models such as LSTM, Bi-LSTM, BERT and T5. These models are trained on large-scale news datasets to enhance their ability to generate coherent summaries from unstructured text. To ensure accessibility and practical usability, our framework is deployed via an interactive web-based interface that allows end-users to upload images and receive concise summaries in real time. Experimental evaluation demonstrates the efficacy of the proposed approach, particularly with transformer-based models, in delivering high-quality summarization from visual text sources

Research article 2025 Vol. 3 · No. 1

Enhancing Optical Coherence Tomography Image Classification Via Swarm Optimization-Based Feature Selection and Machine Learning Models

Zaid Al-Jubouri · Sarah Saadoon Jasim

Optical Coherence Tomography (OCT) greatly facilitates the diagnosis of retinal diseases. However, traditional models based on Convolutional Neural Networks (CNNs) suffer from challenges, most notably high computational cost, sensitivity to noise, and data imbalance. This study aims to compare three hybrid deep learning frameworks, all of which rely on feature extraction using a pre-trained CNN model and then selecting the most important features using intelligent swarm algorithms: the Dolphin Swarm Optimization (DSO), the Particle Swarm Optimization (PSO), and the Ant Swarm Optimization (ACO). The selected features were evaluated using four classifiers: SVM, random forest, XGBoost, and k-NN. Experiments were conducted on a standard dataset from the University of California, San Diego (UCSD) and a local dataset. The comparison results showed that the hybrid framework, which combines the dolphin swarm algorithm and SVM, outperformed the other combinations, achieving a classification accuracy of 93% on local data and 95% on standard data, while also outperforming them in terms of accuracy and computational efficiency.

Research article 2024 Vol. 1 · No. 1

Evaluating The Impact of Feature Extraction Techniques on Arabic Reviews Classification

Hawraa Alshammary · Mohammed Fadhil Ibrahim · Hafsa Ataallah Hussein

With the advent of AI text-based tools and applications, the need to introduce and investigate word-processing tools has also been raised. NLP tools and techniques have developed rapidly for some languages, such as English. However, other languages, such as Arabic, still need to introduce more methods and techniques to provide more explanations. In this study, we present a sample to classify customer reviews which are written in Arabic. The data set (HARD) is used to be certified as a dataset for work. This study adopted four classifications in machine learning and deep learning (CNN, RNN, NB, LR). In addition, the texts were cleaned using data cleaning techniques, and the stemming technique was used, and three types of them were implemented (Khoja Stemmer, Snowball Stemmer, Thashaphyne Stemmer). Moreover, two methods of feature extraction were used (TF-IDF, N-gram). The results of the model provided several explanations. The best performance resulted from the use of (CNN+ Snowball Stemmer +N-gram) with accuracy (%93.5). The results of the model stated that some workbooks are sensitive to the use of different tools, and some accuracy performance can also be affected if there are different methods for extracting the features used. Either feature extraction has an impact on accuracy performance. The model also proved that colloquial Arabic could cause some limitations because different dialects can give different meanings across different regions or countries. The results of the study open the door to exploring other tools and methods to enrich natural Arabic language processing and contribute to the development of new applications that support Arabic content.