Open Access Semi-annual

InfoTech Spectrum: Iraqi Journal of Data Science

· eISSN 3007-5467 · DOI 10.51173/ijds
Vol.1 · Iss.1
Volume 1 · Issue 1 · January 2024

Vol. 1 No. 1 (2024): January 2024

Research article (5)

Research article

Enhancing Malware Detection Through Machine Learning Techniques

Malware detection is important to computer network security since it is the principal attack vector against modern enterprises. As a result, firms must remove viruses from computer systems. Using artificial intelligence, namely machine learning techniques, to function in real-time with an IT system is the ideal solution to this problem. This issue has yet to be fixed, but it is still significant because a lack of processing power and memory constrains these features. The most popular method for evaluating systems and intrusion detection models is using the Application Program Interface (API) calls via the KDD-CUP99 data set to give this solution. KDD-CUP99 has more than three hundred thousand samples, each with 54 features. However, the data set attributes were designed and chosen to provide us with a high malware detection rate. The quality of this data was lowered to produce results. To get the desired results, the attributes of this data were reduced. Data transformation and purification are used in this process. Inaccurate, unnecessary, duplicated, or missing information is eliminated by data cleansing. Data cleaning eliminates inaccurate, excessive, redundant, or lacking information. By comparing this study to earlier research that employed lengthy sequences of software interface (API) calls with the same machine-learning classifiers, data transformation includes discretization, which transforms the continuous process of discretizing continuous data into discrete forms is a type of data transformation. Using more advanced algorithms to do the task at hand with the best precision and the least expense increases accuracy and performance. The data set was divided into two categories using a Support Vector Machine (SVM), Decision Tree (DT), and Iterative Dichotomiser 3 (ID3). The findings revealed that little previous research uses a five-class classification strategy for malware detection. The accuracy of several works is comparable to the accuracy acquired in the proposed work.

Read
Research article

Rotation Invariant Technique for Sign Language Recognition

Sign language recognition is an assistive technology that has garnered significant attention from researchers, particularly with respect to its potential benefits for individuals with hearing impairments. This paper proposes an effective technique for sign language recognition based on the Contourlet Transform (CT) and deep learning. The CT is employed in the pre-processing stage to reduce complexity and processing time, while deep learning is utilized to extract and classify sign language features. The proposed method was evaluated using two sign language databases: a direct feed database and an American sign language database. The experimental analysis demonstrated that the proposed method gives good results in processing time by more than 70% while maintaining high accuracy

Read
Research article

A Hybrid Technique Based on RF-PCA and ANN for Detecting DDoS Attacks IoT

The increasing reliance on smart products has increased vulnerabilities in Internet of Things (IoT) traffic, which poses significant security risks. These vulnerabilities allowed some hackers to exploit them, which led to system performance degradation. Attacks can lead to these vulnerabilities to various undesirable outcomes, including data leakage, economic losses, data breaches, operational disruptions, and damage to the company's reputation. To address these security challenges, network intrusion detection alarms play a crucial role in assessing system security. In recent years, the proliferation of intelligent and soft computing-based algorithmic and structural frameworks has been evident. However, previous studies have faced challenges related to comprehensiveness, zero-day attacks, realism, and data interpretation. In light of these concerns, this study proposes to design a neural network for proactive detection of attacks. Moreover, we propose to use a hybrid system called RF-PCA to facilitate dimensionality reduction and help classifiers. Notably, this is the first application of a BOT-IoT data set in such an approach. The study also includes a discussion of relevant IoT terms in the context of our work. The proposed method uses high-level data features to represent and draw conclusive conclusions. To evaluate its effectiveness, an experiment was conducted using Python as the programming environment, achieving a remarkable detection rate of 99.73%.

Read
Research article

Evaluating The Impact of Feature Extraction Techniques on Arabic Reviews Classification

With the advent of AI text-based tools and applications, the need to introduce and investigate word-processing tools has also been raised. NLP tools and techniques have developed rapidly for some languages, such as English. However, other languages, such as Arabic, still need to introduce more methods and techniques to provide more explanations. In this study, we present a sample to classify customer reviews which are written in Arabic. The data set (HARD) is used to be certified as a dataset for work. This study adopted four classifications in machine learning and deep learning (CNN, RNN, NB, LR). In addition, the texts were cleaned using data cleaning techniques, and the stemming technique was used, and three types of them were implemented (Khoja Stemmer, Snowball Stemmer, Thashaphyne Stemmer). Moreover, two methods of feature extraction were used (TF-IDF, N-gram). The results of the model provided several explanations. The best performance resulted from the use of (CNN+ Snowball Stemmer +N-gram) with accuracy (%93.5). The results of the model stated that some workbooks are sensitive to the use of different tools, and some accuracy performance can also be affected if there are different methods for extracting the features used. Either feature extraction has an impact on accuracy performance. The model also proved that colloquial Arabic could cause some limitations because different dialects can give different meanings across different regions or countries. The results of the study open the door to exploring other tools and methods to enrich natural Arabic language processing and contribute to the development of new applications that support Arabic content.

Read