TY - GEN
T1 - Leveraging Memetic Algorithm and Machine Learning Methods for Email-Based Spam Detection
AU - Al-Ali, Mariam Khalid
AU - Alteneiji, Manal Ali
AU - Hashem, Ibrahim Abaker
AU - Shareef, Omar Salah F.
AU - Hussain, Abir
AU - Turky, Ayad
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Email is the most frequently utilized method of communication between people nowadays. This is because most people in institutions, banks, universities, hospitals, and others depend mainly on e-mail to communicate with each other and transfer information and data effectively and quickly. However, it has been observed in recent years that fraudsters have become smarter in defrauding people, especially via email, and deceiving them by creating fake accounts and impersonation, which has led to an increase in cybercrimes such as phishing and impersonating others to steal their money and bank account data. In this paper, we use a machine learning (ML) based approach in email header analysis as a powerful tool for detecting phishing and spam emails. Specifically, the efficacy of three different machine learning algorithms has been trained and tested in the model, which are: Naive Bayes Classifier (NB Classifier), Multi-Layer Perceptron Classifier (MLP Classifier), and Random Forest. Consequently, our approach, which focuses on using genetic algorithms and simulated annealing for feature selection, effectively detects spam, ham, and phishing emails with high accuracy, precision, and recall. Also, it achieves a 99.69% accuracy for spam detection and a 99.12% accuracy for phishing detection. The findings are compared to some state-of-the-art models and indicate the supremacy of our model.
AB - Email is the most frequently utilized method of communication between people nowadays. This is because most people in institutions, banks, universities, hospitals, and others depend mainly on e-mail to communicate with each other and transfer information and data effectively and quickly. However, it has been observed in recent years that fraudsters have become smarter in defrauding people, especially via email, and deceiving them by creating fake accounts and impersonation, which has led to an increase in cybercrimes such as phishing and impersonating others to steal their money and bank account data. In this paper, we use a machine learning (ML) based approach in email header analysis as a powerful tool for detecting phishing and spam emails. Specifically, the efficacy of three different machine learning algorithms has been trained and tested in the model, which are: Naive Bayes Classifier (NB Classifier), Multi-Layer Perceptron Classifier (MLP Classifier), and Random Forest. Consequently, our approach, which focuses on using genetic algorithms and simulated annealing for feature selection, effectively detects spam, ham, and phishing emails with high accuracy, precision, and recall. Also, it achieves a 99.69% accuracy for spam detection and a 99.12% accuracy for phishing detection. The findings are compared to some state-of-the-art models and indicate the supremacy of our model.
KW - Cybercrimes
KW - Email Header
KW - Machine Learning
KW - Spam Email
UR - https://www.scopus.com/pages/publications/105000472082
U2 - 10.1109/DeSE63988.2024.10912024
DO - 10.1109/DeSE63988.2024.10912024
M3 - Conference contribution
AN - SCOPUS:105000472082
T3 - Proceedings - International Conference on Developments in eSystems Engineering, DeSE
SP - 123
EP - 128
BT - 17th International Conference on Developments in eSystems Engineering, DeSE 2024
A2 - Al-Jumeily, Dhiya
A2 - Assi, Sulaf
A2 - Jayabalan, Manoj
A2 - Hind, Jade
A2 - Hussain, Abir
A2 - Tawfik, Hissam
A2 - Rowe, Neil
A2 - Mustafina, Jamila
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 17th International Conference on Developments in eSystems Engineering, DeSE 2024
Y2 - 6 November 2024 through 8 November 2024
ER -