TY - GEN
T1 - Generative AI with Big Data for Better Detection of Fraud in Medical Claims
AU - El-Enen, Mohamed Ahmed Abo
AU - Tbaishat, Dina
AU - AbdulRazek, Mustafa
AU - Nazir, Amril
AU - Muhammad, Reem
AU - Sahlol, Ahmed T.
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Generative AI refers to a type of algorithms that can generate new content. This can be text, images, or any other form of data. While Large Language Models (LLMs) are a specific type of generative AI that focuses on language, they are basically trained on massive textual input to understand and generate human-like text. This paper addresses the critical challenge of fraud detection in medical insurance claims, a pervasive issue causing significant financial losses in healthcare. This work is focused on devising a robust, automated system for detecting fraudulent activities. Where an integration of Generative AI, specifically LLM is implemented with Big Data processing frameworks to enhancements in fraud detection in medical claims. Each LLM was used as an embedding layer that transforms textual features of a real-world insurance claim data into numerical representations. These claims data has been collected from countries belonging to the Mena region. The results show advantages towards LLMs that were trained on specialized medical contexts as they show better capability of understanding medical expressions which reflects model’s performance. Applying further sampling techniques such as class weight and up-sampling did not have a significant impact on the LLMs performance, with a little better performance for class weight. Gemini showed advantages over BERT medical language models on most experiments by achieving 90.44% of classification accuracy.
AB - Generative AI refers to a type of algorithms that can generate new content. This can be text, images, or any other form of data. While Large Language Models (LLMs) are a specific type of generative AI that focuses on language, they are basically trained on massive textual input to understand and generate human-like text. This paper addresses the critical challenge of fraud detection in medical insurance claims, a pervasive issue causing significant financial losses in healthcare. This work is focused on devising a robust, automated system for detecting fraudulent activities. Where an integration of Generative AI, specifically LLM is implemented with Big Data processing frameworks to enhancements in fraud detection in medical claims. Each LLM was used as an embedding layer that transforms textual features of a real-world insurance claim data into numerical representations. These claims data has been collected from countries belonging to the Mena region. The results show advantages towards LLMs that were trained on specialized medical contexts as they show better capability of understanding medical expressions which reflects model’s performance. Applying further sampling techniques such as class weight and up-sampling did not have a significant impact on the LLMs performance, with a little better performance for class weight. Gemini showed advantages over BERT medical language models on most experiments by achieving 90.44% of classification accuracy.
KW - Big Data
KW - Fraud Detection
KW - Generative AI
KW - Large Language Models
KW - Medical Claims
UR - https://www.scopus.com/pages/publications/85219604589
U2 - 10.1109/HEALTHCOM60970.2024.10880817
DO - 10.1109/HEALTHCOM60970.2024.10880817
M3 - Conference contribution
AN - SCOPUS:85219604589
T3 - 2024 IEEE International Conference on E-Health Networking, Application and Services, HealthCom 2024
BT - 2024 IEEE International Conference on E-Health Networking, Application and Services, HealthCom 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2024 IEEE International Conference on E-Health Networking, Application and Services, HealthCom 2024
Y2 - 18 November 2024 through 20 November 2024
ER -