An Optimized Sub Group Partition based Healthcare Data Mining in Big Data |
||||
|
|
||||
|
||||
BibTeX: |
||||
|
@article{IJIRSTV4I10023, |
||||
Abstract: |
||||
|
Every piece of information learned in human health has the potential to increase the length of patient life. However, the search for more knowledge in the health informatics domain need to analysis vast quantities of data which cannot be easily processed by researchers. Such data are handled by Big data is one of the latest technologies which have the potential for radically changing healthcare data. The main aim of this paper is to introduce efficient data mining techniques for prediction of different diseases like leukemia cancer, lung cancer and heart disease. Initially, a healthcare data is collected and it is divided into partitions and each partition is loaded in a map. Then the firefly optimization algorithm process is carried out on each map which returns subgroups for each partition. However, the subgroups have been obtained by considering best features of partition. In this paper, subgroup is obtained by updating fireflies based on the global best features of each partition. This process reduces the number of features for the classification of healthcare data which saves the time and speed of classifiers without affecting the classification accuracy. The selected features are get into three different classifiers are Naïve Bayes, C4.5 and Random forest which predicts the lung cancer, leukemia cancer and heart disease effectively. The experimental results are conducted to prove the effectiveness of the proposed method in terms of accuracy, precision, recall and F-measure. |
||||
Keywords: |
||||
|
Clinical data mining, big data, Leukemia cancer prediction, lung cancer prediction, heart disease prediction |
||||



