Nithya.Jayaprakash, Ms. Caroline Mary
The paper proposes a method for detecting and deleting distance based outliers in very large data sets. This is based on the outlier detection solving set algorithm. This method introduces parallel computation so as to save more time and having excellent performance. First, weights are assigned to each of the data in the data sets. Based on the weights outliers from all the data sets are obtained by using the distance based method and finally they are all deleted. By deleting the outliers, it increases the space for storing more data.