Weighted Frequent Itemset (WFI) mining is an important model in data mining. It aims to discover all itemsets whose weighted sum in a transactional database is no less than the user-specified threshold value. Most previous works focused on finding WFIs in a transactional database and did not recognize the spatiotemporal characteristics of an item within the data. This paper proposes a more flexible model of Spatial Weighted Frequent Itemset (SWFI) that may exist in a spatiotemporal database. The recommended patterns may be found very useful in many real-world applications. For instance, an SWFI generated from an air pollution database indicates a geographical region where people have been exposed to high levels of an air pollutant, say PM2.5. The generated SWFIs do not satisfy the anti-monotonic property. Two new measures have been presented to effectively reduce the search space and the computational cost of finding the desired patterns. A pattern-growth algorithm, called Spatial Weighted Frequent Pattern-growth, has also been presented to find all SWFIs in a spatiotemporal database. Experimental results demonstrate that the proposed algorithm is efficient. We also describe a case study in which our model has been used to find useful information in air pollution database.
Paper
Full text
Discovering Spatial Weighted Frequent Itemsets in Spatiotemporal Databases
Semantic Scholar · Computer Science · 2019
Abstract
Weighted Frequent Itemset (WFI) mining is an important model in data mining. It aims to discover all itemsets whose weighted sum in a transactional database is no less than the user-specified threshold value. Most previous works focused on finding WFIs in a transactional database and did not recognize the spatiotemporal characteristics of an item within the data. This paper proposes a more flexible model of Spatial Weighted Frequent Itemset (SWFI) that may exist in a spatiotemporal database. The recommended patterns may be found very useful in many real-world applications. For instance, an SWFI generated from an air pollution database indicates a geographical region where people have been exposed to high levels of an air pollutant, say PM2.5. The generated SWFIs do not satisfy the anti-monotonic property. Two new measures have been presented to effectively reduce the search space and the computational cost of finding the desired patterns. A pattern-growth algorithm, called Spatial Weighted Frequent Pattern-growth, has also been presented to find all SWFIs in a spatiotemporal database. Experimental results demonstrate that the proposed algorithm is efficient. We also describe a case study in which our model has been used to find useful information in air pollution database.