Large-scale empirical validation of Bayesian Network structure learning algorithms with noisy data

Numerous Bayesian Network (BN) structure learning algorithms have been\nproposed in the literature over the past few decades. Each publication makes an\nempirical or theoretical case for the algorithm proposed in that publication\nand results across studies are often inconsistent in their claims about which\nalgorithm is 'best'. This is partly because there is no agreed evaluation\napproach to determine their effectiveness. Moreover, each algorithm is based on\na set of assumptions, such as complete data and causal sufficiency, and tend to\nbe evaluated with data that conforms to these assumptions, however unrealistic\nthese assumptions may be in the real world. As a result, it is widely accepted\nthat synthetic performance overestimates real performance, although to what\ndegree this may happen remains unknown. This paper investigates the performance\nof 15 structure learning algorithms. We propose a methodology that applies the\nalgorithms to data that incorporates synthetic noise, in an effort to better\nunderstand the performance of structure learning algorithms when applied to\nreal data. Each algorithm is tested over multiple case studies, sample sizes,\ntypes of noise, and assessed with multiple evaluation criteria. This work\ninvolved approximately 10,000 graphs with a total structure learning runtime of\nseven months. It provides the first large-scale empirical validation of BN\nstructure learning algorithms under different assumptions of data noise. The\nresults suggest that traditional synthetic performance may overestimate\nreal-world performance by anywhere between 10% and more than 50%. They also\nshow that while score-based learning is generally superior to constraint-based\nlearning, a higher fitting score does not necessarily imply a more accurate\ncausal graph. To facilitate comparisons with future studies, we have made all\ndata, raw results, graphs and BN models freely available online.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC