Data Distribution Valuation

Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but that of the distribution from which the dataset was sampled. For example, consider a buyer trying to evaluate whether to purchase data from different vendors. The buyer may observe (and compare) only a small preview sample from each vendor, to decide which vendor's data distribution is most useful to the buyer and purchase. The core question is how should we compare the values of data distributions from their samples? Under a Huber characterization of the data heterogeneity across vendors, we propose a maximum mean discrepancy (MMD)-based valuation method which enables theoretically principled and actionable policies for comparing data distributions from samples. We empirically demonstrate that our method is sample-efficient and effective in identifying valuable data distributions against several existing baselines, on multiple real-world datasets (e.g., network intrusion detection, credit card fraud detection) and downstream applications (classification, regression).

Paper

References (68)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Sj3P6/10 · confidence 4/52024-06-23

Summary

This paper starts by highlighting the importance of accurately assessing the value of data distributions, especially in the growing data economy. To address the problem, this paper introduces a data distribution valuation method based on Maximum Mean Discrepancy (MMD) for comparing the value of data distributions from samples. This paper assumes that each vendor $i$ holds a distribution $P_i$ which follows a Huber model, and is a mixture of ground truth distribution $P^*$ and arbitrary distribution $Q_i$. Based on this assumption, the authors discuss the theoretical foundations and assumptions, providing detailed proofs and derivations for the proposed methods. The study addresses heterogeneity in data distributions and the challenges of combining multiple data vendors' datasets. Experimental results demonstrate the sample efficiency and effectiveness of the proposed methods in ranking data distributions.

Strengths

1. This paper provides a detailed explanation of the theoretical foundations for the valuation of data distributions, including assumptions and proofs, ensuring the rigor and completeness of the theory. 2. This paper studies a meaningful problem: how to compare the values of data distributions from their samples, which can help evaluate the value of data provided by different vendors. 3. This paper is well-organized and easy to follow. The paper introduces a novel method based on maximum mean discrepancy (MMD) for data distribution valuation, offering a fresh perspective in the field.

Weaknesses

1. This paper relies on certain assumptions, such as the Huber model of data heterogeneity, which may not always hold in real-world scenarios. 2. The experiment for ranking data distributions lacks generalizability. 3. Using data samples to represent data distribution can cause issues, such as dealing with malicious data vendors.

Questions

1. In real-world scenarios, data distributions often encompass various complex heterogeneity factors that the Huber model may not accurately capture. Data collected from different sources often exhibit significant variability and may not follow the same distribution. 2. The experiment in the article for ranking data distributions involves mixing two similar datasets, which fails to adequately represent the heterogeneity between datasets provided by different vendors, thus affecting the generalizability of the results. 3. In real-world scenarios, some vendors may provide data samples that do not accurately reflect their true data distribution. Some malicious vendors might even forge samples (for example, using some real data as samples while the rest of dataset is synthetic or useless data). The article does not take this into account.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

NA

Reviewer s86g6/10 · confidence 5/52024-07-13

Summary

The valuation of data is crucial in data marketplaces. Instead of assessing the value of a specific dataset, this paper focuses on the valuation of data distribution behind the dataset itself. For example, several vendors are trying to sell different or even the same datasets, what is the best distribution to purchase when we only observe a sampled dataset? The authors model the problem of data distribution valuation and use the MMD-based method as a metric to evaluate the valuation. They provide theoretically guaranteed policies for buyers to take action and empirically demonstrate it using real-world datasets. The results indicate that the method is sample-efficient and outperforms other valuation metrics.

Strengths

* Originality: This paper is interesting as it evaluates the distribution behind a sampled dataset rather than the dataset itself. This approach is new and can be valuable when dealing with partially sampled datasets. * Quality: The results seem promising and support the use of MMD-based metrics for assessing data distribution. * Clarity: The writing and formulation are clear and easy to understand. * Significance: In data marketplaces, most data is only available for preview and is often sampled. Understanding the value of data distribution is beneficial for the field and users.

Weaknesses

1. The paper's motivation is interesting and contributes to the field. However, I believe there is a missing experiment regarding ranking different error levels of distributions. Suppose we have five distributions ranging from 100% correct to 0% correct. Can the valuation score accurately rank these distributions or reveal their actual error levels? This experiment differs from directly comparing the valuation of the dataset itself. We should see a regression line, where the x-axis is the error level of distribution, while the y-axis is the valuation score. 2. Why do the correlation scores drop when we move the dataset from CIFAR10 to CIFAR100? The description about this is not clear to me. More clarification on this will be helpful. 3. The method of sampling from a distribution to create a dataset can influence its evaluation. Have there been any empirical findings or methods to address sampling bias? This missing experiment can justify the robustness of the valuation function.

Questions

1. The method used to sample from a distribution to construct a dataset can affect its valuation. Are there any empirical results or methods to overcome sampling bias? 2. How does the accuracy of valuation change when a large number of vendors contribute to the mixed reference distribution? Also, does the valuation score reveal the relative level of two distributions or their absolute values? 3. In Equation 1, what is the reason for giving the value function a negative term instead of taking the reciprocal?

Rating

6

Confidence

5

Soundness

2

Presentation

2

Contribution

3

Limitations

I didn't see any potential negative societal impact of their work.

Reviewer YGYg7/10 · confidence 4/52024-07-13

Summary

This paper addresses the problem of data distribution valuation in data markets, where buyers need to evaluate the quality of data distributions to make informed purchasing decisions. The authors formulate the problem and identify three technical challenges: heterogeneity modeling, defining the value of a sampling distribution, and choosing a reference data distribution. They make three key design choices: assuming a Huber model for data heterogeneity, using negative maximum mean discrepancy (MMD) as the value metric, and considering a class of convex mixtures of vendor distributions as the reference. The paper derives an error guarantee and comparison policy for the proposed method and demonstrates its effectiveness on real-world classification and regression tasks. Overall, this work provides a novel framework for data distribution valuation, enabling buyers to make informed decisions in data markets.

Strengths

1. The paper introduces a novel approach to data valuation by focusing on the value of the underlying data distribution from a small sample. This addresses a gap in existing methods, which typically do not formalize the value of sampling distributions or provide actionable policies for comparing them. 2. The paper employs a Huber model to capture data heterogeneity and utilizes the maximum mean discrepancy (MMD) for evaluating sampling distributions. This combination allows for precise, theoretically grounded comparisons of data distributions and allows for sample efficient assessment of the valuation. Assuming a convex combination of the distribution as the reference they provide an error guarantee without making the common assumption of knowing the reference distribution. 3. The authors validate their method through real-world classification and regression tasks. The demonstrated sample efficiency and effectiveness of their MMD-based valuation method, particularly its superior performance in most classification settings compared to existing metrics highlights the practical relevance and robustness of their approach.

Weaknesses

1. Based on Theorem 1 the valuation of $D$ boils down to the samples available from it and the averaged out heterogeneity $d(Q_\omega, P^*)$. In practice, a small sample that looks cleaner could be preferred over a bigger but noisy sample. It looks like the current results do not account for individual noise (heterogeneity), I see it in the Huber model but the theoretical results are averaging out as in Observation 1. 2. The sample complexity of $O(\frac{1}{\sqrt{m}})$ to estimate MMD, might even be sufficient for the learning task at hand, so why would vendors be willing to show such a big sample of the dataset and by seeing samples from each vendor won’t one be able to accomplish the learning objective without even selecting a vendor and buying a larger sample from them.

Questions

See above.

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes.

Reviewer YGYg2024-08-08

Thank you for the rebuttal and congratulations for the nice work! I have increased my score. I presumed that MMD will be computed on the sample shown to the buyer. It would be nice to have some discussion in the paper to avoid this confusion. Minor stuff: please improve the presentation of results in the tables. You don't have to respond to this. I am curious about, how vendors can create the samples without revealing too much. In particular in your setting, as a buyer I can collect the freely visible samples from multiple vendors and train a model. If the vendor distributions are diverse enough then a buyer would still be able to get a very good model from those free samples (freeloader, lol). I am curious about when such freeloading is possible and when it is not, theoretically and can the vendors do something to prevent it. If it requires them to co-ordinate that might introduce conflicts.

Authorsrebuttal2024-08-09

Thank you for increasing the score

We would like to thank Reviewer YGYg for acknowledging our rebuttal and increasing the score. We will definitely take note of your feedback and incorporate it into our revision. As for your question on if/"how vendors can create the samples without revealing too much", and if it requires coordination among vendors and its implications, we believe it to be an interesting avenue to explore an algorithmic solution intersecting statistics/machine learning and game theory and will definitely note this in our revision. Thanks again for your encouraging feedback and helpful comments.

Reviewer s86g2024-08-09

Thanks for conducting additional experiments and clearing my concerns on sampling, more vendors, and different error levels' valuations. Regression lines in your attachment have a negative correlation and enhance the quality of this work. Thanks for clarifying my questions on cifar10 and cifar100. I have increased my rating and support this paper. Thanks for your response!

Authorsrebuttal2024-08-09

Thank you for the acknowledgement and increasing your score

We wish to thank Reviewer s86g for acknowledging our rebuttal and increasing the score. We really appreciate your support!

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC