Rethinking Streaming Machine Learning Evaluation

While most work on evaluating machine learning (ML) models focuses on computing accuracy on batches of data, tracking accuracy alone in a streaming setting (i.e., unbounded, timestamp-ordered datasets) fails to appropriately identify when models are performing unexpectedly. In this position paper, we discuss how the nature of streaming ML problems introduces new real-world challenges (e.g., delayed arrival of labels) and recommend additional metrics to assess streaming ML performance.

Paper

References (10)

09Ease.ML: A Lifecycle Management System for MLDev and MLOps2021 · In Conference on Innovative Data Systems Research
10DABS: a Domain-Agnostic Benchmark for Self-Supervised Learning2021 · In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round

Similar papers

© 2026 NYSGPT2525 LLC