LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision

Curating high-quality datasets of images and annotations is a major bottleneck for deploying robust vision systems over vast, unlabeled data lakes. We introduce Labeling Copilot, the first deep research agent that automates end-to-end vision data curation. A central orchestrator agent, powered by a large multimodal language model, uses multi-step reasoning to coordinate specialized tools across three core capabilities: (1) Calibrated Discovery retrieves relevant, in-distribution images from large repositories; (2) Controllable Synthesis generates novel variants for rare scenarios; and (3) Consensus Annotation produces accurate labels by orchestrating multiple foundation models via a consensus mechanism built on non-maximum suppression and voting. Our large-scale validation proves the effectiveness of Labeling Copilot's components. The Consensus Annotation module excels at object discovery: on the dense COCO dataset, it averages 14.2 candidate proposals per image-nearly double the 7.4 ground-truth objects-achieving a final annotation mAP of 37.1%. On the web-scale Open Images dataset, it navigates extreme class imbalance to discover 903 new bounding box categories, expanding its capability to over 1500 total. Concurrently, our Calibrated Discovery tool, tested at a 10-million-sample scale, features an active learning strategy that is up to $40 \times$ more computationally efficient than alternatives with equivalent sample efficiency. These experiments show that an agentic workflow with optimized, scalable tools provides a robust foundation for curating industrial-scale vision datasets.

Paper

References (62)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC