FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context

We advance sketch research to scenes with the first dataset of freehand scene\nsketches, FS-COCO. With practical applications in mind, we collect sketches\nthat convey scene content well but can be sketched within a few minutes by a\nperson with any sketching skills. Our dataset comprises 10,000 freehand scene\nvector sketches with per point space-time information by 100 non-expert\nindividuals, offering both object- and scene-level abstraction. Each sketch is\naugmented with its text description. Using our dataset, we study for the first\ntime the problem of fine-grained image retrieval from freehand scene sketches\nand sketch captions. We draw insights on: (i) Scene salience encoded in\nsketches using the strokes temporal order; (ii) Performance comparison of image\nretrieval from a scene sketch and an image caption; (iii) Complementarity of\ninformation in sketches and image captions, as well as the potential benefit of\ncombining the two modalities. In addition, we extend a popular vector sketch\nLSTM-based encoder to handle sketches with larger complexity than was supported\nby previous work. Namely, we propose a hierarchical sketch decoder, which we\nleverage at a sketch-specific "pre-text" task. Our dataset enables for the\nfirst time research on freehand scene sketch understanding and its practical\napplications.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC