Object Recognition as Classification via Visual Properties

We base our work on the teleosemantic modelling of concepts as abilities implementing the distinct functions of recognition and classification. Accordingly, we model two types of conceptssubstance concepts suited for object recognition exploiting visual properties, and classification concepts suited for classification of substance concepts exploiting linguistically grounded properties. The goal in this paper is to demonstrate that object recognition can be construed as classification of visual properties, as distinct from work in mainstream computer vision. Towards that, we present an object recognition process based on Ranganathan’s four-phased faceted knowledge organization process, grounded in the teleosemantic distinctions of substance concept and classification concept. We also briefly introduce the ongoing project MultiMedia UKC, whose aim is to build an object recognition resource following our proposed process. 1.0 Introduction Recognition, the acknowledgement of identity of senses across encounters or experiences, is fundamental to human perception. It is most prominently manifested in (human) vision, the fundamental faculty of which is to visually recognize our world in terms of objects, categories, and at a higher level, scenes and events. The inputs of visual recognition are then exploited to build concepts, which, though universally agreed as mental representations, differ widely in their modelling paradigm. The mainstream approach, termed Descriptionism in Millikan (2000), models concepts as classes described intensionally by their linguistically grounded properties, and drives phenomena such as (human) knowledge acquisition, reasoning and communication. An alternative approach termed Teleosemantics (Giunchiglia and Fumagalli 2016), inspired by the work of Ruth Garrett Millikan (Millikan 2000, Millikan 2004, Millikan 2005), was recently proposed which models concepts as abilities capable of implementing functions such as visual recognition and classification. This shift from modelling the means of static representation of the world to modelling the means of continual generation of such representations forms the basis for visual semantics (Giunchiglia, Erculiani and Passerini 2021), namely, the study of how concepts are generated from visual perception. In the context of visual semantics, there are two stratified problems which are of particular concern for object recognition. The first problem is the many-to-many mapping between what is the case in the world, i.e., substances, and the visual input perceived from such substances, i.e., objects (termed Sensory Gap in Smeulders et al. 2000). The second problem results from the many-to-many mapping between the information conveyed by the visual input and its contextual interpretation by a user depending on purpose or objective to be achieved (termed Semantic Gap in Smeulders et al. 2000). These two problems, we argue, are a consequence of the central assumption prevailing in existing computer vision systems that, in visual recognition, objects are hardly organized into classification hierarchies post visual perception. Furthermore, in the few studies they are organized as such (see, for instance, Marszalek and Schmid 2007; Deng et al. 2009), the classification is performed on linguistically grounded properties and not on visually perceived properties, thus generating the many-to-many mapping.

Paper

Similar papers

© 2026 NYSGPT2525 LLC