The Trouble with ‘Seeing’:Two Ethnomethodological Challenges with AI Scenic Descriptions

With the development of computer vision technology, multimodal conversational systems are now being built with the explicit goal of providing spoken ‘descriptions’ of their users’ surroundings (Chang et al., 2024; Cheema et al., 2025), at any moment, in response to these users’ requests (Chang et al., 2025; Čujdíková et al., 2024). We explore two ethnomethodologically significant challenges encountered by computer scientists in producing such systems:<br/><br/>1. The et cetera principle, notably investigated by Sacks (1963) and Garfinkel (1967). In computer vision, this principle is treated as a problem: namely, the impossibility of exhaustively formulating a situation in-so-many-words (Garfinkel &amp; Sacks, 1970; Waismann, 1951)—which can be elaborated indefinitely—and, consequently, the fact that ‘descriptions’ (their length, level of detail, etc.; Stangl et al., 2021) are only relevant in relation to participants’ practical purposes in interaction.<br/><br/>2. The sequential placement of so-called ‘descriptions’ within the temporal order of users’ courses of action. Descriptions, accounts, glosses, formulations, etc., emerge as such (i.e., as practical resources for participants) in being finely timed in relation to a specific ecology of action (Goodwin, 2018; Mondada, 2007).<br/><br/>We ground this analysis in a heuristic contrast between excerpts of blind users interacting, on the one hand, with sighted humans and, on the other hand, with artificial ‘assistants’. Through this comparison, we examine the conditions under which these co-participants’ contributions emerge as resources in relation to blind participants’ ongoing activities (e.g., crossing a road), and when they instead remain “detached” (Dreyfus, 1990) accounts, removed from the rapidly shifting temporal order of these participants’ embodied conduct. Notably, we detail how, in being accountably produced ‘too early’ or ‘too late’ relative to the practical expectancies of an unfolding activity, current multimodal agents’ contributions come to be decoupled from the “vivid present” (Au-Yeung &amp; Fitzgerald, 2023; Garfinkel, 2006) of human participants.<br/><br/><br/><br/>References<br/><br/>Au-Yeung, T. S. H., &amp; Fitzgerald, R. (2023). Time structures in ethnomethodological and conversation analysis studies of practical activity. The Sociological Review, 71(1), 221–242. https://doi.org/10.1177/00380261221103018<br/>Chang, R.-C., Liu, Y., &amp; Guo, A. (2024). WorldScribe: Towards Context-Aware Live Visual Descriptions. Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 1–18. https://doi.org/10.1145/3654777.3676375<br/>Chang, R.-C., Natalie, R., Xu, W., Yap, J. Z. F., &amp; Guo, A. (2025). Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired. Proceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility. ASSETS ’25, Denver, Colorado, USA. https://doi.org/10.1145/3663547.3746319<br/>Cheema, M., Seifi, H., &amp; Fazli, P. (2025). Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals. Proceedings of the 2025 ACM Designing Interactive Systems Conference, 458–474. https://doi.org/10.1145/3715336.3735685<br/>Čujdíková, M., Stankovičová, M., &amp; Vankúš, P. (2024). How The Blind See Chatgpt And Gemini: A Case Study From A Summer School For Visually Impaired Students. 7507–7512. https://doi.org/10.21125/iceri.2024.1813<br/>Dreyfus, H. L. (1990). Being-in-the-World: A Commentary on Heidegger’s Being in Time, Division I (Vol. 102, Issue 2, pp. 290–293). Bradford.<br/>Garfinkel, H. (1967). Studies in Ethnomethodology. Polity Press.<br/>Garfinkel, H. (2006). Seeing sociologically: The routine grounds of social action. Paradigm.<br/>Garfinkel, H., &amp; Sacks, H. (1970). On formal structures of practical action. In J. C. McKinney &amp; E. A. Tiryakian (Eds.), Theoretical Sociology: Perspectives and Developments (pp. 338–366). Appleton-Century-Crofts.<br/>Goodwin, C. (2018). Why Multimodality? Why Co-Operative Action? (transcribed by J. Philipsen). Social Interaction. Video-Based Studies of Human Sociality, 1(2). https://doi.org/10.7146/si.v1i2.110039<br/>Mondada, L. (2007). Multimodal resources for turn-taking: Pointing and the emergence of possible next speakers. Discourse Studies, 9(2), 194–225. https://doi.org/10.1177/1461445607075346<br/>Sacks, H. (1963). Sociological description. Berkeley Journal of Sociology, 8, 1–16.<br/>Stangl, A., Verma, N., Fleischmann, K. R., Morris, M. R., &amp; Gurari, D. (2021). Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision. Proceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility. https://doi.org/10.1145/3441852.3471233<br/>Waismann, F. (1951). Verifiability. In G. Ryle &amp; A. Flew (Eds.), Logic And Language (pp. 35--68). Blackwell.<br/><br/>

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC