LLM Evaluation as Sociotechnical Practice: Experiences from a Research-Practice Partnership in Community Health

As large language models (LLMs) are increasingly deployed in sensitive health domains, evaluation is often framed as a technical problem of measuring accuracy, safety, or bias. Drawing on our experiences in a partnership between HCI researchers and a community health organization in India, we instead examine evaluation as an ongoing sociotechnical practice. We present a qualitative analysis of how we collaboratively developed and implemented a human-centered evaluation framework for an LLM-based sexual and reproductive health (SRH) chatbot. Through interviews with program and technical staff at the health organization, analysis of internal evaluation artifacts, and reflections from the research team, our findings show how designing and doing evaluation is shaped by organizational goals, infrastructural constraints, and the integration of diverse forms of expertise, including medical, social, cultural, and experiential knowledge. We show how research–practice stakeholders, motivated by care and responsibility towards the communities they serve, iteratively tinkered with metrics, roles, and workflows in response to tensions around medical accuracy, linguistic accessibility, cultural appropriateness, and user trust. We draw on our findings to discuss implications for taking a process-oriented approach to LLM evaluation in community health settings, and for strengthening evaluation efforts through research-practice partnerships.

Paper

Full text

PDF

LLM Evaluation as Sociotechnical Practice: Experiences from a Research-Practice Partnership in Community Health

Semantic Scholar · Sociology · 2026

Abstract

As large language models (LLMs) are increasingly deployed in sensitive health domains, evaluation is often framed as a technical problem of measuring accuracy, safety, or bias. Drawing on our experiences in a partnership between HCI researchers and a community health organization in India, we instead examine evaluation as an ongoing sociotechnical practice. We present a qualitative analysis of how we collaboratively developed and implemented a human-centered evaluation framework for an LLM-based sexual and reproductive health (SRH) chatbot. Through interviews with program and technical staff at the health organization, analysis of internal evaluation artifacts, and reflections from the research team, our findings show how designing and doing evaluation is shaped by organizational goals, infrastructural constraints, and the integration of diverse forms of expertise, including medical, social, cultural, and experiential knowledge. We show how research–practice stakeholders, motivated by care and responsibility towards the communities they serve, iteratively tinkered with metrics, roles, and workflows in response to tensions around medical accuracy, linguistic accessibility, cultural appropriateness, and user trust. We draw on our findings to discuss implications for taking a process-oriented approach to LLM evaluation in community health settings, and for strengthening evaluation efforts through research-practice partnerships.

References (55)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC