Evaluating the Effectiveness of Large Language Models for Livestock and Climate-Related Agricultural Advice in Ontario
Recent advancements in artificial intelligence (AI) and large language models (LLMs) offer new opportunities for improving agricultural extension, particularly in communicating livestock and climate-related knowledge. While general-purpose models like ChatGPT have demonstrated potential, their performance in specialized domains such as animal welfare has yet to be fully assessed. Prior studies suggest that domain-specific models outperform general ones on precision and contextual accuracy, yet comparative evaluations with expert-curated content are limited. This study examines the performance of ChatGPT, Claude, and Gemini in answering livestock-related questions relevant to climate change. It evaluates the degree of alignment between AI-generated and expert-developed answers, focusing on five metrics: similarity, faithfulness, context precision, context recall, and answer relevancy. Ten expert-reviewed questions were developed, and corresponding human-curated answers were constructed from recent literature. Responses from the three AI models were collected using standardized prompts. Answers were evaluated using a multi-criteria framework supported by qualitative coding and statistical summaries. Prompt engineering was applied to improve answer quality and comparability across models. AI models—especially ChatGPT and Claude—showed high alignment with expert answers. Their outputs demonstrated strong similarity, faithfulness, and context relevance. While some variation in depth and specificity remained, the overall quality of AI responses was high across most metrics. LLMs show promise for supporting agricultural extension and public knowledge transfer. Ensuring reliability requires continued use of expert oversight, domain-specific data, and refined prompting strategies.
Paper
Full text
Evaluating the Effectiveness of Large Language Models for Livestock and Climate-Related Agricultural Advice in Ontario
Semantic Scholar · 2025
Abstract
Recent advancements in artificial intelligence (AI) and large language models (LLMs) offer new opportunities for improving agricultural extension, particularly in communicating livestock and climate-related knowledge. While general-purpose models like ChatGPT have demonstrated potential, their performance in specialized domains such as animal welfare has yet to be fully assessed. Prior studies suggest that domain-specific models outperform general ones on precision and contextual accuracy, yet comparative evaluations with expert-curated content are limited. This study examines the performance of ChatGPT, Claude, and Gemini in answering livestock-related questions relevant to climate change. It evaluates the degree of alignment between AI-generated and expert-developed answers, focusing on five metrics: similarity, faithfulness, context precision, context recall, and answer relevancy. Ten expert-reviewed questions were developed, and corresponding human-curated answers were constructed from recent literature. Responses from the three AI models were collected using standardized prompts. Answers were evaluated using a multi-criteria framework supported by qualitative coding and statistical summaries. Prompt engineering was applied to improve answer quality and comparability across models. AI models—especially ChatGPT and Claude—showed high alignment with expert answers. Their outputs demonstrated strong similarity, faithfulness, and context relevance. While some variation in depth and specificity remained, the overall quality of AI responses was high across most metrics. LLMs show promise for supporting agricultural extension and public knowledge transfer. Ensuring reliability requires continued use of expert oversight, domain-specific data, and refined prompting strategies.