Prompt Refinement or Fine-tuning? Best Practices for using LLMs in Computational Social Science Tasks

Large Language Models are expressive tools that enable complex tasks of text understanding within Computational Social Science. Their versatility, while beneficial, poses a barrier for establishing standardized best practices within the field. To bring clarity on the values of different strategies, we present an overview of the performance of modern LLM-based classification methods on a benchmark of 23 social knowledge tasks. Our results point to three best practices: select models with larger vocabulary and pre-training corpora; avoid simple zero-shot in favor of AI-enhanced prompting; fine-tune on task-specific data, and consider more complex forms instruction-tuning on multiple datasets only when only training data is more abundant.

Paper

References (14)

08Autoprompt: Eliciting knowledge from language models with automatically generated prompts2020 · ArXiv
09How to fine-tune bert for text classification?2019
102024. SOCIALITE-LLAMA: An Instruction-Tuned Model for Social Scientific TasksArXiv
112024. RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language Processing
122024. The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks

Scroll for more · 2 remaining

Similar papers

© 2026 NYSGPT2525 LLC