We evaluate the zero-shot performance of large language models (LLMs), both reasoning-based and non-reasoning, on financial sentiment classification using the expert-annotated Financial PhraseBank dataset. We compare three proprietary LLMs (GPT-4o, GPT-4.1, o3-mini) under different prompting paradigms that imitate System 1 (intuitive) and System 2 (deliberative) reasoning, alongside two finetuned baselines (FinBERT-Prosus, FinBERT-Tone). We demonstrate that reasoning, whether prompt-induced or built-in, does not improve alignment with human sentiment labels. Notably, GPT-4o without Chain-of-Thought (CoT) prompting achieves the best performance across varying levels of linguistic complexity and annotation ambiguity, suggesting that intuitive, System 1-style prompting better matches human judgment in financial contexts. These findings challenge the assumption that more reasoning universally improves LLM performance across tasks.