Humans quite frequently interact with conversational agents. The rapid\nadvancement in generative language modeling through neural networks has helped\nadvance the creation of intelligent conversational agents. Researchers\ntypically evaluate the output of their models through crowdsourced judgments,\nbut there are no established best practices for conducting such studies.\nMoreover, it is unclear if cognitive biases in decision-making are affecting\ncrowdsourced workers' judgments when they undertake these tasks. To\ninvestigate, we conducted a between-subjects study with 77 crowdsourced workers\nto understand the role of cognitive biases, specifically anchoring bias, when\nhumans are asked to evaluate the output of conversational agents. Our results\nprovide insight into how best to evaluate conversational agents. We find\nincreased consistency in ratings across two experimental conditions may be a\nresult of anchoring bias. We also determine that external factors such as time\nand prior experience in similar tasks have effects on inter-rater consistency.\n