In contemporary society, the issue of psychological health has become increasingly prominent, characterized by the diversification, complexity, and universality of mental disorders. Cognitive Behavioral Therapy (CBT), currently the most influential and clinically effective psychological treatment method with no side effects, has limited coverage and poor quality in most countries. In recent years, researches on the recognition and intervention of emotional disorders using large language models (LLMs) have been validated, providing new possibilities for psychological assistance therapy. However, are large language models truly possible to conduct cognitive behavioral therapy? Many concerns have been raised by mental health experts regarding the use of LLMs for therapy. Seeking to answer this question, we collected real CBT corpus from online video websites, designed and conducted a targeted automatic evaluation framework involving three aspects, namely the evaluation of emotion tendency of generated text, structured dialogue pattern and proactive inquiry ability. Considering limited CBT-related texts in a general chat LLM’s training corpus, we evaluated the CBT ability of the LLM after integrating a CBT knowledge base to explore the influence of introducing additional knowledge. Four LLM variants with exceptional performance are evaluated, and the experimental result shows the great potential of LLMs in psychological counseling realm, especially after combining with other technological means.