Language Models as Critical Thinking Tools: A Case Study of Philosophers

Current work in language models (LMs) helps us speed up or even skip thinking by accelerating and automating cognitive work. But can LMs help us with critical thinking -- thinking in deeper, more reflective ways which challenge assumptions, clarify ideas, and engineer new concepts? We treat philosophy as a case study in critical thinking, and interview 21 professional philosophers about how they engage in critical thinking and on their experiences with LMs. We find that philosophers do not find LMs to be useful because they lack a sense of selfhood (memory, beliefs, consistency) and initiative (curiosity, proactivity). We propose the selfhood-initiative model for critical thinking tools to characterize this gap. Using the model, we formulate three roles LMs could play as critical thinking tools: the Interlocutor, the Monitor, and the Respondent. We hope that our work inspires LM researchers to further develop LMs as critical thinking tools and philosophers and other 'critical thinkers' to imagine intellectually substantive uses of LMs.

Paper

Similar papers

Reviewer veW65/10 · confidence 3/52024-04-29

Summary

While often LLMs are employed to quickly retrieve standard knowledge and arguments, this paper discusses the question to what extent LLMs can serve as tools for critical thinking in the sense of questioning, reorienting, analysing, and developing ideas. As a case study, the paper reports on interviews with 21 professional philosophers about how they engage in critical thinking and on their experiences with LMs. In view of the findings that LLMs are considered to fail to be useful in open, undetermined contexts, and also do not allow intellectual interaction driven by curiosity, the authors model conversations with LLMs along two dimensions – selfhood and initiative. Based on the resulting matrix of high and low selfhood and initiative, they propose three roles of LLMs for philosophy: the Interlocutor (high-selfhood, high-initiative), the Monitor (low-selfhood, high-initiative), and the Respondent (high-selfhood, low-initiative). In conclusion, they take up some suggestions of where LLMs might be improved (e.g. reasoning about “uncommon” sense, or improving their long-range planning), and advocate the use of LLMs in a playful way despite their insufficiencies.

Rating

5

Confidence

3

Ethics flag

1

Reasons to accept

The paper presents an interesting and novel classification of interactions with LLM based chatbots. While the findings of the interviews are not surprising, this classification helps to better understand the reasons behind the interviewed philosophers’ assessments.

Reasons to reject

The paper leaves out important details: (1) By which criteria have the interviews philosophers been selected? Has it been ensured that the interviewed have experience with LLMs? (2) Which LMs and Chatbots have been used? (3) The claims are very general and do not take into account possible modifications of the experimental set-up. In particular, the noted insufficiencies may be a result of just the way how standard LLMs have been set up. An alternative might be databases of pro- and con-arguments that could be added by fine tuning or retrieval augmented generation (RAG) on such databases. (4) From the philosopher’s side, the conclusion also appears too general, missing out differences in the intended gain of knowledge in the philosophical traditions and methods.

Questions to authors

(1) By which criteria have the interviews philosophers been selected? Has it been ensured that the interviewed have experience with LLMs? (2) Which LMs and Chatbots have been used? (3) Have you been considering possible modifications of the experimental set-up?

Reviewer NJTN10/10 · confidence 5/52024-05-04

Summary

This paper explores the potential for language models to serve as tools for critical thinking, using philosophy as a case study. The authors engage with prior work and conduct a sound qualitative study by interviewing 21 professional philosophers, providing valuable empirical insights into how philosophers think and their attitudes toward language models. The work demonstrates originality by examining language models for critical thinking rather than just text generation/editing assistance. The proposed "selfhood-initiative" model offers a practical conceptual framework, and the envisioned roles for language models like the Interlocutor, Monitor, and Respondent represent creative new possibilities. Significantly, the paper highlights significant gaps in current language models for supporting deeper reasoning required for critical thinking and articulates an agenda for developing better-suited models, which could have wide-ranging impacts. The discussion around language models shaping metaphilosophical questions and the discipline of philosophy itself is highly relevant. Overall, this is a remarkable work that makes a valuable contribution by studying an underexplored area at the intersection of AI and philosophy through a combination of good design, conceptual modeling, and creative proposals for future development.

Rating

10

Confidence

5

Ethics flag

1

Reasons to accept

This is a high-quality paper. The authors know their topic well. The writing and organization are very impressive. The key questions and findings are helpful to the community, especially when there is a huge amount of anthropomorphizing of LLMs. LLMs are good at benchmarks, but that sort of reasoning often doesn't imply critical reflective thinking, and then there is the classic case of LLMs being servile, passive, and incurious (thanks to RLHF). This comes time and again in creative writing where writers feel that RLHF limits LLMs in exploring darker difficult topics. I appreciated the design and recommendations for what makes a better critical thinking tool. The SelfHood Initiative model is pretty exciting and will likely have a good impact on interdisciplinary researchers who work with LLMs I would fight for this paper's acceptance.

Reasons to reject

No reason to reject

Questions to authors

Did you try Claude? I have heard it's better than GPT4.

Reviewer Ynvx3/10 · confidence 4/52024-05-11

Summary

This paper is interested in the use of LLMs to assist humans in critical thinking. Academic philosophers, in interviews which are qualitatively coded, are asked to reflect on the use of LLMs to assist them in their work. Overall, the interviewed philosophers find that LLMs are not useful. The paper diagnoses this uselessness in two missing properties in LLMs: they lack a sense of selfhood and initiative. Possible roles for LLMs that vary these features (i.e., having selfhood and lacking initiative, lacking selfhood and having initiative, and having both) are discussed as a way of trying to guide the development of LLMs.

Rating

3

Confidence

4

Ethics flag

1

Reasons to accept

I found the discussion of the roles for LLMs (interlocuter, monitor, and respondent) to be interesting. I think this framing is interesting to consider in an educational context (e.g., how should universities position themselves with relation to AI assistants). Further the paper is well written and does a solid job of situating the current use cases of LLMs in contrast with their use in fostering critical thinking.

Reasons to reject

The link between the paper’s aim (positioning LLMs as critical thinking tools) and the argumentation and methods (the interview of academic philosophers) is not clear. If the aim is to study what makes a tool useful for supporting critical thinking, why are philosopher’s views on LMs (e.g., answers to questions like “What are some risks and weaknesses for language models in philosophy”) relevant for addressing this? At times the paper discusses how critical thinking is deployed in endeavors like “history, science, and philosophy”, wouldn’t it make sense to probe how each of these fields utilizes critical thinking and what tools are useful for doing this more efficiently? Further, the discussion of the things that are useful for philosophers in critical thinking, selfhood and initiative, both require agency (in my understanding) in the definitions given in the paper (“selfhood is a resource’s ability to have certain locally persistent internal states (such as perspectives, beliefs, opinions, memory) and to consistently use them as the basis for judgements” and “Initiative is a resource’s ability to set its own intentions and goals, possibly different from its user’s, and to execute actions oriented towards those intentions.”). This feels not concrete enough to offer practical guidance in developing LLMs.

Questions to authors

Who is the “we” in “we claim all the rights to think?”. I read this as saying we humans claim all the rights to think. What relationship does that have to LLMs? It feels to me, in fact, like that statement is in tension with the aims of the paper (the description of a tool to automate some types of thinking). Perhaps, we will use LLMs to think more critically? But the other uses cases (e.g., coding, generating emails) mentioned in the beginning are taken as useful because they alleviate the need for that type of thinking. Two issues with LLMs, that they are “highly neutral, detached, and non-judgmental” and “servile, passive, and incurious” are mentioned a few times in the paper, however, no empirical results are given to support this. If this is the perception of the interviewees that is fine, but it should be made clear in the paper.

Reviewer ttPJ7/10 · confidence 3/52024-05-15

Summary

The paper discusses the use (potential and actual) of LMs by philosophers as critical thinking tools. From interviews with philosophers, the authors determine two broad groups of reasons why philosophers find current LMs problematic as thinking tools. Based on this, they propose a basic framework to characterise such tools, and suggest how LM outputs and behaviour might be designed to give better support for this kind of use.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

This is a thought-provoking and original piece of work: it's useful to see discussion of the interaction design aspects of LMs, and serious consideration of how they might support (or not) different kinds of end use.

Reasons to reject

The findings about reasons why LMs are not much used in this context are interesting and seem solid, but the analytical framework proposed to explain this and used as the basis for suggestions does not seem to be so clearly justified; and the suggestions are fairly open challenges without much in the way of suggestions as to how (or whether) they could be implemented or addressed.

Questions to authors

The two-dimensional self/initiative model for critical thinking tools is proposed without a great deal of justification. While I can see that it links to the two main themes identified from the analysis of the interviews, I wasn't clear if it has any precedent or builds on / relates to other existing models of critical thinking or of creative assistive technology. There is plenty of prior work in designing tools to assist creative thinkers (artists, designers etc - I know of work dating back at least to Shneiderman (1997)'s Genex framework but I'm no expert in that area so I'm sure there's more), and there seems to be work in categorising critical thinking according to various dimensions. Does this model relate to previous work, and if so how? This would help justify this choice.

Reviewer veW62024-06-04

Dear authors, thank you for your replies! I still think that your paper addresses two topics that are not so closely related as you presuppose. While I do see the merit of your contribution in the introduction of a novel classification of interactions with LLM based chatbots, I still have reservations on the methodology and experimental set-up of your survey of how philosophers interact with LLMs. Therefore, I'll keep my score.

Reviewer Ynvx2024-06-05

Thanks for the engagement with the review. [Q1] Thank you for the clarification. [Q2] I think the evaluation lacked some detail. It is unclear how often the interviewed people used LLMs, in what domains they used them, and which ones they used. It does not seem that any of the questions in A.1 ask for how participants use or have used LLMs for critical thinking (the responses to Reviewer veW6 further add confusion to this. In what ways were interviewees new to LLMs interacting with LLMs, for example?). It would be interesting to get a sense of this and how it relates to their responses, for example. [RR1] Thanks for the clarification. [RR2] I don't see how the definitions given do not require agency, given the quoted definitions from the paper I included in the review. Overall, I still find the paper flawed and the main argumentation unclear. If the aim is to do a study like those in HCI work, the survey is lacking details about the users and their experiences which limit the inferences we can draw from their study. I will keep my score.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC