Agentic LLMs for Argument Mining in Philosophical Texts

Logics and Argumentation for New-Generation AI

Authors Maximilian Noichl & Jan Broersen

Affiliation Theoretical Philosophy, Utrecht University, The Netherlands

Venue & date CSIC, Barcelona
14 September 2026

Overview

  • Research question and motivation
  • Argumentation in philosophy
  • Choosing argumentation frameworks
  • Pilot study & results

Research Questions

  • Can we use LLMs to analyse and map out the argumentative structures of philosophical texts, leveraging the symbolic knowledge captured by formal argumentation frameworks?
  • How can such a system work, and how do we validate it?

Motivation

  • A tool for students and researchers to get to the core of a paper more quickly
  • Validity and ‘completeness’ checking of arguments in philosophical texts
  • A metric for further exploration of the body of knowledge that forms philosophy
  • Answering questions on differences between areas in philosophy
  • Exploring if philosophical argumentation is in some sense ‘special’
  • A building block for AI-generated philosophy

What does a formal structure add?

  • An exact format (a formal language)
  • A level of abstraction
  • A formal semantics (an exact theory, in terms of the formal language, on what arguments are valid)

Criteria for choosing a formal structure

  • Do we, apart from a formal syntax, require a formal semantics?
  • Are positions/statements (targets or starting points for defence or attack) atomic or structured?
  • Do we allow for hierarchy?
  • Does a scheme only have attack relations or also support relations? Do we distinguish between different kinds of them?

The structure of argumentation

  • Dung: abstract (only arguments and attacks)
  • ASPIC+: structured (premisses, different kinds of support, etc.) Attack types
  • Carneades: statements and arguments
  • Free: just suggest a graph structure.

Pilot study: Data

  • Corpus: Mind Design III — classic texts at the interface of philosophy, psychology, and AI: Searle, Dennett, Churchland, Fodor, …
  • 21 articles, median length 32 pages.
  • Excluded: Turing’s Computing Machinery and Intelligence: The imitation game triggers GPT-5.4’s safety features.

Pilot study: Workflow

  • Agentic loop, not single-pass: arguments are distributed nonlinearly across philosophical texts.
  • GPT-5.4 in OpenAI’s Codex CLI: tool use, repeated text lookups, iterative work.
    • Pass 1: work through the whole text (Markdown) & annotate argument passages.
    • Pass 2: build a JSON argument graph, according to guidelines — generic, Dung- & Carneades-style.
  • 63 graphs (21 papers × 3 schemes), laid out with Graphviz.

01 / Generic
Generic argument graph of Searle 1980
02 / Dung
Dung-style argument graph of Searle 1980
03 / Carneades
Carneades-style argument graph of Searle 1980
Searle (1980) · scroll to zoom · drag to pan · double-click to reset

Pilot study: Validation & Results

  • Human validation: 46 of 50 sampled links appear fully justified.
  • LLM-assisted audit (Opus 4.6): no critical faithfulness deviations in any scheme; 1–2 major deviations at the median. Validation
  • But: Carneades is hard to follow: sig. more guideline deviations.
  • WIP: LLM audit catches introduced errors.
  • Graph structure tracks the schemes: generic → fewest components; Dung → denser, shallower; Carneades → deeper, bipartite. Graph features

Conclusion & further work

  • Parsing complex arguments with agentic LLMs appears viable.
  • But much remains to be done!
  • Evaluation at high complexity is tough & expensive.
  • What are the optimal argumentation frameworks & workflows?
  • How to integrate solvers & verifiers?

Literature

Cayrol, Claudette, and Marie-Christine Lagasquie-Schiex. 2005. “On the Acceptability of Arguments in Bipolar Argumentation Frameworks.” In Symbolic and Quantitative Approaches to Reasoning with Uncertainty, edited by Lluís Godo, 3571:378–89. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. https://doi.org/10.1007/11518655_33.
Chen, Guizhen, Liying Cheng, Anh Tuan Luu, and Lidong Bing. 2024. “Exploring the Potential of Large Language Models in Computational Argumentation.” In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2309–30. Bangkok, Thailand: Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.126.
Cohen, Andrea, Sebastian Gottifredi, Alejandro J. García, and Guillermo R. Simari. 2014. “A Survey of Different Approaches to Support in Argumentation Systems.” The Knowledge Engineering Review 29 (5): 513–50. https://doi.org/10.1017/S0269888913000323.
Dung, Phan Minh. 1995. “On the Acceptability of Arguments and Its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and N-Person Games.” Artificial Intelligence 77 (2): 321–57. https://doi.org/10.1016/0004-3702(94)00041-X.
Gordon, Thomas F., Henry Prakken, and Douglas N. Walton. 2007. “The Carneades Model of Argument and Burden of Proof.” Artificial Intelligence 171 (10–15): 875–96.
Haugeland, John, Carl F. Craver, and Colin Klein, eds. 2023. Mind Design III: Philosophy, Psychology, and Artificial Intelligence. New edition. Cambridge, Massachusetts London, England: The MIT Press.
Hintikka, Jaakko. 1970. “Information, Deduction, and the a Priori.” Noûs 4 (2): 135–52. https://doi.org/10.2307/2214318.
Li, Hao, Yizheng Sun, Viktor Schlegel, Riza Theresa Batista-Navarro, and Goran Nenadic. 2025. “Large Language Models in Argument Mining: A Survey.” https://arxiv.org/abs/2506.16383.
Lippi, Marco, and Paolo Torroni. 2016. “Argumentation Mining: State of the Art and Emerging Trends.” ACM Transactions on Internet Technology 16 (2): 1–25. https://doi.org/10.1145/2850417.
Malaterre, Christophe, and François Lareau. 2022. “The Early Days of Contemporary Philosophy of Science: Novel Insights from Machine Translation and Topic-Modeling of Non-Parallel Multilingual Corpora.” Synthese 200: 242. https://doi.org/10.1007/s11229-022-03722-x.
Modgil, Sanjay, and Henry Prakken. 2014. “The ASPIC+ Framework for Structured Argumentation: A Tutorial.” Argument and Computation 5 (1): 31–62. https://doi.org/10.1080/19462166.2013.869766.
Noichl, Maximilian. 2021. “Modeling the Structure of Recent Philosophy.” Synthese 198 (6): 5089–5100. https://doi.org/10.1007/s11229-019-02390-8.
Peldszus, Andreas, and Manfred Stede. 2013. “From Argument Diagrams to Argumentation Mining in Texts: A Survey.” International Journal of Cognitive Informatics and Natural Intelligence 7 (1): 1–31. https://doi.org/10.4018/jcini.2013010101.
Petrovich, Eugenio. 2024. A Quantitative Portrait of Analytic Philosophy: Looking Through the Margins. Cham: Springer. https://doi.org/10.1007/978-3-031-53200-9.
Ruiz-Dolz, Ramon, Zlata Kikteva, and John Lawrence. 2025. “Mining Complex Patterns of Argumentative Reasoning in Natural Language Dialogue.” In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7421–35. Vienna, Austria: Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.368.
Searle, John R. 1980. “Minds, Brains, and Programs.” Behavioral and Brain Sciences 3 (3): 417–24. https://doi.org/10.1017/S0140525X00005756.