Towards AI generated Philosophy through Autoformalization

AI4Humanities & Social Sciences

Author Maximilian Noichl

Affiliation Utrecht University

Venue & date AI 4 Science, Mainz
5 October 2026

Introduction

  • Doctoral candidate at Utrecht University in theoretical philosophy.
  • Data-driven philosophy, philosophy with AI, complex collective cognition, philosophy of science.
  • To follow along: maxnoichl.eu/talk
  • Towards AI generated Philosophy through Autoformalization

Why would we want this?

  • Better (AI-assisted) philosophy?
  • Understand better how we do philosophy!
  • Understanding the differences between AI-generated and human philosophy!
    • Alignment purposes?

How do we get good AI-generated philosophy?

  • Out-of-the-box AI-generated philosophy tends to be cliché/slop/repetitive…
  • AI tends to shine in “verifiable” domains (coding & math), which provide signal for ‘hillclimbing’.
  • How can we phrase philosophy that way?

Philosophical theories

What is knowledge?

  • Let’s say: Knowledge = justified true belief (JTB).
    • As a formula: knows ⇔ true ∧ believes ∧ justified
  • Gettier cases: something is wrong with justification alone.
    • Russell’s clock: true, believed & justified → knows, but we disagree!
  • How about: Knowledge = JTB + no false lemmas.
  • Philosophy, under this view: finding the simplest principles that match our conceptualizations.
Painting of a long-case clock in an empty, flooded room, with a bentwood chair and a paper boat
Russell's stopped clock: true and justified, but not knowledge.

Our workflow

  1. Research the philosophical landscape: theories & desiderata (OpenAI Deep Research or Codex + database).
  2. Evaluate theories on desiderata using independent formalization (Z3). Heatmap
  3. But: “Monster-theories”!
    • Force a simplicity constraint.
  4. Pareto frontier: simpler theories that solve more problems. Pareto
  5. Run research flow against the criteria: come up with a theory (verbalized sampling), get feedback, revise. Workflow
  6. Optimize along the frontier. Pareto run

Early results

  • For 10 domains (Knowledge, Scientific Explanation, Moral Responsibility, etc.): 26 of 95 literature theories are at the frontier (27% vs. 19% under chance). Domains
  • Does it work?
  • 52 of 100 AI-proposals make it past the frontier. Domains run
  • Some results are interesting, but we still get a lot of “cheating”; literature theories are under-optimized.

Problems & future work

  • Lack of “deep creativity” (dependent on landscape → counterexample mining)
  • Evaluation is hard (but cf. PhilosophyBench; AI Philosophy Competition)
  • No tieback to literature (novelty)
  • Understanding our results? Back to essays?
  • Thank you!

Literature

Clark, Michael. 1963. “Knowledge and Grounds: A Comment on Mr. Gettier’s Paper.” Analysis 24 (2): 46–48. https://doi.org/10.1093/analys/24.2.46.
Gettier, Edmund L. 1963. “Is Justified True Belief Knowledge?” Analysis 23 (6): 121–23. https://doi.org/10.1093/analys/23.6.121.
Goodsell, Zachary, and Elliott Thornley. 2026. “AI Philosophy Competition: 1st Edition.” September 1, 2026. https://www.zacharygoodsell.com/ai-philosophy-competition.
Guo, Daya, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, et al. 2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs Through Reinforcement Learning.” Nature 645 (8081): 633–38. https://doi.org/10.1038/s41586-025-09422-z.
Kirk, Robert, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2024. “Understanding the Effects of RLHF on LLM Generalisation and Diversity.” In International Conference on Learning Representations (ICLR 2024). https://arxiv.org/abs/2310.06452.
PhilosophyBench. 2026. “PhilosophyBench: AI and Philosophy Research at Stanford.” September 2026. https://philosophybench.org/.
Russell, Bertrand. 1948. Human Knowledge: Its Scope and Limits. London: George Allen & Unwin.
Wu, Yuhuai, Albert Q. Jiang, Wenda Li, Markus N. Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy. 2022. “Autoformalization with Large Language Models.” In Advances in Neural Information Processing Systems 35 (NeurIPS 2022). https://arxiv.org/abs/2205.12615.
Zhang, Jiayi, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, and Weiyan Shi. 2026. “Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.” In International Conference on Machine Learning (ICML 2026). https://arxiv.org/abs/2510.01171.