Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5

arXiv:2510.07024v3 Announce Type: replace-cross
Abstract: LLMs are remarkable artifacts that have revolutionized a range of knowledge-intensive tasks. A significant contributor is their factual knowledge, which, to date, remains poorly understood, and is usually analyzed from biased samples. In this paper, we provide a framework and the results of a multi-dimensional analysis of GPTKB v1.5 (Hu et al., 2025a), a recursively elicited Knowledge Base (KB) of 100 million facts (or beliefs) of a frontier LLM, namely, GPT-4.1. Given the scale of the elicited facts, we provide a multi-dimensional approach to qualitatively and quantitatively analyze these facts as opposed to the mainstream fact completion benchmarks, which are prone to availability bias. We find that the models' factual knowledge differs quite significantly from established knowledge bases, and that its accuracy is significantly lower than indicated by previous benchmarks. We also find that inconsistency, ambiguity and hallucinations are major issues, shedding light on future research opportunities in neuro-symbolic AI concerning extraction, consolidation and verification of factual LLM knowledge.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: