In November, Ars Inquirendi convenes over thirty scholars from around the world to answer two questions: In the era of large language models, what can we now discover about the pre-print past that we could not before, and what would justify believing it? Join us at St Edmund Hall, Oxford, and online, 20–22 November, to explore the rapid advances in what models can do and how scholars use them. The conference moves from charting the evidence of the pre-print world and considering how we make it available to LLMs that have so far seen only a fraction of it, as machines learn with startling speed to interpret our artefacts en masse, through historical inquiry at machine scale, to scholarly sovereignty over the new technology.
All talks except the live keynotes and workshops are pre-recorded and released a week ahead, leaving the live days for discussion.
Keynotes Marieke Meelen (University of Cambridge) · Peter Turchin (Complexity Science Hub Vienna & University of Oxford)
The pre-print graphosphere – everything written, daubed, etched, carved or otherwise laden with semantic freight in a culture before movable-type printing spread through it, along with all its interrelations, the vast majority of it now lost and the rest surviving unevenly – and indeed all other records and evidence: come and hear about the unprecedented power that LLMs are giving us to explore and depict that totality – especially now that machines have begun to interpret manuscript images and other artefacts at scale – and the attendant challenges.
Artificial Intelligence and Large Language Models in particular appear to rapidly change the modern world, but how can we use them to unlock the past of pre-print cultures? In this talk I'll show how and when LLMs can be useful in the process of digitising handwritten manuscripts and turning them into searchable corpora. I'll argue that when these corpora are well-annotated in particular, they are unique resources not only for linguists, but also for researchers in the fields of philology, history, literature, religion or other social sciences and humanities. Using examples from an extremely difficult body of historical Tibetan texts, I will furthermore show the limits of LLMs and where exactly in the pipeline from manuscript to answering research questions other computational methods should be considered as well.
In the Regional Great Russian Dictionary of 1852, красота ('beauty') denotes a bride's ribbon, placed by the priest into the Gospel book during the wedding rite. In the Suprasl Codex of the 10th–11th centuries, its root renders Greek κόσμος 'ornament, order, universe'. The familiar aesthetic sense—'attractiveness'—is a late and derived development. Historical lexicography preserves this stratigraphy; large language models, trained overwhelmingly on post-print text, may not. This pilot study tests whether LLMs can read pre-print-era dictionary definitions without projecting modern semantics backwards.
The material comprises entries for красота in eleven Church Slavonic and historical Russian dictionaries documenting the language of the pre-print era—from the earliest Slavonic manuscripts of the 10th–11th centuries onwards—in editions from the eighteenth to the twenty-first century. The pre-print past becomes machine-readable only through this later lexicographic mediation, which is precisely where models may substitute their own training distribution for the historical record. A completed manual componential and Greek–Slavonic analysis of these sources serves as the gold standard: humanistic scholarship as the evaluation benchmark for AI. Definitions are presented to frontier models under three conditions—unattributed, source-attributed, and with an explicit question about the relation to modern meaning—separating inference from the text, recognition of the source, and knowledge of later semantics. Errors are classified as epistemic upgrades, anachronisms, fabricated components, fabricated Greek glosses, and invented continuity narratives.
Two single-model series (42 responses) show fabrication migrating to the continuity condition: the model refused to reproduce unseen dictionary content in direct probes, read the authentic regional label correctly, yet invented two contradictory expansions for its corrupted variant, produced an anachronistic attribution, and inserted a nonexistent word into a Septuagint quote. The paper proposes an anti-fabrication protocol for LLM-assisted work with historical lexicography—a contribution to defending pre-print scholarship against dangerously plausible AI output.
This paper presents an experiment in using a commercially available frontier LLM to extract, normalize, and critically verify Slavic toponymic evidence for Livonia and its adjoining eastern Baltic borderlands before 1700. The corpus belongs to a transitional manuscript and early-print ecology, in which Latvian and Estonian vernacular textual traditions remained comparatively sparse while German, Latin, Polish, Ruthenian, and Russian documentary practices intersected.
The Slavic evidence comes principally from three early modern traditions—Middle Russian, Polish, and Ruthenian—with a smaller Old East Slavic chronicle layer. Place-names survive in chronicles, legal acts, boundary descriptions, maps and atlases, and later editions of manuscript material. The workflow combines candidate extraction, scholarly transliteration, separation of attested historical forms from modern translations, geographical filtering, duplicate detection, and comparison with an existing historical-toponymic database. The current working table contains 538 attestations and preserves uncertain identifications, competing normalizations, and editorial decisions.
A preliminary normalized-Levenshtein analysis asks whether orthographic affinity between Slavic forms and German or modern local comparison forms varies geographically and by source tradition. Aggregated by identified place, Slavic forms are significantly closer to German comparison forms in Estonia (paired p=.003) and the Daugava corridor (p=.028), while northern Latvia shows no comparable asymmetry. The balance between German and local affinity also differs significantly among Middle Russian, Polish, and Ruthenian groups (Kruskal–Wallis p=.004), with Polish forms showing the strongest German affinity. The smaller Old East Slavic layer is treated separately.
The paper examines LLM failure modes including hallucinated identifications, over-inclusion, confusion between ethnonyms and toponyms, and false equivalence between historical and modern forms. The LLM serves as an interface for querying a fragmented multilingual archive under explicit linguistic control.
This paper discusses an in-progress project to create a list of works cited in Arabic books in the period 700-1800. The goal is to create a list for every book in the OpenITI corpus (8,810 Arabic books, exceeding 1 billion word-tokens) detailing citations, including of titled books and more ephemeral texts (such as notes). Such compiled lists will function as new metadata to accompany the books that survive today in the OpenITI corpus, and will also contain further information per work cited, including authorship attribution, where possible; frequency of citation in the OpenITI book; locations within the OpenITI book of the work citations; and judgments about the certitude of identifications. These lists can be used to address crucial questions about the Arabic tradition as a whole, such as:
The project will be iterative — producing, reproducing and improving the lists over time. All items on a list should be considered candidates, requiring subsequent scholarly judgment. Datasets will be released with the OpenITI corpus and can be ingested into the KITAB/OpenITI web application (kitab-project.org/explore) and other applications.
The method is a pipeline that teaches a model to recognise work references, treating detection — does this span refer to a specific written work? — separately from identification of which work it is, and evaluating the two separately. The pipeline will be discussed in the presentation, and likewise, the contribution that conversations with Claude made to its development.
The Dunhuang manuscripts are among the richest sources for the written cultures of medieval Central Asia and the Silk Road, yet more than ninety percent survive only as fragments scattered across collections worldwide. For a century, rejoining them has relied on the chance encounters of scholarly memory. We turn reassembly into a reproducible computational process: it first perceives what survives on each fragment, then synthesises the pieces into whole pages, with a large language model proposing layouts and scholars making the final decision. We first build a perception layer to make fragments machine-readable. Boundary-geometry matching identifies sibling fragments; patch-level handwriting recognition groups leaves by scribal hand; a codebook of 512 learnable visual primitives captures fine-grained style; and glyph-augmentation networks restore damaged characters into readable form. Where material is lost, a diffusion-based simulator, Fate Twin, generates plausible degradation paths, so each reassembly decision rests on simulated evidence rather than intuition. From these fragment-level signals, SRP performs global reassembly. It fuses fragment images, edge maps and OCR text to predict adjacency and relative direction, then a large language model reasons under placement constraints to reconstruct whole-page layouts. Reassembly is thus lifted from local pairwise matching to constrained whole-document inference. Every proposed placement is then scrutinised by domain experts, who make the final determination. Our methods extend beyond Chinese manuscripts to the endangered Khotanese script, for which we have assembled a dataset of 256 character classes and 201,452 images through self-supervised contrastive learning and iterative clustering. This human-machine collaborative approach offers a reproducible and generalisable path toward the digital reassembly of scattered manuscripts worldwide.
The Saharan manuscript tradition inverts misguided intuitions about "low-resource" regions of intellectual production. In fact, the Sahara is a region of extraordinary manuscript wealth and scholarly production, with thousands of volumes across family libraries and a nomadic civilization rooted in mobile knowledge transmission. Notwithstanding, the Saharan archive remains minimally catalogued, scarcely digitized, virtually unattested in LLM training corpora. This poverty of learning is the machine's, not the archive's. Drawing on in situ research on Saharan intellectual history, this paper reports two lines of inquiry conducted with frontier models, set against an exemplary archive of digitized microfilms at the University of Illinois Urbana-Champaign.
I begin with a stress test. Trained overwhelmingly on printed and born-digital text, models probed on reception history, scholarly networks, and bibliography in this tradition fail in a characteristic way: fluency inversely tracks reliability. They produce inauthentic if mimetic Islamicate prose, plausible-enough transmission chains, and confected dialectic precisely where training data is thinnest. In a field of plausible-sounding lineages and attributions, deadly nonsense can escape notice. The philologist's inherited disciplines of source criticism and transmission-evaluation turn out to be the working method for using these systems at all, after all.
Moreover, the Stewart microforms of Mauritanian mss at Urbana-Champaign, now digitized and with (occasionally) attributed hands, make a paleographic pilot testable. Can vision-capable models, taught explicit diagnostic criteria, triage manuscripts by regional script type (Andalusī, Maghribī, ṣaḥrāwī/shinqīṭī)? Harder still, can they sort unattributed leaves by scribal hand better than chance, at measurable levels of confidence, against adversarial cases like teacher-student stylistic continuity? What becomes findable matters: women's copying and annotation, legal-network geographies, visual graphs of modal logic applied to theology. I present this as experiment design and early probing, not results, offered by a scholar still learning how to put models to work beyond the chat window, and arguing that neglected traditions are the true test of whether a pre-print AI ecosystem serves the whole pre-print world or only its well-digitized provinces.
How do you apply universal, categorical, and computationally inspired annotation guidelines to language varieties that are historical, fragmented, and abound with variation? This is the question we, the PARSEME Ancient Greek team, have been faced with for years. In a nutshell, we use an existing Universal Dependencies treebank and make manual adjustments where necessary for the task at hand. We then enhance this treebank by means of manual annotation based on the PARSEME 2.0 universal annotation guidelines for multi-word expressions (MWEs). MWEs are expressions made up of multiple words, such as in front of or to make a suggestion. For PARSEME, words are orthographically defined; linguistically, the question is more complicated (see e.g. Taylor 2014; Dixon & Aikhenvald 2021; Haspelmath 2023). The PARSEME 2.0 guidelines are universal, in that they provide comparative concepts which can be applied to a range of language varieties (cf. Haspelmath 2010). Decision-trees based on the hypothesis that semantic idiomaticity goes hand in hand with morpho-syntactic inflexibility ensure replicability. Thus, diversity can be measured inter-lingually, but what about intra-lingual diversity? This is where even high-resource corpus varieties like classical Greek (5th c. BCE) and Latin (1st c. BCE to 1st c. CE) have challenged us (e.g. Fendel 2025). From sampling, through finding the so-called neutral form of each MWE token, to assessing the (in)flexibility of this neutral form during annotation by means of corpus queries, we have had to adapt (see Fendel, Squeri & Platanou 2026). As an illustrative example, I will draw on the Shared Task 2025 GRC data (classical literary Attic Greek courtroom oratory) (cf. Savary et al. 2026) and the UniDive WG1 Sub-task 1.6 Latin data (classical literary Latin historiography). Both datasets show a significant skew in the MWE tokens towards verbal MWEs rather than nominal, adverbial, adjectival, or functional MWEs. Both samples also have faced us with significant issues when distinguishing between categories of MWEs and between MWEs and fully compositional, flexible structures. The skew towards verbal MWEs can be explained diachronically in the languages’ history. The difficulties in distinguishing between structures can be explained not only synchronically in each language but also based on the annotation process. Finally, I will raise open questions that relate to future work we are planning.
How can we create reliable machine-readable sources for studying the global history of print? This paper presents our preliminary work on a project entitled ‘Cross-Cultural Analytics: Using Big Data, Large Language Models, Knowledge Graphs and Rare Books to Chart the Influence of Chinese Culture in the West’. Studying European-language texts on Chinese culture and its reception in the West, from the sixteenth to the nineteenth centuries, faces a basic obstacle: no corpus exists that is both relevant to this question and structured with the standardized, machine-readable metadata that computational analysis requires. This paper presents a pipeline built to close that gap.
The initial corpus draws primarily on public digital repositories, especially HathiTrust, Internet Archive, EBBO, ECCO, Gallica, focusing on those texts in early modern intellectual and imperial lingua francas (e.g. Latin, Spanish, French, and English). We have developed a workflow for retrieving bibliographic records, filtering material by date and language, identifying duplicate records and converting heterogeneous source files into a consistent structure. Each text is linked to bibliographic metadata, including title, author, language, date, source and repository identifiers. This creates a foundation for later semantic search and intertextual analysis.
OCR quality remains a major obstacle, particularly in older books with variable typefaces, damaged pages or inconsistent spelling. We have therefore established a correction process that combines rules-based and LLM-based methods with targeted scholarly review. Our experiments with LLM-assisted correction indicate considerable promise, but they also show the need to carefully optimize prompts, chunk lengths and other parameters. We discuss how an appropriate balance can be found between improving readability and preserving evidence about the printed page.
By concentrating on corpus creation and OCR correction, this paper argues that the pre-print past becomes machine-readable only as the result of significant human thought and labour. Without this laborious initial stage, later analytical work of the sort prized in the humanities is impossible.
Reading, understanding, and evaluating account books from late medieval Italy can be a daunting task, even for researchers who have spent years in local archives or other manuscript-holding institutions. While entries vividly depict the everyday realities of medieval Italian life, entries frequently include specialized vocabularies and expressions of exchange, whether industry-specific or based on local linguistic norms. Likewise, accounts feature idiosyncratic vocabularies, personal names, nicknames, or highly localized toponyms, since the goal was to memorialize a specific transaction that took place between two discrete parties. Similarly, specialized abbreviations are common and stand in for the denominations of local currencies (eg., lire, soldi, denari, scudi, etc.), units of measure for certain commodities (eg, braccia of cloth, staia of grain, botte of wine, etc.), or other oft-repeated terms. Additionally, since accounting documents were often created solely for a record-keeper’s or company’s internal use, handwriting was usually informal, as there was no expectation that anyone beyond those who had made the records or their close associates would read them. Financial records are also repetitive, featuring many of the same terms of exchange over and over, meaning that researchers may eventually tire of reading them and question their ultimate value for scholarly use.
At the same time—due in part to what scholars have termed “the documentary revolution”—financial records from the Italian peninsula represent one of the most abundant and detailed types of sources surviving from the later Middle Ages, and thus offer an unparalleled view into the societies in which they were produced. Past scholars, however, have primarily used financial records to study business and accounting history or to launch an investigation into one specialized industry, like the wool or silk trade. The University of Pennsylvania’s Medici-Gondi archive—a large collection of financial records featuring 13th- to 18th-century manuscripts tracking the business dealings of two prominent Florentine families—is one example of a prominent collection that has attracted little scholarly attention since it was acquired by the university some six decades ago. As a fellow at Schoenberg Institution for Manuscript Studies in August 2026, I examined a selection of fifteenth-century account books from the Medici-Gondi collection to test whether HTR technology could help to decipher these complex historical sources. This paper will describe my workflow, examine some of the roadblocks I encountered, and outline the potential for HTR tools to unlock the potential of the fifteenth-century account books, which offer the reader and intimate look into the personal lives of their Florentne creators.
Large Language Models are becoming a common point of entry for students and researchers working with premodern Islamic texts. They can locate a reported saying, explain a legal or theological position, and suggest a source within seconds. But an answer may be broadly correct while its evidence is not. A quotation can be slightly altered, a statement assigned to the wrong scholar, or a convincing reference given to a text in which it does not actually appear.
This paper examines this problem through a small comparative experiment focused on source fidelity. A set of questions will be selected from premodern Islamic materials in hadith, law, theology, and Qur'anic interpretation and submitted to several widely used language models under the same conditions. The references will first be checked against the relevant primary sources, providing a basis against which the models' answers can be assessed.
The analysis will distinguish between different kinds of reliability. Does the model give the right information? Does it attribute a statement to the right person or school? Does the cited work exist, and does it actually contain the material attributed to it? When the model presents words as a quotation, how closely do they correspond to the source? Particular attention will be given to references that sound credible but cannot be verified.
These questions matter especially for Islamic textual traditions, where the provenance of a statement is often inseparable from its scholarly value. The study therefore suggests that methods familiar from textual criticism and source verification can also help us evaluate AI-generated scholarship. Rather than treating accuracy as a single measure, it proposes source fidelity as a distinct category for assessing the use of LLMs in the study of premodern texts.
The Pseudo-Isidorian Collectio Decretalium (c. 830–850) is the pre-print era’s most ambitious enterprise of forgery: a vast canonical collection interweaving authentic and fabricated material, from conciliar acts to papal letters. Its opening part consists of letters attributed to the thirty pre-Nicene bishops of Rome, from Clement I to Miltiades – fabricated in their entirety, stitched from thousands of authentic biblical, patristic and legal excerpts to sound credible to contemporaries and posterity. This section – the field of my current research – most fully lays bare the workshop of a forger who deceived readers for seven centuries. The corpus offers a unique laboratory for one question: how can machines that generate plausible text today assist in studying plausible text “generated” in the ninth century?
The paper demonstrates a pilot LLM-assisted source inventory on a sample of the corpus: automatic detection of biblical quotations and allusions, distinguishing Vulgate wording, Vetus Latina readings and borrowings mediated by the Fathers, alongside patristic and canonistic excerpts. Results are validated against Hinschius’s apparatus and Karl-Georg Schon’s Pseudoisidor transcriptions, measuring precision and recall and typologizing errors. I will show where the model outperforms the traditional toolkit (scale, unflagged quotations, paraphrase) and where it fails (hallucinated attributions, mistaking a quotation’s intermediary for its source).
The demonstration is paired with a hermeneutic reflection: medieval forgery as a mirror of today’s anxieties about synthetic text. Pseudo-Isidore proves that dangerously plausible text is no invention of LLMs – and philology has long possessed the tools for unmasking it, worth translating into standards for working with language models. I close by situating the method within a wider pre-print ecosystem – retrieval over editions, verification on MDZ scans, the Clavis Canonum – as a working model for the lone scholar without programming skills or grant infrastructure.
LLMs are opening new possibilities for reconstructing lives, networks and societies at previously impractical scales, giving new reach to large-scale explanatory approaches. Yet only a fraction of the extant pre-print evidence, let alone the lost, is currently available to these systems, and that overwhelmingly through post-print editions, translations, records and analyses. Machine-Scale History asks what can be inferred at scale, how those inferences can be tested, and how historical understanding might be transformed by bringing previously inaccessible or underused primary evidence into systematic analysis.
"Vibe-coding", the practice of directing large language models to write, run, and improve code through natural-language instruction rather than active human programming, has moved quickly from a curiosity to a genuinely viable research method. This workshop teaches medievalists, whatever their coding background, how to vibe-code effectively and safely, and to recognise when the technique may give trustworthy results and when it will not.
A central question of responsible use: when can we trust the outputs of vibe-coded tools? The answer is predominantly one of verifiability — tasks with an easily checkable output, such as valid TEI-XML encodings automatically produced from a manual transcription, are far safer territory than tasks whose correctness cannot easily be confirmed. The practical core covers tool and environment selection, precise prompt specification, planning phases, architectural pre-decisions, incremental testing and review, and safe version control — consolidated through a guided project building a simple TEI-encoding web application.
What to bring. Participants need a laptop with access to a coding environment and an integrated agentic LLM. VS Code with GitHub Copilot is recommended, and participants with an institutional affiliation are encouraged to obtain free academic access to GitHub Copilot Pro beforehand (the approval process can take a few days). If you are unfamiliar with this kind of software, don’t worry: get an LLM to talk you through the setup (good practice for the session) and you will learn how to use it on the day.
Cliodynamics is the new transdisciplinary field that combines analysis of historical data with the tools of complexity science. One major question is, why do human societies experience recurrent waves of social turbulence and political instability that often end in an outbreak of internal warfare? A major effort of the Social Complexity and Collapse project at the Complexity Science Hub, which I lead, is to collect quantitative data on the dynamics of socio-political instability in past societies. We need such data to empirically test a variety of theories that propose explanations for why collective violence within polities waxes and wanes. My talk will report on the current work by our team on combining AI and human expertise to generate instability data for past societies.
Background reading: Turchin, Peter. 2023. End Times: Elites, Counter-Elites, and the Path of Political Disintegration (Penguin Random House, New York and London).
The contestation of religious authorities – direct verbal challenges to ecclesiastical legitimacy, clerical conduct, and sacramental power – is dispersed thinly across medieval inquisition records and overshadowed by other topics. Scholarship recognizes its presence in dissident milieus but has never quantified its frequency or examined the social and temporal factors shaping its expression. This dispersal across thousands of documents in dozens of registers made systematic analysis technically unreachable for unassisted scholarship.
We deployed Anthropic's Claude Sonnet 4 to classify and extract relevant passages from 4,357 testimonies spanning 20 inquisition registers (South-Western France, North-Central Italy, Switzerland, England; 1243–1522). LLM classification was validated against independent human coding on random samples of 200 testimonies per variable, achieving 85–99% agreement. We then examined whether four factors correlate with contestation frequency: religious culture, urban versus rural setting, temporal period, and gender.
Results reveal distinct patterns: reformistic dissidents (Waldensians, Beguins, Lollards) articulated authority contestation much more frequently than separatistic groups (Cathars, Apostles, Guglielmites). Urban-dominated registers paradoxically contain fewer such contestations. Temporal analysis demonstrates genuine growth in authority contestation toward the Reformation, affirming expanding lay engagement with ecclesiastical legitimacy. Gender showed no significant effect.
This study illustrates how LLM-assisted extraction, when grounded in rigorous prompt design and human validation, transforms what was previously dispersed beyond scholarly reach into a dataset amenable to quantitative cultural-historical analysis.
A database of some 75,000 picture descriptions from nearly 11,000 European manuscripts (800–1600), used to estimate the prevalence of medieval capital goods. Deliberately non-LLM, manual-extraction work — the baseline the machine-assisted programmes must beat.
Drawing on ongoing research with the St. Augustine Jewish Historical Society, this paper explores how large language models can assist in reconstructing Jewish and converso kinship networks connecting Iberia, Spanish Florida, Mexico City, and the Caribbean during the sixteenth and early seventeenth centuries.
Military musters, inquisitorial proceedings, wills, and genealogical inquiries provide the basis for investigating relationships sustained through marriage, patronage, and officeholding. LLM-assisted transcription, translation, and comparison help identify possible connections across dispersed records, which require verification against original manuscripts. The inquiry distinguishes documented kinship from plausible association, and Jewish ancestry from religious practice or accusation.
Collaboration between historians and volunteer researchers grounds these methods in shared questions of belonging and historical memory. The paper considers how AI can support historical inquiry while preserving uncertainty and the complexity of lives shaped by migration, conversion, and colonization.
For fifty years, machines have generated conjectures and killed them against evidence, one domain at a time: from Graffiti in graph theory, through BACON and Eureqa in physical law, to the Robot Scientist in yeast genetics. QUINCUNX uses LLMs to assemble such precedents into a general loop: it mints conjectures, selects them against evidence, grades the mechanisms behind the survivors, promotes the strongest into reusable inference instruments, and uses those instruments to infer what the surviving record does not contain. This talk shows what QUINCUNX has inferred from pre-print-era evidence such as Roman inscriptions, and how far its registered tests justify believing it, given the particular challenges of pre-print-era studies: sparse and unevenly surviving evidence, no possibility of experiment, and ground truths that are rarely knowable. It also touches on QUINCUNX's applicability to any field where decisive experiments are impossible, and to re-modelling the relationships and joins between fields themselves.
Long-run growth is driven by new ideas, yet the cultural environment shaping their production is difficult to measure over time. We use large language models to read 23,000 books from the Western canon and score whether each endorses, rejects, or merely depicts positions along six cultural dimensions. We accumulate the scores into inherited stocks and summarize them with an Innovation Wedge measuring cultural resistance to new ideas. Between 1000 and 1920 the wedge falls by 52 percent. We check the measure against blinded expert readings and modern surveys; an independent 5,000-book archive reproduces the decline. In a calibrated semi-endogenous growth model, the falling wedge raises 1920 productivity to between 1.5 and 1.8 times its counterfactual level, explains between half and two-thirds of the first sustained acceleration in productivity growth from 1500 to 1700, and accounts for 27–40 percent of productivity growth in 1920.
We design a method for measuring the risk preferences of agents in the deep past. The method combines a structural model of crop choice as a portfolio allocation with machine-learning prediction of expected crop returns, using historic agronomic and climate data. We estimate county-level risk preferences for the United States and farmer-level preferences in Kansas from 1889 to 1929. More risk averse farmers leveraged less, were less likely to purchase novel WWI Liberty Bonds, and were more likely to participate in local risk-sharing institutions. We show that higher risk aversion predicts slower tractor adoption and farm mechanization during the 1920s.
The palace-memorial (zouzhe) system reached its fullest articulation under the Yongzheng emperor (1723–1735): memorials bearing vermilion rescripts (zhupi) circulated between the throne and provincial officials, forming a vast administrative record. Historians have relied on close reading of selected exemplars. We ask what changes when the corpus becomes computable. In a system whose bottleneck was imperial attention itself, these documents record how a state managed overflow: prioritisation, verification, and the reproduction of administrative knowledge. We present a computational-history framework that fuses SikuBERT, a BERT model pre-trained on classical and historical Chinese, with the three-level coding of grounded theory. SikuBERT supplies dense semantic representations of memorial texts; open, axial and selective coding organises these into historically meaningful categories, from genre and subject to the tone and urgency of imperial feedback. The two components run iteratively, not sequentially: codings are checked against model outputs and re-derived from the representations, keeping distant reading accountable to close reading. The framework makes systematic analysis of massive historical collections tractable, and keeps interpretation firmly in the historians' hands, replacing exemplar-driven narrative with corpus-scale evidence. Applied to the Yongzheng corpus, the framework traces how information was generated, fed back and reproduced under the memorial system, offering new empirical evidence for information behaviour in Qing state governance. The framework is designed to extend across reigns and document types; we outline how large language models will enter the pipeline next, through assisted coding, corpus-wide hypothesis generation, and dialogue with the archive. The memorial system thus offers an instructive precedent for the computational study of pre-print cultures: a state-run information regime whose “language model” was the emperor himself.
Medieval genealogy is exceptionally vulnerable to the structure of its archive. Surviving documents privilege inheritance, title, property and institutional continuity, while cadets, cousins, wives, clerics, household associates, witnesses, patrons and other lateral actors frequently disappear from conventional narratives. Later pedigrees can therefore impose linearity upon relationships that may once have been experienced as far more distributed communities of kinship, obligation, patronage and memory.
This paper asks whether generative AI can make a longstanding but previously impractical form of historical inquiry more achievable: reconstructing relationships dispersed across people, places, archives and documentary traditions without collapsing possibility into proof. Developed through The Invisible House, a study of Ridel/Rudel networks across eleventh- and twelfth-century western France, southern Italy and England, the method treats the aristocratic family not simply as a succession of inheritors but as a distributed social system.
An LLM-assisted research environment interrogates charters, cartularies, editions, prosopographical databases, spatial relationships and scholarship iteratively and at scale, allowing weak connections between otherwise separated actors and archives to become visible. Rather than treating nominal similarity as evidence of identity, the method distinguishes source assertions, relational inferences and identity hypotheses, preserving the epistemic status of each claim. Generated hypotheses must survive human source inspection and attempts at falsification before entering the historical argument.
The paper argues that generative AI's most productive role in medieval prosopography is therefore not automated genealogy, but an interrogative layer between archive and scholar. By recovering overlooked actors and relationships, it may also expose the limits of later pedigree traditions themselves: beneath the linear genealogies through which medieval families were subsequently remembered lie older, dispersed networks whose surviving traces suggest that kinship, identity and family memory followed more complicated paths.
Large Language Models (LLMs) are extremely powerful generative machines that compress a huge amount of knowledge into a searchable embedding. These machines are great for processing multimedial data and perform many tasks at human level, but they contain subtle biases and we need techniques for probing, evaluate and control the knowledge extracted from their embedding space.
In this talk I will present three use cases of historical data processing using LLMs: one on data generation and two about data annotation.
The data generation case is the perspectivist expansion of the Seshat databank using Deepseek (China), Llama (US) and Mistral (EU) to compare proxies of LLMs' worldviews. The quality of generated data (which is possible only with models larger than 100 billion parameters) is validated against human-compiled data by means of semantic similarity metrics using proper embedding models. Results show interesting divergencies, especially on the interpretation of elite power type.
The data annotation cases are about the comparison of humans and LLMs in the classification of Structural-Demographic phases from short texts describing historical events and the LLM-driven annotation of toxic language in texts from the XIII to the XX century to understand how much it is related to the Structural Demographic cycles. In these cases we will see the inter-annotator agreement and the perplexity scores as metrics to evaluate the annotation tasks. Results show that humans disagree in data annotation mainly for biases due to their personal views and knowledge, while LLMs disagree by causal slicing, assigning different weights to causes of multi-causal events.
There are three take-home messages. 1) the good: LLMs can be used a research tool that allows replicability and scientific evaluation in digital humanities; 2) the bad: LLMs with more than 100 billion parameters are expensive, companies have access to them, universities not always; 3) the bias: yes, LLMs are biased, but there are prompting techniques to probe their internal knowledge, and a perspectivist approach should be used rather than relying on a single model to mitigate these biases.
How do scholars exert authority over a transformative technology that they did not build, and whose breathtaking development perhaps no one can control, predict, or even fully understand?
Coding Agents such as Anthropic’s Claude Code or OpenAI’s Codex have significantly expanded the ways humanists can address the retrieval, analysis and presentation of cultural heritage. At the same time, the latest iterations of the LLMs they are based on exhibit astonishing capabilities in Handwritten Text Recognition (HTR), often rivalling dedicated HTR applications and models. However, these benefits come at the cost of well-known drawbacks of commercial LLMs (problematic business model of the large US providers, carbon footprint, data centers, issues of data protection, intellectual property, and dependence on external services, among others). Additionally, due to their biases towards modern standard languages, LLMs used for HTR purposes often silently normalize or modernize their transcriptions, thereby distorting the results and making certain investigations impossible.
In my presentation, I provide an example of how to navigate this complex landscape. I present the web version of the multi-source HTR tool Polyscriptor and discuss how and under what conditions to replace commercial agentic models with open-weight models, thus fostering digital sovereignty. The archival material I will provide as an example is a corpus of postcards written by Soviet Ukrainian people deported to Nazi Germany for forced labor. Many of the postcards are written in complex layouts in a mixture of Russian and Ukrainian, requiring specific solutions for HTR.
This presentation outlines a pragmatic approach to balancing the capabilities of commercial models with the goal of greater digital sovereignty.
Speakers and participants include
Organising committee
Mirrored live from arsinq.com.