Top

On this page

DANES Newsletter - October 2026

Computational methods are expanding steadily beyond cuneiform and the classical languages, and this month Ancient Egyptian stands out. Hieratic now has a diagnostic dataset for sign recognition and detection from Deir el-Medina, while machine translation is moving across the language’s whole history, from hieroglyphic-to-German to a new state-of-the-art for Coptic. Both papers also raise the question of data contamination and how far reported scores can be trusted.

The foundations of digital classics are expanding even further, among other things due to new OCR capabilities. The Open Greek and Latin Project now holds 41 million words of Classical Greek, and a companion article traces how Perseus data has been reused across treebanks, annotated texts, and named entity resources. New datasets on metre, recipes, and Latin–Greek alignment join a new issue of Digital Classics Online. All of this comes together in the newly published Oxford Handbook of Digital Classical Studies, whose 63 chapters map a field that has clearly come of age.

The study of the ancient world is also becoming more visible in the wider disciplines DANES draws on. The ACL 2026 proceedings include twelve papers on ancient languages and scripts, from Sumerian to Indus, and the DH2026 Book of Abstracts has around twenty, along with a mini-conference on AI and ancient scripts. The calls for papers for NAACL/COLING 2027, ACL 2027, and DH2027 openly invite work on under-resourced languages and linguistic diversity, so this is a good moment to submit.

The rest of this newsletter includes new datasets which make archaeological, environmental, and linguistic evidence open for reuse, and the special mentions ask how AI is changing research practice itself and how the past can be documented and protected. Upcoming seminars, workshops, and training series offer ways to build these skills. The job openings show institutions investing in long-term data infrastructure for the field.

Table of Contents

img

Recent Academic Publications

The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) took place in July 2026 in San Diego, California, and its proceedings, including Findings and the co-located workshops, are now available open access in the ACL Anthology. Among more than 6,000 papers, the following contributions deal with the languages and scripts of the ancient Near East and Mediterranean:

Five new articles in Digital Classics Online (volume 12, issue 1, 2026) apply computational methods to Greek and Roman sources. Several of the contributions discuss making interpretive criteria explicit and reproducible in digital workflows.

Additional publications of relevant to DANES from the past four months are:

Special Mention

Digital Humanities 2026: Book of Abstracts (conference proceedings, Zenodo), edited by Song-yi Jung, Gayeon Kim, Seonyeong Park, Iro Lim, and Haein Ji, collects more than 500 abstracts from DH2026, the annual conference of the Alliance of Digital Humanities Organizations (ADHO), held as a hybrid event in Daejeon, South Korea under the theme “Engagement.” Around 20 contributions deal with the ancient Near East and Mediterranean, including a linked open data Elamite dictionary, an IIIF-based edition platform for cuneiform tablets, large language models (LLMs) for Egyptian–Coptic translation, computational paleography of Greek papyri, bias in digital epigraphy, and a mini-conference session on AI and ancient scripts.

Large language models can predict the results of social science experiments (article, Nature), by Ashwini Ashokkumar, Luke Hewitt, Isaias Ghezae, and Robb Willer, used GPT-4 to simulate how representative samples of Americans would respond in 70 preregistered survey experiments, with 469 effects and 119,330 participants. The predicted effects correlated strongly with the actual ones, reaching the accuracy of pooled human forecasters, although they systematically overestimated effect sizes. This has interesting ramifications for AI’s ability to correctly predict real human responses. Data and code are available on Code Ocean.

A Theory-First Approach Towards Using Generative AI for Humanities Research (article, Computational Humanities Research), by Andrew Piper, argues that humanities research with generative AI should be theory-first rather than tool-first, because prompts act as measurement instruments and small changes to them can alter conclusions. He proposes a three-step framework in which theory guides problem formation, prompt validation, and interpretation, with robustness testing across prompt variations.

Reimagining research papers as interactive and reliable AI agents (article, Nature), by Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou, introduced Paper2Agent, a framework that automatically turns a research paper and its code into an AI agent built on the Model Context Protocol (MCP), which reproduces the paper’s results and answers new queries with validated tools. The code is available on GitHub.

Remote Sensing-Based Analysis of Archaeological Site Damage in Syria: Revisiting a Post-War Landscape (article, Heritage), by Jesse Casana, Jasper A. Clayton, Mary Lamberth, and Carolin Ferwerda, used high-resolution satellite imagery to reassess damage at around 200 archaeological sites across Syria since the early years of the civil war, documenting a novel form of mechanized looting in the north, expanded militarization and refugee camps on prominent sites, and small-scale looting at 59% of the sites observed.

A Review of Artificial Intelligence Techniques in Oracle Bone Inscriptions (article, ACM Journal on Computing and Cultural Heritage), by Jiaze Cai, Hengyi Li, Bang Li, and Lin Meng, reviews artificial intelligence (AI) methods and datasets for Shang-dynasty oracle bone inscriptions, from duplicate detection and fragment rejoining to character recognition and decipherment support. The authors identify data scarcity, weak generalization across photographs, rubbings, and facsimiles, and the gap between visual matching and cultural interpretation as the main bottlenecks.

img

Datasets Published

The 1909–1910 Gebelein Manuscript Inventory (Inventario Manoscritto): Archival Data from the M.A.I. Excavations (Journal of Open Archaeology Data), by Aneta Skalec, presents a digital edition of Ernesto Schiaparelli’s handwritten register of 1,906 artifacts from the Italian Archaeological Mission’s 1909–1910 season at Gebelein, Egypt, which were transferred to the Museo Egizio in Turin and date from the Predynastic period to Late Antiquity. The dataset includes a diplomatic transcription of the Italian text, English translations, standardized object descriptions, and a comparison of each entry with the museum’s collection database (SiME), organized by provenance area. It is available in the RADOGOST repository.

Archaeozoology of the Bronze Age Herders: An Open Zooarchaeological Dataset from the South Urals and Northern Kazakhstan (Journal of Open Archaeology Data), by Alexey Rassadnikov, presents data on 123,937 animal bones from 65 Bronze Age settlements, burial complexes, and mines in the Southern Urals and Northern Kazakhstan (c. 3500–1100 BCE), mostly from Sintashta and Alakul contexts. Analyzed over nearly two decades, the records cover species, skeletal element, fragmentation, measurements, surface modifications, pathologies, tooth wear, and X-ray age estimates for cattle metapodials. The dataset is available on Zenodo.

Conceptual Metaphors for Jealousy in Latin: A Corpus-Based Dataset (Journal of Open Humanities Data), by Roberta Grazia Leotta, presents 831 metaphorical uses of five Latin jealousy terms (invidia, livor, aemulatio, obtrectatio, and malevolentia), identified among 1,324 occurrences extracted from the Antiquitas section of the Library of Latin Texts (2nd century BCE–2nd century CE). The dataset is available on Zenodo.

C-Turkey: A Comprehensive Radiocarbon Dataset from Türkiye (Journal of Open Archaeology Data), by N. Ezgi Altınışık, presents 3,390 radiocarbon dates from 59 provinces of Türkiye, spanning from the Epipalaeolithic (c. 22,000 cal BCE) to the Late Ottoman period, including 885 dates previously absent from global datasets. Skeletal samples are annotated by species and linked to ancient DNA identifiers. The dataset and R scripts are available on Zenodo, with a searchable Dataset Explorer and a calibration tool online.

Emotion Vocabulary in Anatolian Languages: Etymology and Conceptual Structure (Journal of Open Humanities Data), by Maria Molina and Ksenia Uvarova, presents 72 etymological entries and 150 derivational nests (word families) documenting emotion vocabulary across seven Anatolian languages: Hittite, Luwian, Palaic, Lydian, Lycian A and B, and Carian. Compiled from the eDiAna corpus and the Hethitologie-Portal Mainz, the entries group cognates by emotional domain with Proto-Anatolian and Proto-Indo-European reconstructions and semantic derivation patterns. The dataset is available on Zenodo.

Gazetteers of Latin Authors and Works for Chronological Modelling in the LiLa Knowledge Base (Journal of Open Humanities Data), by Matteo Pellegrini, Francesco Mambrini, Giovanni Moretti, and Marco Passarotti, presents gazetteers of 173 Latin authors and 539 works, from the 3rd century BCE to present-day ecclesiastical Latin, for the LiLa knowledge base. Each author is linked to Wikidata and CTS URN identifiers and dated with new start- and end-date properties. The gazetteers are available on Zenodo.

Geochemical Reference Data on Ancient Iron Industry from Ariège, Aude, Pyrénées-Orientales (Occitanie, France) and Andorra (Journal of Open Archaeology Data), by Gaspard Pagès and Alexandre Disser, presents chemical analyses of smelting slag and iron ore collected through surveys and excavations in the eastern Pyrenees and the Montagne Noire, spanning from the late Roman Republic (late 2nd–1st century BCE) to 1876. The dataset is available on Zenodo and supports provenance studies of iron circulation in the ancient western Mediterranean.

Lead White in Context Across Greco-Roman Sources: The First TheSu XML Annotation Dataset of Arguments and Recipes, with Graph Visualisations and Discussion of their Design (Journal of Open Humanities Data), by Daniele Morrone, applies a new annotation schema to passages on lead white (ψιμύθιον/cerussa) in Plato, Theophrastus, Dioscorides, and Plutarch, alongside two modern studies replicating the ancient recipes. Arguments are mapped across sources, and recipes are broken down into their individual steps, which a Python pipeline converts into argumentation maps with Graphviz and network graphs with Gephi. The dataset is available on Zenodo.

M/OTHERING_Gr: A Database of Terracotta Figurines Representing Adults with Subadults from Ancient Greece (Journal of Open Archaeology Data), by Giulia Pedrucci, Ricardo Fernandes, and Carlo Cocozza, presents a database of 730 terracotta figurines depicting adults with children from mainland Greece, the Aegean islands, and Illyria, dating from about 700 BCE to 50 CE and found in sanctuaries, tombs, and other contexts. Compiled from publications, museum visits, and unpublished material identified with curators, the records combine iconographic, typological, contextual, chronological, and geospatial metadata, including caregiving interactions such as breastfeeding, eye contact, and body contact. The dataset is available on Hebe.

Open Digital Data on Funerary Landscapes and Settlement Organisation in Pre-Roman Apulia: Monte Sannace and Vaste (Journal of Open Archaeology Data), by Dominik Hagmann and Matthias Hoernes, presents a dataset of 379 funerary features, representing about 474 buried individuals, mainly from the late Classical and early Hellenistic periods (4th–3rd centuries BCE). Compiled from published literature and earlier datasets and mapped in QGIS, it includes spatial, chronological, depositional, and burial information in CSV and KML formats, with data dictionaries and bibliographies, and is available on Zenodo.

The Palatine Anthology – Book 1: A Metrical and Prosodic Statistical Study (Journal of Open Humanities Data), by Simona Nicolae, Cristian Șimon, and Ioan-Andrei Nicolae, presents a metrical and prosodic analysis of the 123 Greek epigrams (500 lines) in Book 1 of the Palatine Anthology, composed between the 5th and 10th centuries in hexameter, elegiac, and iambic verse. It records syllable quantity, metrical feet, caesurae, diaereses, bridge violations, and the position of the accent, and shows a transition from quantitative to stress-based versification. The dataset is available on Zenodo, with an interactive HTML viewer whose code is on GitHub.

T’OMIM: A Morphologically Annotated Dataset of Parallel Passages in the Hebrew Bible (Journal of Open Humanities Data), by David M. Smiley, presents 554 narrative parallels between Chronicles and Samuel–Kings and 256 poetic parallels, drawn from published scholarly catalogs and aligned at verse and word level to the ETCBC’s BHSA morphological database. The resulting tables contain 28,009 word-level rows with full morphological annotation, supporting the training and evaluation of models for semantic similarity and text reuse in Classical Hebrew. The dataset is available on Zenodo.

img

Events

Talks and Conferences

The Digital Classicist Berlin seminar series, organized by the Berlin-Brandenburg Academy of Sciences and Humanities, continues this academic year (2026/27) with the theme “Maße, Modelle und Muster in der Altertumsforschung” (Measures, Models, and Patterns in Ancient Studies), held fortnightly on Tuesdays at 4:15 PM (Berlin time) in hybrid format (zoom link). Talks in the next couple of months include:

Unicode Technology Workshop 2026 (UTW 2026): Unicode in the World will take place on 20–23 October 2026 in Nancy, France. The Unicode Consortium is co-hosting it with the partners of the Missing Scripts Program, including UC Berkeley’s Script Encoding Initiative (SEI). This year’s edition reaches out specifically to digital humanities researchers and script researchers. The first two days are hands-on tutorials on topics such as script shaping, font engineering, and converting legacy encodings to Unicode, and the last two days are talks and panels. Sessions of particular DANES interest include a tutorial on a digital typeface for Book Pahlavi by Amir Moslehi (20 October) and a talk on Elamite scripts: philology, typography, and encoding by Sina Fakour, Kaveh Ashourinia, and François Desset (22 October). Registration can cover the tutorial days, the session days, or all four days, with discounts for students and members.

Training Opportunities

DARIAH-CH Training Series 2026–2027: Building Digital Skills for the Arts and Humanities is a series of 11 online sessions held on Fridays from 1:30 to 3:00 PM between September 2026 and May 2027, and aimed at early-career researchers, doctoral and master’s students, and newcomers to digital humanities. Each 90-minute session combines a theoretical introduction, hands-on activities, a Q&A, and Open Educational Resources. Upcoming sessions include:

The Institute of Classical Studies in London is offering workshops in the upcoming months that are of interest to the DANES community:

The Humanities and Social Sciences National Team of the Digital Research Alliance of Canada offers a free online training series in 2026–2027 introducing researchers and students in the humanities and social sciences to digital tools for each stage of the research process, with no prerequisites for most sessions. Registration is through the Alliance’s Explora platform. Upcoming sessions include:

Call for Papers

NAACL 2027 and COLING 2027 share an ACL Rolling Review (ARR) submission cycle for long (8-page) and short (4-page) papers. NAACL 2027, the conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (ACL), will take place on June 1–5, 2027 in San Francisco, USA. Its theme track is “Language as a Medium for Agentic Communication,” and its areas include machine translation, multilinguality and language diversity, and low-resource methods for natural language processing (NLP). COLING 2027, the 32nd International Conference on Computational Linguistics, will take place on May 9–14, 2027 in Macau, China. Its special theme, “NLP for Linguistics,” invites work that uses NLP and large language models (LLMs) to advance linguistic research, with a focus on linguistic diversity and under-resourced languages, and its topics include NLP for digital humanities and cultural analytics, language resources, morphology and syntax, and language evolution. Submission deadline is October 12.

Theorising Networks: Relational Frameworks in Archaeology and History is a workshop organized by the Interdisciplinary Research Network for Ancient Near East that will take place on March 18–19, 2027 at the University of Helsinki, Finland (in person), examining the theoretical and epistemological implications of network approaches as frameworks for understanding an interconnected ancient world. Abstracts (300 words max) should address the implications of the “relational turn” in archaeology and history, network analysis as a conceptual framework versus a methodological tool, or interdisciplinary engagement between network science, archaeology, and history. Limited travel and accommodation support is available for selected speakers in financial need. Submission deadline is November 16.

DH2027, the annual conference of the Alliance of Digital Humanities Organizations (ADHO), will take place on June 28–July 3, 2027 in Galway, Ireland (partly hybrid) under the theme “Creativity,” and has opened its call for proposals to work across the full range of digital humanities, including work that does not engage directly with the theme. Submissions may be long papers (2,500–3,000 words, published in the Anthology of Computers and the Humanities), short papers (1,000–1,250 words), posters (500–750 words), half-day workshops, or day-long mini-conferences, and may be designated for the Technical Track or the new Digital Art(s) Track; all submissions must be in English and are made through ConfTool. Submission deadline is November 20.

Reading the Past with AI: Computational Approaches to Historical Artifacts will take place on April 29–30, 2027 at Sorbonne Université, Paris, France. Topics include handwritten text recognition (HTR) and optical character recognition (OCR) and computational paleography; computational codicology; computational restoration of fragmentary manuscripts, papyri, and inscriptions; multispectral and hyperspectral analysis of pigments, inks, and supports; computer vision for seals, coins, and ceramics; and knowledge representation and entity extraction from epigraphic, papyrological, and codicological corpora, while digitization projects and digital editions are outside the scope. Proposals of up to 500 words with a brief biographies should be sent to sophie.robert@sorbonne-universite.fr and maria-victoria.eyharabide@sorbonne-universite.fr, and the proceedings will be published as a peer-reviewed special issue of Digital Medievalist. Submission deadline is December 1.

DH Benelux 2027 will take place on July 6–9, 2027 at the Université libre de Bruxelles, Brussels, Belgium, as an in-person conference (pre-conference workshops on July 6), under the theme “Where Is Digital Humanities?”, which examines the disciplinary identity, public credibility, and engagement with generative AI of digital scholarship after large language models (LLMs). Topics include digital humanities methodology and institutional futures, computational ethics, LLMs and AI systems, reproducibility, multilingual DH, digital archives, and pedagogy. Submissions may be oral presentations, panels, or workshops (1,000–1,500 words), demonstrations (500–750 words), or lightning talks and posters (about 500 words), with accepted authors invited to submit full articles to the DH Benelux Journal. Submission deadline is December 7.

ACL 2027, the 65th Annual Meeting of the Association for Computational Linguistics (ACL), will take place on August 17–22, 2027 in Kyoto, Japan, as a hybrid conference. Its theme track, “Homogenization and knowledge collapse in LLMs,” invites empirical, theoretical, survey, and position papers on how training and alignment practices can reduce diversity in the outputs of large language models (LLMs), creating a “generative monoculture.” Topics include metrics for homogenization, the mechanisms behind it, human versus model diversity, and uniformity across languages. The general areas include language diversity, multilingualism, machine translation, phonology and morphology, syntax, and linguistic theory. Long (8-page) and short (4-page) papers are submitted through ACL Rolling Review (ARR). Submission deadline is January 4.

Fellowships, Scholarships, and Job Opportunities

Data Scientist at Johannes Gutenberg University Mainz is a full-time position (pay grade TV-L E13, can be split between two part-time employees) in the DFG-funded project “Beyond Categories – Towards Materiality. Wooden Figures from Ancient Egypt” in Egyptology. The duration of the position is for three years starting on January 1, 2027. The project combines Egyptological research with material analyses of wooden figures, in collaboration with wood specialists in Turin, and digital methods. The data scientist will design the project’s data and knowledge infrastructure, model research data relationally and then in the Semantic Web framework ResearchSpace, co-develop an ontology for archaeological data based on CIDOC CRM, and evaluate relational against semantic data models. Applicants should hold a master’s degree in computer science, information science, data science, digital humanities, or a related field, with experience in developing database models, programming skills in Java, C, or similar, and English at B2 level or above. German and experience in ontology development are desirable, and no background in ancient studies is expected. Applications, including a motivation letter, CV, and academic transcripts, should be submitted through the online portal. Application deadline is November 2.

eBL Data Steward at LMU Munich is a permanent, full-time civil-service position (Akademische Rätin / Akademischer Rat auf Lebenszeit, pay grade A13) at the Institute for Assyriology and Hittitology of Ludwig-Maximilians-Universität München, starting on October 1, 2027. The data steward will prepare and maintain data for the Electronic Babylonian Library (eBL). Further duties include contributing to research projects and grant proposals, supervising student assistants, and teaching two hours per week. Applicants should hold a degree in Assyriology or a related field with excellent grades and a completed doctorate, or be close to completing one. Demonstrated experience with data cleaning tools and large datasets is required. Applications, including a CV, transcripts, and a one-page motivation letter, should be sent as a single PDF (max. 5 MB) to sekretariat.jimenez@assyr.fak12.uni-muenchen.de with the subject line “eBL Data Steward.” Application deadline is December 18.

img

Did we miss relevant articles published in the previous month? Did we miss upcoming events in the next month? Would you like to ensure your news will appear in the next newsletter? Please send us an email at digpasts@gmail.com! Corrections to published Newsletters will be sent via the DANES mailing list.