In humanities research — archaeology, history, historical geography, landscape studies — the spatial dimension is almost always part of the argument. Knowing which places a contribution discusses, and how often, is as valuable as knowing its subject or its chronology. Extracting that by hand from hundreds of PDFs, however, is not feasible.
geoNamesFromPdf was built to automate this step. It is a command-line tool (with an optional graphical interface) developed at LAD – Laboratorio di Archeologia Digitale of Sapienza Università di Roma, which identifies the toponyms in a PDF document and returns a sorted list, optionally enriched with coordinates. The project is open source (MIT license) and available on GitHub.
The tool was first presented in October 2025. This article describes the version revised in August 2026, which introduces substantial changes in both architecture and functionality.
What changed in the 2026 revision
The first version handled the essential use case well — give me a PDF, I return the place names — but it showed its limits when used on real work with large bibliographic volumes and with archaeological documentation. The revision addresses five fronts.
Fixes. The function that matched the text against an external gazetteer contained an error in its regular expression (a doubled escape sequence) that effectively prevented any match: from the command line, the --gazetteer option never found anything. The error message for a missing file was misleading, language detection was not reproducible across runs, and very long documents exceeded spaCy’s internal limit, aborting the analysis. All of these have been fixed.
Architecture. All extraction logic has been moved into a single, side-effect-free module (core.py), shared by the command line and the graphical interface, which previously duplicated — and had begun to diverge on — the same code. Text is now processed page by page, the language pipeline loads only the components needed for entity recognition, and models are cached: analysis is faster and requires less memory. The graphical interface runs the analysis in a separate thread and no longer freezes while it waits.
Structured output. Beyond the plain-text list, the tool now produces CSV, JSON and GeoJSON. Each toponym carries its number of occurrences, the pages on which it appears, and its provenance (automatic recognition, gazetteer, or both). The result is no longer a list of strings but a small dataset, ready for a GIS or a database.
The gazetteer as geographic data. The gazetteer can now be a .csv or .tsv file with columns for an identifier and coordinates: where present, the coordinates are attached to the recognised toponyms and carried into the GeoJSON. A gazetteer-only mode is also available, which disables automatic recognition and is intended for Latin, Ancient Greek and the other languages for which no reliable language models exist.
Zotero integration. A second graphical tool, zotero_assistant.py, applies the extraction directly to a Zotero library: it walks the records one by one, proposes geographic tags and — once the user approves — writes them onto the item. It is described further below.
One underlying choice was left unchanged, and in fact reinforced: the tool runs entirely offline and does not depend on any external AI service. All models run locally on the user’s machine.
How it works
The processing flow has six stages:
- Text extraction from the PDF with PyMuPDF, page by page, optionally limited to a range (
--pages "10-50"). - Language detection with langdetect, with a fixed seed for reproducible results; the language can be forced with
-l. - Model loading: a spaCy pipeline reduced to entity recognition only is loaded and cached.
- Recognition of geographic entities (
GPE,LOC,FAC), collected per page with a count of occurrences. - Gazetteer matching, if one is provided: a single pass over the text, respecting word boundaries, insensitive to case and diacritics, able to recognise multi-word names.
- Post-processing: removal of names on the exclusion list, de-duplication, sorting, and serialisation in the requested format.
A typical run on an Italian text:
python geoNamesFromPdf.py articolo.pdf -l it🌐 Detected language: it🧠 Using engine 'spacy' / model: core_news_lg v3.8.0
📍 Toponyms found in the PDF (12 total):
- Atene- Bologna- Egitto- Firenze- Grecia- Italia- Mediterraneo- Milano- Napoli- Roma- Sicilia- TevereWith --details, the same list reports the label, the number of occurrences and the pages for each toponym; with -f csv or -f json the same information is exported in tabular or structured form.
Recognition engines
The default engine is spaCy’s “large” CNN models, a good compromise between accuracy, speed on CPU and size. For cases that need higher precision on modern text, --engine spacy-trf selects the transformer-based models; --engine gliner uses GLiNER, a local model to which you state the categories to extract (for example “ancient settlement”, “hydronym”, “archaeological site”) without any retraining. Both alternatives require installing extra packages but, like the default engine, they run without a connection and without external services.
For text in ancient languages the advice is different: rather than relying on automatic recognition — which on Latin and Greek tends to get even the language wrong — it is better to use the gazetteer-only mode described below.
The gazetteer: from name lists to geographic data
A gazetteer is, in its simplest form, a list of place names, one per line:
PompeiErcolanoStabiaOplontisProvided with --gazetteer list.txt, the names present in the text are recognised and merged with the results of automatic recognition. Matching respects word boundaries — “Como” is not recognised inside “Comodo” — and is insensitive to case and diacritics.
The gazetteer can also be a .csv or .tsv file with additional columns for an identifier and coordinates:
name id lat lonButrint pleiades:530798 39.7456 20.0206Epirus pleiades:991380 39.5000 20.5000In this case the coordinates are attached to the recognised toponyms and, with -f geojson, exported as point geometries. This makes it straightforward to use local extracts of historical gazetteers such as Pleiades for the ancient world or GeoNames for modern geography: the PDF goes in as text and comes out as a map layer.
Finally, the --no-ner --gazetteer list.txt mode disables automatic recognition entirely and limits itself to matching against the list. It is deterministic, requires neither language models nor language detection, and is the preferred route for corpora in Latin, Ancient Greek or other languages not covered by the models.
Output formats
The default output remains the plain-text list, convenient for a quick read. The other three formats exist to integrate the results into a workflow:
- CSV — one row per toponym, with name, label, occurrences, pages, provenance and, where available, coordinates. Suitable for spreadsheets and database imports.
- JSON — the full result, including run metadata (language, model, processed pages).
- GeoJSON — a collection of point features for the toponyms that have coordinates, which opens directly in QGIS.
The provenance of each toponym (ner:GPE, ner:LOC, gazetteer) makes it possible, for instance, to distinguish the places found by the model from those confirmed by one’s own control list.
Graphical interface
For those who prefer not to use the terminal, gui.py offers an interface built with PyQt5 that exposes the same functions: PDF selection (including by drag and drop), page range, interactive management of the inclusion list (the gazetteer) and the exclusion list, and a view of the results with buttons to move a toponym quickly into one list or the other. In the revised version processing happens in a separate thread: the window stays responsive even on documents of several hundred pages.

Zotero assistant
Many research groups already organise their bibliography in Zotero with thematic tags. zotero_assistant.py builds on that practice, but it needs one firm convention: tags that denote a place must carry a leading @ (@Butrint, @Çuka e Ajtoit). The tool reads every @ tag in the library as its reference vocabulary, checks for each article which of those places occur in the attached PDF, and writes any new places with the same prefix. Without the convention the assistant has no vocabulary to reuse, so it must be adopted — for instance by batch-renaming the place tags in Zotero — before running it.
The tool works offline, through Zotero’s local API (version 10 or later). Once a library is chosen — personal or a group — the records not yet marked as done are listed. Selecting one, the assistant reads its PDF and shows two lists: the existing @ tags found in the text, pre-checked, with occurrence count and page numbers; and the new places found by automatic recognition, each with an editable field to normalise the form and, when it resembles an existing tag, a button to adopt it. The user approves what looks right; the tool sorts the item’s tags alphabetically, writes them and marks the record as done, so it does not come back in later sessions. Every action is written to a log file.


For articles in languages with no recognition model — Albanian is the typical case in Balkan archaeological literature — a selector lets you force the language or switch to a gazetteer-only mode that matches the text against the existing tags alone, removing the false positives of automatic recognition.

For PDFs with no text layer (scanned offprints), if ocrmypdf is installed the assistant can OCR the file on the fly and re-analyse; the resulting file can be saved to disk, while the attachment in Zotero is left unchanged.

Use cases in research
The most direct application is the geographic organisation of a bibliography. By processing a collection of PDFs in bulk and importing the resulting CSVs into a database, one obtains a table linking each publication to the places that appear in it; the Zotero assistant described above does the same thing directly inside the reference manager. From there the data can be queried geographically: which contributions deal with a given area, which mention two distant places together, how one’s digital library is distributed on the map. With GeoJSON output the same material becomes a map of bibliographic coverage, useful for spotting under-studied regions.
On larger corpora, systematic extraction of toponyms enables analyses of the geographic distribution of research: which areas receive more attention, how interest in a territory changes over time, which associations of places recur in the literature of a given field. For those conducting systematic reviews, a preliminary pass over abstracts and full texts quickly provides a picture of the geographic coverage of the studies examined.
A plausible workflow, starting from a collection of 500 PDFs on classical archaeology:
for pdf in library/*.pdf; do python geoNamesFromPdf.py "$pdf" -f csv -o "toponyms/$(basename "$pdf" .pdf).csv"doneThe CSVs produced this way are imported into SQLite or PostgreSQL and queried with ordinary SQL; the coordinate columns, if a georeferenced gazetteer was used, allow one to move straight to cartographic visualisation.
Installation and use
The project requires Python 3.11 or later.
git clone https://github.com/lad-sapienza/geoNamesFromPdf.gitcd geoNamesFromPdf
python3.12 -m venv venvsource venv/bin/activate # on macOS/Linux
python geoNamesFromPdf.py document.pdfOn first launch the tool checks the dependencies and the essential language models (Italian and English) and offers to install them automatically. Later runs do not repeat the check.
Main commands:
python geoNamesFromPdf.py article.pdf # automatic language detectionpython geoNamesFromPdf.py -l it document.pdf # forced languagepython geoNamesFromPdf.py document.pdf -p "10-50" # selected pages onlypython geoNamesFromPdf.py document.pdf -f geojson -o out.geojson --gazetteer pleiades.tsvpython geoNamesFromPdf.py --no-ner --gazetteer ancient.txt document.pdfpython geoNamesFromPdf.py --list-languages # installed modelspython geoNamesFromPdf.py --install-language es # adds Spanish
python gui.py # graphical interface (one PDF)python zotero_assistant.py # Zotero assistantSupported languages
The tool is preconfigured for ten languages — Italian, English, Spanish, French, German, Portuguese, Dutch, Modern Greek, Polish and Romanian — each installable with --install-language. The list is defined in the core.py module and is easily extended to other languages for which a spaCy model with entity recognition exists. Multilingual coverage is particularly useful in international research settings, where bibliography is often in several languages.
Reliability and AI-assisted development
The revision added a suite of automated tests (pytest, more than fifty) that verify, offline and without downloading models or requiring a Zotero instance, the behaviour of the pure functions: parsing of page ranges, gazetteer matching, reading files in various formats, serialisation, and the exchange with Zotero’s local API against a mock. It is a test of this kind that would have caught, at the time, the error in the gazetteer’s regular expression.
It is worth being explicit about the working method. The first version of geoNamesFromPdf, in October 2025, was written entirely with the assistance of Claude Sonnet 4.5 via GitHub Copilot. The August 2026 revision was carried out with Claude Sonnet 5 via Claude Code, in a session that began with a critical review of the existing code — which is how the gazetteer bug and the other problems surfaced — and continued with the redesign of the architecture, the addition of the new features, and the writing of the tests and documentation. It is a concrete example of how these tools can lower the technical threshold for those who, in humanities research, need bespoke software: not only for writing new code, but for reviewing and consolidating what already exists.
Conclusions
geoNamesFromPdf remains a deliberately narrow tool: it does one thing — find the places in a PDF — and tries to do it well, while staying simple to install and to use. The 2026 revision, however, also makes it suitable for work at a larger scale: processing of entire volumes, output that integrates with GIS and databases, gazetteers that carry coordinates, a mode designed for ancient languages, and an assistant that brings all of this inside Zotero. All without giving up offline operation and without depending on external services.
Links and resources
- Repository: github.com/lad-sapienza/geoNamesFromPdf
- License: MIT
- Technologies: Python, spaCy, PyMuPDF, langdetect; optionally spacy-transformers and GLiNER
- Reports and contributions: Issues on GitHub


