Trustworthy Metadata in the Age of AI: Foundations and Practice
Ontologies, SHACL validation, automated subject indexing, and AI-assisted cataloging: moving metadata professionals from cautious observers to active knowledge engineers.
Date & Time
UTC
Duration
2 hours
Format
Online
Course Overview
In response to critical practitioner concerns about professional deskilling, training ownership, and the risk of catalogers becoming mere proofreaders for chatbots, this course is designed to transform metadata professionals from cautious observers into active knowledge engineers. Specifically, participants will learn how to move flat, disconnected legacy records out of siloed databases into open, connected knowledge graphs. Since automated systems and Large Language Models (LLMs) can produce fluent but inaccurate, hallucinated, or inconsistent metadata, a technical validation layer is essential to enforce semantic accountability. Traditional cataloging skills, such as authority control, classification, controlled vocabulary, and data curation, are exactly the competencies needed to make AI-assisted metadata functional and equitable. This course provides an overview of the workshops in this metadata and AI course series.
Learning Objectives
By the end of this course, participants will be able to:
- Describe how ontologies, knowledge graphs, and rule-based data checks such as SHACL help connect separate systems and maintain consistent, trustworthy metadata.
- Compare traditional automated subject indexing tools with newer large language model approaches, and identify which is most useful or limited in different metadata workflows.
- Discuss how well generative AI tools handle key cataloging tasks, including authority control, linking to correct URIs, and creating BIBFRAME-compliant metadata, and explain why human review remains essential.
Course Contents
Introduction
Charlene Chou (10 mins)
Session 1: Building Semantic Infrastructure: Ontologies & Knowledge Graphs
Josh Falconer, Lead Data Taxonomist & Ontologist, The New York Times (30 mins)
Josh shares an applied methodology for semantic architecture, synthesizing principles from facet analysis theory, semantic domains, philosophy of language, cross-linguistic typology, NLP, and ontology engineering. Drawing on his experience connecting data silos and constructing knowledge graphs at an enterprise scale, he demonstrates how validation tooling such as SHACL and Avo enables knowledge engineers to write deterministic validation and inference rules. Participants will gain practical strategies for enforcing data quality and establishing a grounded and explainable meaning representation layer for both AI systems and human users.
Session 2: Automated Subject Indexing in Practice: Experiences of the Annif System Development
Mona Lehtinen, Chief Information Specialist, Finto team, National Library of Finland (30 mins)
This presentation will focus on the development of the Annif system for automated subject indexing. It will demonstrate an approach that combines traditional natural language processing (NLP) and machine learning (ML) techniques implemented in the Annif toolkit, with LLM-based methods for translation and synthetic data generation, and merging predictions from monolingual models.
Session 3: Can AI Catalog? From Longitudinal Assessment to Standards-Aware Metadata Production
Myung-Ja (MJ) K. Han, Andrew Turyn Professor and Metadata Librarian, University of Illinois Urbana-Champaign (30 mins)
This presentation will showcase a longitudinal evaluation of four generative AI models (ChatGPT, Copilot, DeepSeek, Gemini) and large-scale testing of BIBFRAME reconciliation workflows. Focus areas include authority control, URI accuracy, entity identification, and legacy data quality, which leads to a human-in-the-loop framework for standards-aware metadata production pipeline.
Interactive Q&A
Charlene Chou with all speakers (20 mins)
Schedule
| City | Time |
|---|---|
| New York, USA | Wed, 23 Sep 2026 at 09:00 – 11:00 EDT |
| Toronto, Canada | Wed, 23 Sep 2026 at 09:00 – 11:00 EDT |
| Chicago, USA | Wed, 23 Sep 2026 at 08:00 – 10:00 CDT |
| Helsinki, Finland | Wed, 23 Sep 2026 at 16:00 – 18:00 EEST |
| Berlin, Germany | Wed, 23 Sep 2026 at 15:00 – 17:00 CEST |
| London, United Kingdom | Wed, 23 Sep 2026 at 14:00 – 16:00 BST |
| Beijing, China | Wed, 23 Sep 2026 at 21:00 – 23:00 CST |
| Perth, Australia | Wed, 23 Sep 2026 at 21:00 – 23:00 AWST |
Course Details
Course structure and schedule information
- Course Code
- DCA0010
- Duration
- 2 hours
- Schedule
-
- Status
- Upcoming
- Course Fee
-
DCMI, ASIST & ISKO members
$25
Students
$25
Professionals
$50
Your Instructors
Charlene Chou
Associate Director, Dublin Core Academy
New York University Libraries
Charlene Chou is Associate Director of the Dublin Core Academy and Head of Knowledge Access at New York University Libraries.
View profile
Josh Falconer
Lead Data Taxonomist & Ontologist
The New York Times
Josh Falconer is Lead Data Taxonomist & Ontologist at The New York Times, joining from Bloomberg where he was Senior Ontologist modeling events and states for an enterprise knowledge graph.
View profile
Mona Lehtinen
Chief Information Specialist, Finto team
National Library of Finland
Mona Lehtinen is Chief Information Specialist of the Finto team at the National Library of Finland, the national vocabulary and thesaurus service behind the General Finnish Ontology (YSO) and the Finto.fi portal.
View profile
Myung-Ja (MJ) K. Han
Andrew Turyn Professor and Metadata Librarian
University of Illinois Urbana-Champaign
Myung-Ja (MJ) K. Han is Andrew Turyn Professor and Metadata Librarian at the University of Illinois Urbana-Champaign, researching data interoperability, metadata studies, and digital humanities.
View profile