What Is RAG and How Could It Support Digital Collection Search?

What Is RAG and How Could It Support Digital Collection Search?

June 29, 2025

Our team has been researching and testing retrieval-augmented generation, or RAG, as a potential new discovery experience for the Veridian platform. RAG could allow users to ask questions in natural language, discover relevant collection material, and receive a concise response with clear links back to the original sources.

People increasingly expect to search by asking complete questions rather than constructing precise keyword queries. Everyday experiences with search engines, ChatGPT, and other digital tools have made intuitive, context-aware search feel familiar.

Those expectations are also shaping how people want to explore the digital collections held by libraries, archives, cultural heritage organizations, and educational institutions. We are therefore actively exploring retrieval-augmented generation, or RAG, as a new search option within the Veridian platform.

What is retrieval-augmented generation?

Retrieval-augmented generation is an approach that combines two processes:

  • Retrieval: The system identifies material from a digital collection that is relevant to the user’s question.
  • Generation: A large language model uses the retrieved material to produce a natural-language response.

In the approach Veridian is exploring, semantic search can help identify collection material that is conceptually related to the question, even when it uses different wording.

The retrieved records are then provided to the language model as context. The model uses this material to produce a response supported by links or citations to the underlying newspaper pages, documents, photographs, or other collection items.

In simple terms, RAG could allow someone to ask a digital collection a question in plain English and receive a concise response grounded in relevant source material.

Related reading: Semantic Search vs Keyword Search


Why could RAG improve digital collection discovery?

Traditional keyword search remains essential for finding known names, dates, places, phrases, and historical terms. However, researchers do not always know which words appear in the collection.

Historical records may use terminology that differs from modern language. Important evidence may also be distributed across several publications, dates, or items rather than appearing in a single search result.

RAG could help by:

  • Allowing users to search with natural-language questions
  • Retrieving relevant material from across a collection
  • Summarizing information found in multiple records
  • Providing links back to the original sources
  • Supporting follow-up questions and further exploration
  • Helping less experienced researchers begin exploring unfamiliar subjects

The intention is not to replace keyword search. RAG could provide an additional route into a collection for questions that are broader, more exploratory, or difficult to express as a precise keyword query.

How would RAG use collection content?

The large language model component of RAG does not learn from or train itself using collection content. Instead, it works in real time:

  • It generates answers only using the specific content retrieved from a collection in response to a user’s query.

  • It ensures transparency and makes it easier for users to engage directly with the original source content (and the collection).

  • And it does not fall back on general internet knowledge (as ChatGPT might if asked the same question).

These distinctions are important—as RAG preserves the integrity, ownership, and copyright protections of a digital collection.

What does a RAG search look like? 

A researcher exploring the 1918 influenza pandemic in New Zealand might ask: What was the Impact of the 1918 New Zealand influenza pandemic? 

The system could retrieve relevant newspaper reports and other collection records, then produce a concise overview based on that material.

RAG Search - Retrieval

The response could begin with links to the records used, enabling the researcher to inspect the original pages and assess the evidence directly. It could then summarize recurring themes in the retrieved material, such as public health measures, school closures, pressure on hospitals, disruption to businesses, or community responses.

The researcher could ask follow-up questions to narrow the topic, investigate a particular location, or explore how the effects changed over time.

Generate-LLM-Response

The generated response would serve as a starting point for discovery, not as a replacement for reviewing the original material and its historical context.

Why do links to original sources matter?

A fluent AI-generated response can appear authoritative, even when it is incomplete or inaccurate. For digital collections, users therefore need a clear connection between the generated response and the records that support it.

The approach Veridian is exploring places source discovery at the center of the experience. Users should be able to move easily from a generated response to the relevant newspaper page, document, photograph, or other collection item.

This allows them to:

  • Review the original wording
  • Understand the surrounding context
  • Compare different accounts
  • Identify the institution responsible for the material
  • Continue exploring the wider collection

RAG should guide people toward trusted collection content rather than separate an answer from the evidence behind it.

Related reading: Retrieval-Augmented Generation (RAG) and Copyright


Developing RAG as a Veridian search option

Veridian Software currently provides keyword-based search, browsing, viewing, and discovery capabilities designed for digitized historical and cultural heritage collections.

Our team is researching and testing how RAG could add another discovery pathway to the platform. The approach being explored combines semantic retrieval with generative AI so users could ask broader questions and receive responses grounded in relevant collection records.

This work is intended to complement Veridian’s existing search capabilities rather than replace them. Keyword search remains important for precise and reproducible research, while semantic retrieval and RAG could help users uncover related material when they do not know the exact terminology used in the sources.

The research is also focused on maintaining clear pathways back to the original collection content. Any future implementation would need to reflect the collection’s content, audience, copyright protection, technical requirements, and institutional priorities.

RAG is still an emerging area of development for the Veridian platform. We will share further information as the research and testing progress.

 

Interested in the future of digital collection search?

Related reading