Understanding Veridian’s digitization workflows and how they differ from some other projects is important, especially for those accustomed to using products like CONTENTdm.
For example, when developing a digital collection with CONTENTdm librarians often upload one digital file at a time, manually cataloging them as they go. That approach can work well for many collection types, but processing newspaper collections containing hundreds of thousands or millions of pages usually requires a more automated, batch-based workflow.
With Veridian, newspaper pages are typically scanned and processed in batches containing hundreds or thousands of pages. The resulting images, OCR, metadata, and structural data are prepared as digital objects and ingested into the Veridian platform together.
Our team manages the technical ingestion, platform configuration, software maintenance, and hosting, freeing library staff to focus on collection development, research support, community engagement, and other priorities.
“I love the fact that I can set a project in motion and just leave it alone. We derive a great deal of peace-of-mind reassurance from the deep support we receive from the Veridian team and the fact that everything runs smoothly without any intervention on our part.”
CLIFF WULFMAN, PRINCETON UNIVERSITY LIBRARY
This article explains the typical stages involved, from preparing and scanning the source material to making the finished collection searchable and accessible through the Veridian platform..
1. Plan the project
Before scanning begins, the institution and project team define the collection’s scope, requirements, and intended outcomes.
Planning may include:
- Identifying the newspaper titles, dates, editions, and source materials
- Assessing whether to digitize from original newspapers, microfilm, or existing digital files
- Confirming image, OCR, metadata, and structural-data requirements
- Defining responsibilities across the institution, Veridian, and other vendors
- Establishing delivery schedules and quality standards
- Determining access, copyright, licensing, and preservation requirements
Planning the full workflow at the outset helps avoid inconsistent outputs, missing metadata, and costly reprocessing later in the project.
2. Scan the newspaper pages
The source material is scanned to create high-quality digital images.
Scanning may be completed by the institution, or by trusted scanning partners. The most appropriate approach depends on the condition and format of the source material, available equipment, project scale, required output quality, and budget.
Newspaper images are scanned in batches, either from microfilm or paper originals, to produce TIFF images. Some libraries choose to do the scanning in-house, while others choose local scanning vendors, or we can recommend one of our scanning partners.
Original newspapers may offer strong image quality, but they can be fragile, difficult to handle, and expensive to transport. Microfilm may provide a more practical production source, although its quality depends on how and when the film was created.
Before large-scale scanning begins, representative samples should be tested to confirm that image resolution, contrast, cropping, orientation, compression, and file formats support the next stages of processing.
More information about the scanning process can be found in The Process of Scanning Newspapers.
3. Process the images and create digital objects
The scanned images are processed to create the files and data required for search, navigation, and online access.
The recommended digital objects for newspaper digitization are based on METS/ALTO XML and require specialized software to produce. Depending on requirements the batches of scanned TIFF images are usually shipped either to us at Veridian or to a selected vendor/partner for processing. Alternatively, some large projects choose to purchase the necessary processing software and produce the digital objects themselves.
The digital objects produced during this process include all the appropriate metadata, usually embedded within METS XML files.
The organization responsible for producing the digital objects would also usually do the image clean-up work, including splitting any “two-up” images, cropping borders, de-skewing, de-speckling, etc.
4. Ingest the collection into the Veridian platform
The prepared content is loaded into the Veridian digital collections platform.
Veridian Software provides the public search, browsing, viewing, and discovery environment. The Veridian team manages the technical ingestion process and configures the collection according to the institution’s approved requirements.
Depending on the project, configuration may include:
- Publication and issue browsing
- Full-text search
- Page-level viewing
- Search-result highlighting
- Metadata display
- Article-level access
- Download options
- Access controls
- Collection branding
- Mobile and accessibility considerations
The available functionality depends on the structure and quality of the source data, the collection type, and the agreed implementation.
5. Complete quality assurance
Quality assurance, or QA, takes place throughout the project and again before public launch.
QA can be done in several different ways, depending on the library’s preference. For some projects we set up a second “staging” copy of Veridian, and load all new batches there first. Library staff can then check the data on the staging site before copying it to the live site. Other projects are comfortable to load new data batches directly to the live site, and do their QA on the live site. If a vendor was chosen to produce the digital objects it is often possible for that vendor to load the objects to their own copy of Veridian, so online QA can be carried out by library staff as soon as the digital objects are completed, and before they are shipped.
Errors found during the QA process usually result in the entire digital object (i.e. a newspaper issue) being reprocessed. Or in rare cases an entire data batch containing many digital objects may be sent back for reprocessing.
For more information see QA ALTO XML to prevent common hidden errors.
6. Preservation
As well as being loaded to Veridian for online access, and regardless of where the Veridian site is hosted, a copy of the finished and QA verified digital objects is always shipped to the library.
For more information see Long-term preservation of digitized newspapers.
7. Monitor and maintain the collection
A digital collection requires ongoing attention after launch.
Our team never consider a digitization project to be “finished”. Having the newspapers digitized, preserved, and accessible online is only the first step. After that the collection needs to be nurtured and maintained, it needs to attract visitors and encourage those visitors to engage with the digitized content. Veridian features like Crowdsourced User Text Correction (UTC) and metadata editing allow both librarians and online patrons to contribute to and improve the quality of the collection, long after it’s first posted online.
We offer hosting, support, and maintenance services to support digitization projects in the long term. For most Veridian projects we install software updates with new features at least once every year, to keep them as fresh and up to date as possible.