Data conversion services
Transform scanned historical materials into searchable, structured data
Data conversion transforms scanned images into structured information that can be searched, organised, filtered and explored online.
For more than 20 years, Veridian has helped libraries, archives and cultural heritage organisations convert newspapers and other complex historical materials into standards-based outputs that support online access, interoperability and long-term collection growth.
Once data conversion is complete, Veridian Presentation Software can make the resulting collection searchable and accessible online.
What does data conversion produce?
High-resolution master images, typically uncompressed TIFF files from the scanning phase, are converted into structured, standards-based formats that prepare collections for online access and discovery.
Depending on the collection and project requirements, outputs can include searchable OCR text, metadata, METS/ALTO XML, PDFs and JPEG 2000 images.
How do metadata standards support digital collections?
Data quality directly affects how usable, searchable and sustainable a digital collection will be. Where appropriate, Veridian uses METS and ALTO XML, widely adopted standards hosted by the Library of Congress.
METS describes the structure of a digital object and connects its associated files and metadata. ALTO records OCR text and its position within a scanned page. Used together, they create a structured relationship between the original page image, searchable text and the publication’s physical and logical organisation.
METS and ALTO can help collections:
- Store searchable text at page and word level
- Record the position of words, lines and text blocks
- Preserve logical structures such as articles, headlines and bylines
- Connect OCR text with the corresponding page image
- Support interoperability and future platform migration

What data conversion options are available?
Veridian offers three data conversion approaches to suit different collection types, content structures, budgets and user needs—from automated page-level processing to detailed article segmentation and headline cleanup.
Automated Page-Level METS/ALTO
This automated option produces page-level METS/ALTO and OCR, making content full-text searchable while retaining a page-based structure. It is the most cost-effective approach for large collections where article-level segmentation is not required.
Estimated cost: US$0.15 per page
Best suited for:
- Page-based, text-heavy materials
- Newspapers, magazines, journals, and reports
- Collections with good-quality scans or microfilm
- High-volume projects where article-level structure is not required
Page-Level METS/ALTO with Text Block Auditing
Text blocks are manually audited to identify issues such as incorrect reading order or inaccurately captured content. This improves OCR usability while retaining a page-based structure.
Estimated cost: US$0.28 per page
Best suited for:
- Page-based materials with more complex layouts
- Newspapers and magazines with multiple columns or dense content
- Books and journals where improved reading order is important
- Projects requiring greater OCR quality assurance
METS/ALTO with Article Segmentation & Headline Cleanup
This option uses human review to identify individual articles and improve headline data, preserving more of the publication’s logical structure. It provides a richer search and browsing experience but requires more detailed processing.
Estimated cost: US$0.71 per page
Best suited for:
- Newspapers and periodicals where content is organised into distinct articles
- Magazines and journals where section-level navigation is valuable
- Projects prioritising article-level search and browsing
- Collections where preserving publication structure is important
How is data-conversion quality assured?
Quality assurance is a core part of our data conversion process. Before files are delivered, we validate outputs to identify issues that could affect search performance, usability or data integrity.
Our quality checks focus on structure, consistency and completeness, helping ensure the converted data meets agreed specifications. This process reduces downstream issues and helps ensure collections perform as expected once published online.
What does data conversion cost?
Data conversion is generally priced per page. The cost depends on the content structure, source-image quality, level of automation and amount of manual review required.
Automated page-level conversion is the most cost-effective option, while text-block auditing, article segmentation and headline cleanup require additional human review. We help each organisation balance cost, structure and usability according to its collection and access goals.