Skip to content
Document intelligence

Digitising a 40-year paper archive

An extraction pipeline for several million scanned pages of mixed quality, typed and handwritten, across two languages.

Client under NDASample
Sector
Public sector
Year
2025
Duration
5 months
Filter by discipline
Document intelligence

Results

3 days → 2 minMedian retrieval time
97.4%Field-level accuracyMeasured against a 5,000-page human-labelled holdout set.
11%Routed to human review

The challenge

Decades of records existed only as scans of varying quality — skewed pages, faded carbon copies, marginalia and stamps. Retrieval meant a physical search taking days, and no part of the collection was queryable.

Our approach

We measured the existing retrieval process first, then built a layout-aware extraction pipeline with per-field confidence scoring. Anything below the agreed threshold routes to a human review queue rather than being silently accepted, so the error rate is bounded by design rather than by hope.

The outcome

The collection became fully searchable, with a documented confidence level attached to every extracted field and a review queue that shrinks as the models improve.

Before

Physical retrieval request, average three working days, no full-text search, no way to answer aggregate questions about the collection.

After

Full-text and structured search in seconds, aggregate reporting across the whole collection, and a confidence score attached to every field.