Digitising a 40-year paper archive
An extraction pipeline for several million scanned pages of mixed quality, typed and handwritten, across two languages.
- Sector
- Public sector
- Year
- 2025
- Duration
- 5 months
- Filter by discipline
- Document intelligence
Results
The challenge
Decades of records existed only as scans of varying quality — skewed pages, faded carbon copies, marginalia and stamps. Retrieval meant a physical search taking days, and no part of the collection was queryable.
Our approach
We measured the existing retrieval process first, then built a layout-aware extraction pipeline with per-field confidence scoring. Anything below the agreed threshold routes to a human review queue rather than being silently accepted, so the error rate is bounded by design rather than by hope.
The outcome
The collection became fully searchable, with a documented confidence level attached to every extracted field and a review queue that shrinks as the models improve.
Before
Physical retrieval request, average three working days, no full-text search, no way to answer aggregate questions about the collection.
After
Full-text and structured search in seconds, aggregate reporting across the whole collection, and a confidence score attached to every field.