Forbes
    All case studies
    PROTOTYPE · B2L by Elevon

    AI Fills the Catalogue Record, a Human Checks Only the Doubtful Fields

    Digitizing a collection means turning scans into structured records, one field at a time. We prototyped both halves of that: a public catalogue with faceted search and series pages, and a back office where AI extraction is measured field by field and a review queue puts only the uncertain fields in front of a person. Sent with the offer, running on sample data.

    Client

    Collection platform

    Industry

    Agencies & consultancies

    Solution

    Catalogue with AI data extraction and a review queue (prototype)

    Deployment

    Interactive prototype

    Part of an offer, not a delivered project. The accuracy figures on the screens are sample data, not measured results.

    01
    The Brief

    The Brief

    The expensive part of digitizing a collection is not scanning it. It is turning each scan into a record: country, year of issue, denomination, series, catalogue number, motif. Typed by hand, the same handful of fields gets retyped hundreds or thousands of times, and the person doing it is the bottleneck for everything downstream: search, series pages, print.

    "AI will read it" is not an answer a client can check. The question that decides whether such a system is usable is what happens with the fields the model gets wrong. A single accuracy number hides exactly that: a system can read the country almost perfectly and the motif badly, and the average of the two tells you nothing about how much manual work is left.

    And the collection has a second audience. Without faceted filters, sorting and series pages, a large catalogue is technically online and practically unusable. One dataset has to serve both a curator working through a review queue and a visitor looking for one item.

    02
    How We Designed It

    How We Designed It

    The design principle throughout: the system may pre-fill anything, but it never quietly publishes a field it is unsure about.

    01

    Confidence per field, not one accuracy number

    The back-office overview breaks accuracy down field by field, and the review queue shows a confidence indicator next to each of the seven extracted fields. The fields the system did not determine reliably are framed in amber, so the reviewer's eye goes straight to them.

    02

    Human-in-the-loop by design

    The review queue puts the scan on the left and the extracted fields on the right, with three actions: confirm and next, mark as unclear, or skip. The queue only ever contains what the system was not sure about, so a person spends their time on the hard records instead of retyping the easy ones.

    03

    Export that survives typesetting and print

    Catalogue data leaves the system as XLSX or CSV and as a PDF overview, so a printed catalogue does not start with someone retyping the database into a layout file. The export section is in the prototype for the same reason as the rest: so the client can see the shape of the output.

    03
    How Delivery Would Run

    How Delivery Would Run

    The prototype covers two surfaces. The public catalogue has four screens: the home page with the collection's headline numbers and a quick search, results with faceted filters: country, series, years, motif: and sorting, an item detail with its series and a table of facts, and a series page. The back office has six sections: the overview with KPIs and per-field AI accuracy, batch upload of scans with progress per file, the review queue, records with full-text search and bulk editing, series and events, and export.

    Delivery would start where the risk is, not where the screens are easiest. The first milestone would be the extraction fields and the review queue on the client's real scans, because that is the only way to find out which fields the system reads reliably and which will keep needing a human. The catalogue, the search and the export are conventional work whose scope is already visible in the prototype.

    Nothing here is deployed. The records, the scan filenames, the accuracy percentages and the collection totals on the screens are sample data used to show the format of the interface. The bilingual catalogue and the roles in the back office are in the prototype for the same reason: so the client can see them rather than take them on trust.

    What this looks like in practice

    The client opens the review queue and works one record: the scan on the left, seven fields on the right, four of them confident and three flagged. They correct the two that are wrong, confirm, and the next record appears. Then they open the overview and see which fields the system reads reliably and which ones will keep a person in the loop.

    Sample output

    Review queue

    Anonymized preview, running on blind sample data.

    04
    What the Prototype Proves

    What the Prototype Proves

    This is a prototype, so there is nothing deployed to report on. What the client can check by clicking is this.

    A review queue with seven extracted fields, each with its own confidence indicator

    Three reviewer actions on every record: confirm and next, mark as unclear, skip

    Accuracy presented field by field instead of as one number

    Batch upload of scans with progress shown per file

    Faceted filters and sorting on the results screen, plus a series page

    An export path for typesetting and print: XLSX, CSV and a PDF overview

    05
    Why We Send a Prototype, Not a Deck

    Why We Send a Prototype, Not a Deck

    Every offer involving AI extraction says the same thing: the model will read the data and save the work. The client cannot verify that from a document, and they know it. What they can verify is a review queue: how a record looks when the system is unsure, how many clicks a correction takes, and which fields keep coming back.

    Showing accuracy per field is the honest version of the same claim. It admits up front that some fields will be read almost perfectly and others will not, which is what the client needs in order to plan how much human review the project actually requires. A single headline percentage would have looked better and helped less.

    The second reason is scope. A catalogue project quietly contains two products: the ingestion side and the public side. Putting both in one prototype makes that visible before anyone signs anything, and it makes the sequencing conversation possible, which half is worth building first.

    Want a similar transformation in your organization?

    Let's talk about how Elevon can help your team too.

    Book consultation
    Contact us
    Contact us

    We use essential and analytics cookies by default to ensure proper functionality and understand site usage. Marketing cookies are off unless you opt in. Privacy Policy