Summary: A head librarian is told the new AI assistant should “search the collection”. The collection is two thousand printed books, half of them Arabic, none of them digital. This is what it takes to turn a million pages into something a retrieval pipeline can actually use: non-destructive scanning, honest Arabic OCR targets, metadata that carries rights, and a QA loop per batch. Plus the two numbers that decide the cost before a single page is scanned.
The requirement was one line in a sixty-page tender: digitise up to 2,000 books so the digital library assistant can search them. The clarifications added a second line. Books average 500 pages. Nobody had counted the Arabic titles, the illustrated ones or the fragile ones. So the head librarian at a federal entity in the UAE was looking at roughly a million pages, of unknown condition, in two scripts, with a delivery date attached.
Her first instinct was the sensible one: this is a scanning job, get three quotes. It’s a scanning job the way building a house is a bricks job. The scanning is the part you can price per page. Everything that makes the pages useful to an AI assistant happens before and after.
Two numbers decide the cost
Before anyone talks about OCR engines, two assumptions set the budget, and both live with the library, not the vendor.
The first is pages per book. Five hundred is a reasonable average for a reference collection, and it multiplies everything: a million pages at 300 dpi colour is tens of terabytes of master images before derivatives.
The second is the fragile share. A book that can go on an overhead scanner with a V-cradle at normal speed costs one thing per page. A book whose spine can’t open past 90 degrees, or whose pages need handling one at a time, costs about double and takes longer. In this programme we assumed up to 25% of pages needed fragile or illustrated handling, wrote the assumption into the proposal, and priced volumes beyond it as a separate unit of work. If the real share turns out to be 10%, the customer saves. If it’s 40%, nobody argues about whose problem it is.
Scan without damaging anything
Non-destructive is a hard requirement for a library, which rules out the fastest method (guillotine the spine, feed the sheets). The capture runs on overhead or V-cradle book scanners at 300 dpi colour, in batches of 500 books scheduled with the library, by a UAE scanning bureau under our management. The library chooses whether the stations sit on its premises or the books travel with secure transport; both are priced.
Every book is listed first, with title, condition and language, and fragile items are flagged at intake rather than discovered on the cradle. That intake list becomes the spine of the QA process later.

Arabic OCR: set the target you can defend
The OCR engine is Azure AI Document Intelligence, Read model, which lists Arabic among its supported languages for printed text and runs inside the entity’s Azure tenant. That keeps a million pages of the national collection in-country, which matters for the same reason it matters for the assistant that will read them.
The targets we committed to are 95% character accuracy for English and 90% for Arabic, on printed text. The gap is deliberate. Arabic script is cursive, letters change shape by position, and older typesetting adds diacritics and ligatures that trip any engine. Promising parity and missing it is worse than promising 90% and hitting it. Handwritten and decorative text is captured as an image and flagged, not OCR’d and pretended.
Two things make the targets achievable rather than hopeful. The first batch is used for tuning: sample pages across the printing styles in the collection, measure, adjust. And the targets are stated as subject to document quality, because a foxed 1960s reprint is not the same input as a 2015 hardback, and the tender clarifications said as much.
Metadata is what the assistant actually searches
A pile of OCR text is not a library. Each book gets a record: title, author, ISBN, subject, language, page count, and a rights flag, in Arabic and English. The rights flag is the one people forget. The digital library assistant can search full text where rights permit and summarise where rights permit, and “where rights permit” is a field on the record, not a policy document someone remembers to check. If the flag says no, the assistant returns the catalogue entry and the shelf, and nothing else.
| Step | What happens | What goes wrong without it |
| Intake | Every title listed with condition and language; fragile items flagged; batches of 500 scheduled | Fragile books found on the cradle, schedule slips, disputes over the fragile share |
| Capture | Overhead or V-cradle scanners, 300 dpi colour, non-destructive, on site or at the bureau | Damaged spines, unusable low-resolution masters |
| OCR | Document Intelligence Read, in-tenant; 95% EN / 90% AR on printed text; handwriting flagged as image | Search that misses a third of Arabic queries and nobody knows why |
| Metadata | Title, author, ISBN, subject, language, pages, rights flag, in both languages | An assistant that summarises a book it had no right to open |
| QA per batch | Page-count check, image quality check, OCR sampling, re-scan of failed pages, QA report | Errors discovered at month six across all 2,000 books instead of at batch one |
| Output | Master TIFF or JPEG 2000, access PDF/A with text layer, metadata record; stored in the repository, synced to the data lake, indexed | A folder of PDFs nobody can search or preserve |
From pages to answers
The output that matters to the AI assistant isn’t the PDF. It’s the text, chunked with Arabic-aware segmentation so a passage doesn’t split mid-word, embedded with a multilingual model so one index serves both scripts, and stored with its metadata: source, language, page, rights. Queries run hybrid keyword and vector search with a semantic ranker, and every answer carries a citation back to the book and page. A librarian can check it. So can the visitor.
Sequencing matters too. Scanning starts in week four, and the first 500 books are searchable by the end of month three, alongside the catalogue search that already works from the library system’s own records. The assistant doesn’t wait for the last book. It gets more useful each batch.
What the librarian signed
Not a scanning quote. A programme with an intake list she controls, a fragile-share assumption she can challenge with her own shelves, OCR targets she can measure on batch one, a rights flag that protects the collection from its own assistant, and masters in an archival format that will outlive the assistant entirely.
Three quotes would have priced the bricks. If your collection is about to be “searched by the assistant”, talk to 10ⁿ Tech about the house.
Frequently asked questions
What accuracy can you expect from Arabic OCR on printed books?
Set separate targets per script. On printed text, 95% character accuracy for English and 90% for Arabic are defensible with Azure AI Document Intelligence Read, which lists Arabic as a supported print language, provided the first batch is used to tune on the collection’s actual typesetting. Handwritten and decorative text should be captured as images and flagged rather than OCR’d.
How do you digitise books without damaging them?
Use overhead or V-cradle book scanners at 300 dpi colour, never spine-cutting sheet feeders. List every title with condition and language at intake, flag fragile items before they reach the scanner, and work in scheduled batches with a QA report per batch.
What decides the cost of a library digitisation project?
Two assumptions: average pages per book and the share of pages needing fragile or illustrated handling, which costs roughly double per page. State both in the contract and price volumes beyond them as a separate unit of work.
What format should digitised books be stored in?
Master images in TIFF or JPEG 2000 for preservation, an access PDF/A with a text layer for reading, and a metadata record with title, author, ISBN, subject, language, page count and rights flag in both languages. Store in the institutional repository and sync to the data lake for indexing.
How does an AI assistant search digitised books?
OCR text is chunked with script-aware segmentation, embedded with a multilingual model and indexed with its metadata. Queries use hybrid keyword and vector search with a semantic ranker, and each answer cites the book and page. A rights flag on the record decides whether the assistant may return full text and summaries or only the catalogue entry.
Related resources
- Data residency on Azure UAE North depends on the deployment type, not the region name
- A bilingual AI avatar for government services that never leaves the UAE
- Sovereign AI in the GCC: what “in-country” actually requires
- Self-hosted cloud AI
Photo credits: Jean Vella and Mikołaj on Unsplash.