S Syntheta Historic Document Archive

Help & Documentation

Everything you need to know about using the Syntheta historic document archive.

🚀 Getting Started

Syntheta uses Google's Gemini AI to read, transcribe and help you explore handwritten documents from the archive. All documents are organised into collections which you can browse in the left-hand panel.

  1. Sign in with your authorised Google account.
  2. Browse the Archive panel on the left — click a collection to expand it.
  3. Navigate through sub-folders until you reach the individual page scans.
  4. Click any page to view the scan and request a transcription.

📁 Navigating the Archive

The archive is organised in up to three levels:

📁 Collection (e.g. Journals)
📁 Sub-collection (e.g. Journal 1852–1856)
📄 Page scans (.jpg)

A breadcrumb trail at the top of the panel shows where you are. Click 🏠 to return to the top level at any time.

🔍 Viewing Documents

  • Click any file name to open the scan in the centre panel.
  • Use ‹ Prev and Next › to step through pages sequentially.
  • Toggle Show scan to hide the image and give the transcription more space.
  • Use Copy or Download .txt to save a transcription.
Note: Some documents have been scanned with inverted (negative) colours to improve legibility — Syntheta handles both normal and inverted scans automatically.

✍️ Transcription

Transcriptions are generated by Google Gemini 2.5 Pro, trained to read handwriting from 1750–1950.

Single page

Type "transcribe this page" in the chat box while a page is selected, or simply ask any question about the document.

Page range

Enter start and end page numbers in the Pages fields in the sidebar, then click Transcribe Range.

Entire collection

Click Transcribe Entire Collection in the sidebar. This queues all pages in the current folder — processing time depends on the number of pages.

⚠️ Limitations
  • Transcription is performed one page at a time — batch jobs process pages sequentially.
  • Very large collections may take several minutes to complete.
  • Transcriptions are cached for 90 days. After that they must be regenerated.
  • Documents must be in .jpg format (PNG and WEBP are also supported).
  • Handwriting outside the period 1750–1950 may be less accurately transcribed.
  • Heavily damaged, faded, or folded documents may produce partial transcriptions — check [illegible] markers in the output.
  • Summaries of large collections are capped at 20 pages of context at a time.

💬 Using the Chat

The chat bar at the bottom of the page lets you ask questions about the currently selected document or collection. The context pill above the chat box shows what is in scope.

Example queries:

  • "Transcribe this page"
  • "Summarise this journal entry"
  • "Who is mentioned in this letter?"
  • "What date was this written?"
  • "List all place names mentioned in this collection"

Click the × on the context pill to clear the current context and ask a general question.

🎯 Transcription Accuracy

Gemini 2.5 Pro is highly capable but not infallible. Always verify transcriptions against the original scan, especially for:

  • Proper nouns (names, places, titles)
  • Dates and numbers
  • Words marked [?] (uncertain reading) or [illegible]
  • Abbreviations common to the period

If you spot a consistent error, contact the Archive Administrator using the form below.

🔐 Access & Privacy

  • Access requires a Google account that has been granted permission by the Archive Administrator.
  • All documents are private — only authorised users can view them.
  • Transcriptions are stored securely in Google Cloud and are not shared with third parties.
  • To request access for a colleague, use the contact form below.

📤 Adding Documents to the Archive

Documents are stored in a private Google Drive folder managed by the Archive Administrator. To have new documents added:

  1. Scan your documents as .jpg files at a minimum resolution of 300 dpi.
  2. Name files sequentially (e.g. page_001.jpg, page_002.jpg).
  3. Contact the Archive Administrator using the form below, describing the collection and providing a Google Drive sharing link to your scans.

The Administrator will organise the files into the archive structure and notify you when they are available.

✉️ Contact the Archive Administrator

For upload requests, access queries, or transcription corrections — we aim to respond within 2 business days.

Or email directly: info@syntheta.ai