Automatic Transcription of Ottoman Documents Using Deep Learning

dc.authorid0000-0003-2740-7946
dc.authorid0000-0001-8758-443X
dc.contributor.authorTasdemir, Esma F. Bilgin
dc.contributor.authorTandogan, Zeynep
dc.contributor.authorAkansu, S. Dogan
dc.contributor.authorKizilirmak, Firat
dc.contributor.authorSen, M. Umut
dc.contributor.authorAkcan, Aysu
dc.contributor.authorKuru, Mehmet
dc.date.accessioned2025-05-10T19:54:05Z
dc.date.issued2024
dc.departmentİstanbul Medeniyet Üniversitesi
dc.description16th IAPR International Workshop on Document Analysis Systems (DAS) -- AUG 30-31, 2024 -- Athens, GREECE
dc.description.abstractWith the accelerated pace of digitization, a vast collection of Ottoman documents has become accessible to researchers and the general public. However, most users interested in these documents are unable to read them, as the text is Turkish written in the Arabic-Persian script. Manual transcription of such a massive amount of documents is also beyond the capacity of human experts. With the advancements in deep learning, we have been able to provide a solution to the long-standing problem of automatic transcription of printed Ottoman documents. We evaluated three decoding strategies including Word Beam Search that allows to use a recognition lexicon and n-gram statistics during the decoding phase. Furthermore, the effect of lexicon size and coverage and language modelling via character or word n-grams are also evaluated. Using a general purpose large lexicon of the Ottoman era (260K words and 86% test coverage), the performance is measured as 6.59% character error rate and 28.46% word error rate on a test set of 6, 828 text lines.
dc.description.sponsorshipIAPR,DFKI,Univ W Attica,Natl Tech Univ Athens,La Rochelle Univ,Nust Seecs,Int Medias Data Serv,Natl Univ Sci & Technol, Sch Elect Engn & Comp Sci,EON,Metsob,Nonvtexneion
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [122E399]; TUBITAK
dc.description.sponsorshipThis study was supported by Scientific and Technological Research Council of Turkey (TUBITAK) under the Grant Number 122E399. The authors thank TUBITAK for their support.
dc.identifier.doi10.1007/978-3-031-70442-0_26
dc.identifier.endpage435
dc.identifier.isbn978-3-031-70441-3
dc.identifier.isbn978-3-031-70442-0
dc.identifier.issn0302-9743
dc.identifier.issn1611-3349
dc.identifier.scopus2-s2.0-85204528834
dc.identifier.scopusqualityQ3
dc.identifier.startpage422
dc.identifier.urihttps://doi.org/10.1007/978-3-031-70442-0_26
dc.identifier.urihttps://hdl.handle.net/20.500.14730/12940
dc.identifier.volume14994
dc.identifier.wosWOS:001334866300026
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherSpringer International Publishing Ag
dc.relation.ispartofDocument Analysis Systems, Das 2024
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WOS_20250302
dc.subjectOttoman Document Recognition
dc.subjectTurkish
dc.subjectDeep Learning
dc.titleAutomatic Transcription of Ottoman Documents Using Deep Learning
dc.typeConference Object

Dosyalar