Published on February 18, 2026.
Reading the Unreadable: AI & Early Arab-American Stories
Deciphering handwriting is a nightmare. Even in our own language, making sense of a scribbled note or an old family recipe can be a struggle.
Now, imagine trying to decode a letter from the 19th century written in a different alphabet, where the ink has faded, and the dots distinguish one letter from another have vanished.
If that letter is in Arabic, specifically in the rapid, stacked Ruq’ah script (ุฎูุทูู ุงูุฑููููุนุฉู), it becomes nearly impossible for non-native speakers. While most of us learn the neat, printed Naskh (ุฎูุทูู ุงููููุณูุฎู) style in school, the reality of historical archives is a chaotic web of vertical ligatures and scrawled shorthand.
This gap has left thousands of stories locked away in archives – until now. A project called Fihris is changing the game.
Developed by the Moise A. Khayrallah Center for Lebanese Diaspora Studies at NC State University (North Carolina, USA), it doesn’t just scan documents; it uses a custom-trained AI model to actually read them.
By adapting a machine learning model originally designed for German script, the team has created a tool that deciphers the daily lives, legal battles, and family gossip of early Arab immigrants to the Americas.
For learners, this is a goldmine: you can finally see the scanned handwriting next to a clear digital transcription, bridging the gap between textbook Arabic and the real world.
We sat down with Prof. Dr. Akram Khater and Dr. Mehreen Saeed to understand how they taught a computer to read what many humans can’t.
How was the idea for “Fihris” born?
Khater As part of the Khayrallah Center’s efforts to preserve the history of Arabic-speaking immigrants to the Americas, we have collected thousands of pages of letters, manuscripts, diaries, and other handwritten Arabic archival material.
As this collection grew, we came to recognize that digitizing it is not enough; we also had to make it accessible and discoverable. These two interrelated goals would open the archive to a far larger public and make the archive truly interactive, thus helping to generate new histories, all while helping to reconnect Arab Americans with the stories of their ancestors.
After doing an extensive search into the field of machine learning and the Arabic language, we realized that despite many impressive and laudable efforts, the tools we were seeking had yet to be developed. So, two years ago, we set about doing exactly that.
Picture Credit: The Hoda Z. Nassour and Herbert R. Nassour Jr., MD, Archive of Lebanese
Diaspora, The Khayrallah Center
Who is behind the project and how is it funded?
Khater The funding for this project comes from a combination of sources. In large part, it is self-funded by the Khayrallah Center. But we also received grants from the National Endowment for the Humanities, as well as FamilySearch.
The project PIs are Dr. Akram Khater (director of the Khayrallah Center and professor of Middle East History) and Dr. Chau-Wai Wong (Professor of Computer Science at NC State University). The main lead scientist is Dr. Mehreen Saeed, who works at the Khayrallah Center.
In addition, we worked with a team from Holy Spirit University in Kaslik (Lebanon) on the transcription of the dataset we used to train the AI model. We also had a number of undergraduate students who worked in one way or another on this project.
When and why did these immigrants arrive in the Americas?
Khater The first major wave of Arab migration (โค read more here) to the Americas came between the 1880s and early 1920s.
Like many immigrants from that time period, they were drawn by burgeoning job opportunities in North and South America, and were greatly helped in their migration by the transnational networks of information and funding that their compatriots created.
Example of a letter written by the Saudi King Abdulaziz Al Saud
This document from 1922 displays correspondence from King Abdulaziz Al Saud (ุงูู ูู ุนุจุฏ ุงูุนุฒูุฒ ุขู ุณุนูุฏ; 1876โ1953 – the first King of Saudi Arabia) to Amin Al-Rihani (ุฃู ูู ุงูุฑูุญุงูู; 1876โ1940 – a noted Arab-American author and diplomat).
Writing with warmth, the King welcomes Al-Rihani and apologizes for a communication mix-up regarding his travel schedule from Basra (ุงูุจุตุฑุฉ – a major port city in southern Iraq). He confirms that specific arrangements have now been made for Al-Rihani’s reception and his subsequent passage to Aqeer (ุงูุนููุฑ- a historic coastal port in the Eastern Province of Saudi Arabia).”
Picture Credit: The Hoda Z. Nassour and Herbert R. Nassour Jr., MD, Archive of Lebanese
Diaspora, The Khayrallah Center
In fihris.org, you can see the above scan of the letter and next to it the extracted Arabic text. Note that the Arabic text has not been edited or corrected.
ุจุณู ุงููู ุงูุฑุญู ู ุงูุฑุญูู
ู ู ุนุจุฏ ุงูุนุฒูุฒ ุจู ุนุจุฏ ุงูุฑุญู ู ุงู ููุตู ุงู ุณุนูุฏ ุงูู ุญุถุฑุฉ ุงููุทูู ุงูุบููุฑ ูุงูู ุตูุญ ุงููุจูุฑ ุงู ูู ุงููุฏู ุงูุฑูุญุงูู ุงูู ุญุชุฑู ุฏุงู ุงูุถุงูู
ุงู ูู ุณูุงู ุงู ูุดููุงู ูุจุนุฏ ูุจุง ุงุดุฑู ุทุงูุน ูุฑุฏูู ูุชุงุจูู ุงููุฑูู ุงูู ุชุธู ู ูุตูููู ุงูุจุญุฑูู ูุงููู ู ุฒู ุนูู ุงูุชูุฌู ุงูู ุทุฑููุง : ุงููุงู ูุณููุงู ุนูู ุงูุฑุญุจ ูุงูุณุนู ุชุงุงููู ููุฏ ุณุฑุฑุฉ ุฌุฏุงู ุจุฐุงูู ูุทุงูู ุง ููุช ู ุดุชุงูุงู ููููุงูู ููุฏ ุญููุฉ ุงูุงูุงู ุดููู ูุงูุญู ุฏููู:
ุงูู ุง ูุง ูุณุนูู ุงูุง๏ฑฃ ุงู ุงุถูุฑ ุดุฏูุฏ ุชุฃุณูู ูุนุฏู ุงุดุนุงุฑูู ููุง ุชูุบุฑุงููุงู ูู ุญูู ุชูุฌููู ู ู ุงูุจุตุฑู ุฐุงูู ุงูุงู ุฑ ุงูุฐู ุงูุฌุจ ูุชูุฑูุง ููููุงู ูู ุนุฏู ุงุฎุจุงุฑูุง ูู ุญุณูุจูุง ูู ุงูุจุญุฑูู ูู ูุงูุงุชูู ูุงูู ุณุฆูุช ุงูุฎุจูุฑูู ูู ู ุนุฑูุฉ ูุตูู ุงููุงุฉ ุงูู ุฑุงูุจ ุงูู ุงูุจุญุฑูู ูุนูู ุช ู ููู ุงู ุงูู ุฑูุจ ุงููุงุฏู ู ู ุฑุจู ุง ูุชุฃุฎุฑ ููุฐุง ูููุฐุง ูุญุฏู ุญุตู ุฐุงูู ูุฃุฑุฌููู ุงูู ุณุงู ุญู:
ูุญู ุจุงูุชุถุงุฑูู ููุฏ ุงู ุฑูุง ู ุญุณูุจูุง ุงููุตูุจู ุงู ูููุก ููู ุณูููุฉ ุชูููู ุงูู ุงูุนููุฑ ูุจูุตูููู ุงูููุง ุชุฌุฏูู ู ุญุณูุจูุง ุงูุณูุฏ ูุงุดู ุจุฃูุชุถุงุฑูู ูุจุงูุฎุชุงู ุชูุถููุง ุจูุจูู ุงูุงุญุชุฑุงู ูุฏู ุชู
ูกูฃูคูก ูขูง ุฑุจูุน ูก ูกูง ุชูข ูกูฉูขูข
In the name of God, the Most Gracious, the Most Merciful
From Abdulaziz bin Abdul Rahman Al Faisal Al Saud to the presence of the zealous patriot and the great reformer, Amin Effendi Al-Rihani, the respectedโmay his virtues endure.
Amin, peace and longing. To proceed: In the most noble of fortunes, your generous letter reached me, containing news of your arrival in Bahrain and your resolve to head towards us. Welcome and welcome, with all spaciousness and ease. By God, I was indeed very pleased by this, for I have long yearned to meet you, and the days have finally fulfilled my longing. Praise be to God.
However, I cannot but express my intense regret for not notifying you by telegram at the time of your departure from Basra. This matter caused a slight delay on our part in not informing our agent in Bahrain to meet you. I had asked the experts regarding the arrival times of boats to Bahrain, and I learned from them that the boat coming [from Basra] might be delayed. For this reasonโand for this reason aloneโthis occurred, so I beg your forgiveness.
We are awaiting you. We have ordered our agent [Abd al-Aziz] Al-Gosaibi to prepare a ship to transport you to Al-Aqeer. Upon your arrival there, you will find our agent Mr. Hashim waiting for you. In conclusion, please accept [my] respect, and may you remain [well].
27 Rabi’ al-Awwal 1341 17 November 1922
What topics do these manuscripts cover?
Khater The manuscripts ranged from Amin Rihani’s Arabic manuscripts, to Emily Nasrallah’s lectures, to a travel diary from Yabroud (Syria) to Argentina, to letters from a son in West Africa to his father in Bayt Shabab (Lebanon).
The topics covered range broadly. For example, family letters contained exchange of greetings, but also information about the separated family members’ lives. Sometimes their relatives in Lebanon or Syria included discussion of the situation in the village, town, or country. Other times, they sent requests for help in one way or another.
One of the most interesting collections is a series of letters that focus on a legal case wherein an immigrant living in Shreveport, Louisiana pursued legal action all the way to the Ottoman Supreme Court in Istanbul to reclaim land he alleged was illegally taken from him through the collusion of lawyers, judges, and people in his village of ‘Amsheet (Lebanon).
Quick reference: people and locations
LebaneseโAmerican writer, poet, and intellectual of the Mahjar (Arab diaspora) movement; wrote in Arabic and English and is known for The Book of Khalid and for bridging Arab and Western literary cultures.
Lebanese novelist, shortโstory writer, and activist; celebrated for novels and essays about village life, migration, and womenโs experiences in Lebanon.
A historic city in Syriaโs Rif Dimashq region, north of Damascus; noted for ancient sites and long settlement history.
A mountain village in Mount Lebanon, north of Beirut, known for traditional crafts and a historic bell foundry.
How did you train the model and how large is the dataset?
Saeed Collecting data and training our handwriting recognition model was an ongoing process. The first 500+ images were annotated and transcribed manually. After that, the AI would predict the annotations/transcriptions, and expert Arabic speakers cleaned them to prepare more data for training.
Now, we have 4000+ images for training our model. These images also include open-source Arabic datasets specifically meant for Arabic handwriting recognition.
โค We have published our data generation pipeline here.
What architecture does the AI use and how long was the training?
Saeed By “backend”, we assume the AI model used to recognize Arabic handwriting. It is adapted from the handwriting recognition model “Start, Follow, Read,” which was originally trained to recognize German handwritten text. (โค You can access the PDF paper about the model here.)
As mentioned earlier, the training was an ongoing process that included cleaning AI predictions to generate data for training, training the AI model, and making predictions from the AI model. It took about 12 months for the model to achieve 85%-90% accuracy, the main bottleneck being the manual intervention for data cleanup.
Picture Credit: The Hoda Z. Nassour and Herbert R. Nassour Jr., MD, Archive of Lebanese
Diaspora, The Khayrallah Center
Is there a human verification process?
Saeed Yes, there are still humans in the loop. All the transcriptions published on Fihris have been reviewed manually. With enough data, we hope to reach a point where we wonโt have to clean up the data manually.
How does the model handle the complex stacking in Ruqโah script?
Saeed Our handwriting recognition model has a CNN (Convolutional Neural Network) backbone connected to biLSTM (Long Short Term Memory).
The CNNs are capable of capturing different types of low-level and high-level features from images. Given enough examples, a network based on CNN can recognize complex handwritten patterns even with vertical overlaps.
โค You can learn more about CNN and biLSTM here.
Which combinations were the hardest to decipher?
Saeed The model had errors on letters having similar shapes, e.g, ุฏ was often confused with ุฑ and vice versa. Also, the model at times missed the dots or diacritics below or above the letters, e.g., ( ุง , ุฅ , ุฃ ). It also had problems with detecting word boundaries. So it would often join two words together instead of adding a space between the two words.
Does the AI standardize the text or preserve colloquialisms?
Khater The AI preserves the original spellings, including the colloquial terms as well as the misspellings of standard Arabic words. We want to show the user the exact original rather than flatten it into an MSA variation. This preserves the historical flavor and value of the manuscripts.
Picture Credit: The Hoda Z. Nassour and Herbert R. Nassour Jr., MD, Archive of Lebanese
Diaspora, The Khayrallah Center
How does the search function work?
Khater At this first stage in the project, the search is conducted on a word-for-word basis. However, in the next two iterations of this database, we will first introduce semantic search, which permits conceptual searching, and then we will add a “chatbot” that allows users to ask questions of the whole dataset.
How does the multilingual translation feature work?
Khater For translations, Fihris makes an API call to an outside source. However, we are also undertaking the creation of a lexicon of words and their translations because much of the language used in these late 19th century and early 20th century manuscripts is historical and colloquial in its usage. To improve on the translations, the lexicon will contextualize the words properly and render a more accurate meaning.
Will the model and dataset be available to the public?
Saeed We have published a part of the dataset that we used for training in a repository on GitHub (โค link: ู ุญุฑู Manuscripts of Arabic Handwriting (Muharaf) Dataset).
A subset of the dataset is available on HuggingFace (โค link: aamijar/muharaf-public). The model is also posted on HuggingFace Spaces (โค link).
Thank you, Dr. Khater and Dr. Saeed, for taking the time to share your expertise.
You can easily access the archive without registering on fihris.org.
Wanna cite this article?
Driรner, G. (2026, February 18). Reading the Unreadable: AI & Early Arab-American Stories. Arabic for Nerds. https://arabic-for-nerds.com/interviews/arabic-handwriting-fihris/
Published: February 18, 2026
Please check the citation before use, especially if your style guide has local requirements.
Publication ISSN: 3055-1730
Join the discussion: Sign in with social media for instant commenting, or register with email (comments require approval):