
DHI Seed Funding Project 2025/26
Akha Voices in One Script: A Project of Digitalization and Standardization of Oral Texts
Learning and Growth Through the Project
by Mariia ROGOZINA
As a second-year English major, I was responsible for the digitization of Paul Lewis’s dictionary, which contains around 700 pages with approximately 40 words per page, its transcription into a new spelling system, and research into a potential translation app. This project offered me an invaluable opportunity to experience what it is like to work with the system of a language, transcribing and converting it without focusing primarily on meaning, as well as to understand how humanities students can contribute to seemingly coding-based projects.
My role mainly involved digitizing and transcribing Akha materials, checking OCR output, and researching the requirements for developing translation technology for low-resource languages.
1. My Work on the Project
1.1 Digitizing Akha Materials
The most crucial task was to scan the Akha materials and transform them into digital PDF texts. The original scanned pages of Paul Lewis’s dictionary contained specialized symbols and formatting that conventional OCR systems, including Tesseract, Google Drive OCR, Google Cloud Vision OCR, Azure OCR, and multimodal AI tools such as ChatGPT and Google Gemini, could not always recognize correctly.
This was especially the case with tone characters and tabbed spacing before the lines. Some of the symbols were not easy to find in Unicode, so I used the closest available equivalents:
U+02C7 ˇ: caron
U+02C6 ˆ: modifier letter circumflex accent
U+032D ̭ : combining circumflex accent below
U+032C ̬ : combining caron below
Tuning the LLMs based on the mistakes they made and purifying the original scans did not solve the problem. The most accurate method was simply to split each page into two halves and process them separately using AI models, especially Google Gemini.
I then manually compared the OCR-generated text with the original pages and corrected errors in symbols, spacing, punctuation, and layout in DOC files before converting them into PDFs. Some pages required particularly careful checking because the software could confuse similar-looking characters or change the original structure of the text.

Fig. 1 A typical page of Paul Lewis’s dictionary
1.2 Transcription and Text Organization
I also transcribed Akha written materials into a new spelling system. Because the archival materials had been produced by different researchers, they did not always follow the same writing conventions. My task was therefore to convert Paul Lewis’s transcription into the newer CAO system.
This made consistency in the digitized files especially important. A character, tone marker, or line break that appeared insignificant at first could affect how the text was interpreted or converted into another spelling system later.
The work required patience and close attention to linguistic detail. It also helped me understand why clean and consistently organized data are necessary before more advanced technologies can be applied.

Fig. 2 A page in two different spelling systems
1.3 Researching Translation Technology
Another part of my work was to research how a translation application for Akha might be developed. I researched the resources and processes that would be needed for translation between Akha and languages such as English or Chinese.
My research considered questions such as:
- What amount and types of texts would be needed to create a useful dataset?
- What role could dictionaries, aligned translations, transcription rules, and audio recordings play?
- What technologies and software would be required?
- How many hours of work would be involved?
I summarized this research in a field analysis report. Rather than developing the application myself, my responsibility was to examine what linguistic data, software, computing resources, and working time would be needed to make such a tool possible.
2. What I Learned
2.1 The Limits of OCR
Working with Akha materials showed me how dependent OCR performance is on the availability of suitable training data.
For widely digitised languages, many scripts, fonts, and document types are already represented in OCR systems. For a language such as Akha, specialised tone characters or page layouts may not be recognised at all.
The most reliable tool was Google Gemini because it produced the most accurate results among the tools I tested. It also clearly informed me when I had reached the usage limit. At the time when I was working on the project, it was the only tool that provided this information clearly.

Fig. 3 The Difference in Input and Output
2.2 Language Technology Begins with Data Preparation
The project also changed my understanding of translation applications. Previously, I focused mainly on the final product: a user enters a sentence and receives a translation. During this project, I learned how much preparation must happen before that stage becomes possible.
Data must first be collected, digitized, corrected, standardized, and translated. Audio tools would additionally require recordings and accurate transcripts. This helped me understand that transcription and data cleaning are not secondary administrative tasks.
The required and recommended working hours, as well as the software, frameworks, data types, and computing power described in our collective report, also showed that data collection and refinement take up a major part of the process. For example, one hour of audio may require more than 30 hours of manual transcription.
3. Conclusion
I am very grateful for this opportunity to merge my interests in languages and technology while improving my skills in both. I am excited to carry forward my deeper understanding of how important it is to place technology within its cultural and linguistic context in order to make it meaningful.
Most importantly, I realized that technology should adapt to the complexity of the real world, not the other way around.

Fig. 4 Samples of the Converted to CAO Paul Lewis Dictionary

