A working corpus of OCR-derived Tibetan texts for discovery, comparison, and downstream research
rKTs presents an extended repository of machine-generated e-texts. These digital renditions have been created using BDRC’s OCR application for Tibetan, released in March 2025. The models for character recognition were developed by Eric Werner in collaboration with Élie Roux, Pentsok W. Rtsang, and the Monlam AI team.
The present corpus largely constitutes raw OCR output without manual correction or editorial review. As such, researchers should note:
• Character recognition accuracy varies according to script type, image quality, and textual layout.
• The OCR models employed are in continued development.
• Results may exhibit inconsistencies in complex passages or where source materials present degradation.
• Occasional character substitutions, omissions, or misrecognitions may be present.
• Levels of quality are indicated by the following labels: Input (manually entered and reviewed); OCR+ (machine-generated with multiple proofreading steps); OCR (raw output without further verifaction).
Users are advised to verify critical passages against original sources when employing these materials for substantive philological or doctrinal analysis.
The rKTs team welcomes scholarly contributions in the form of corrected or edited versions of these texts. All contributions will be properly attributed to their original authors.
If you find errors in these e-texts or if you have good-quality e-texts to contribute, please contact us.
This short essay describes the history of creating Tibetan e-texts for Tibetan Buddhist studies (Daniel Wojahn, 2025).