Handwritten Text Recognition (HTR) has been applied to digital humanities at high speed to improve the accessibility of historical manuscripts; however, research on how well HTR handles bilingual and code-switched materials, which can affect interpretation, is currently lacking. The above deficiencies are especially severe for cross-cultural archival collections; specifically, only 1.6 per cent of existing HTR research has focused on Chinese-English mixed-script materials, and none have used a controlled human comparison system with domain-expert transcribers. This paper presents the first controlled, empirically supported comparison of HTR and expert human transcription for Lin Yutang’s bilingual (Chinese-English) handwritten notebooks. A targeted review of recent SSCI-index studies on HTR evaluation (2018–2025) was conducted. This was followed by a 2 × 2 × 3 controlled comparative experiment in which 50 purposively selected manuscript pages (12,847 characters, 38.2% code-switching) were transcribed by three HTR systems (Transkribus base, Transkribus fine-tuned, and a transformer-based model) and three expert human transcribers with complementary expertise in paleography, Lin Yutang studies, and digital humanities. Transcriptions were coded and analyzed across four dimensions: accuracy, error typology (5 main categories, 17 subcategories), interpretive fitness (4 dimensions rated by a blind panel of three Lin Yutang scholars), and efficiency (via 10,000-iteration Monte Carlo simulation). Findings show that: (1) HTR systems have achieved a mean CER of 17.84% compared to 3.18% for humans (ηp2 = .80, p < .001, d = 2.84), and performance drops significantly on code-switching passages (CER increases to 32.83%); (2) errors in HTR are mainly “semantic-blind” (cross-script confusion: 13.1%; code-switching boundary misrecognition: 22.2%; semantic-contextual errors: 19.5%), whereas human errors are primarily “mechanical” (punctuation: 22.5%; cross-linguistic normalization: 32.9%); (3) HTR interpretive fitness scored 2.84/5, lower than 4.74/5 for humans (r = 0.72, p < .001); (4) only 42.6% accuracy was achieved by HTR on culturally significant terms (e.g., “Christo-Confucian” was never recognized correctly) compared to 100% for humans; and (5) a hybrid workflow of HTR + confidence-based human post-correction reached near-human performance (2.14% CER) at 66% lower cost and 83% higher throughput. This paper presents the first controlled HTR-human comparison of Chinese-English bilingual manuscripts, proposes a distinction between semantic-blind and mechanical errors based on error analysis theory, develops an Interpretive Fitness Audit framework for assessing transcription quality beyond aggregate CER, and shows that algorithmic transcription was more likely to misrecognize neologisms and code-switching boundaries, thus demonstrating that automated transcription may contribute to the loss of culturally significant meanings and code-switching boundaries embedded in bilingual manuscripts.
