This study proposes a multidimensional metric—Semantic Leakage Indicator (SLI) —to quantify meaning loss in AI-generated translations of Taiwanese into Mandarin. Drawing on a corpus of 26,046 Taiwanese sentences sourced from diverse spoken and written contexts, the analysis identifies five morpho-pragmatic features highly susceptible to semantic erosion: aspect markers, sentence-final particles, reduplication, kinship terms, and directional compounds. The SLI integrates entropy reduction, back-translation divergence, and semantic embedding shifts to detect untranslatable constructions that evade surface-level fluency metrics. Results reveal that even when lexical equivalents are present, high-SLI instances exhibit pragmatic distortion and affective dissonance. Cluster analysis delineates patterns of leakage across grammatical categories, while qualitative examples illustrate genre-specific vulnerabilities. The study advances translation studies by offering a scalable, language-sensitive diagnostic tool for evaluating AI performance on low-resource, morphologically rich languages. It contributes empirically to discussions on untranslatability, pragmatics, and meaning preservation in AI translation, with direct implications for algorithm design, language preservation, and corpus-informed evaluation in underrepresented languages.
【中文摘要】
本研究提出一項多維度指標—語意流失指數(Semantic Leakage Indicator, SLI),用以量化臺語翻譯為華語時,人工智慧譯文中之語意流失。研究使用之語料庫共計26,046組臺語句子,涵蓋多元的口語與書面語境。分析聚焦於五項高度易受語意侵蝕之形態語用特徵:體標記、句尾助詞、重疊形式、親屬稱謂詞,以及動向複合詞。SLI 結合語意熵減、回譯差異及語意嵌入位移三項演算法,以識別那些表層流暢度指標難以偵測之「不可譯」結構。研究結果顯示,即便譯文表面上存在詞彙對應,高 SLI 值案例仍常伴隨語用扭曲與情感不協調。群聚分析則勾勒出不同語法範疇中的語意流失樣態;另輔以質性實例,呈現各文體類型所固有之翻譯脆弱性。此研究為翻譯學提供一套可擴充、具語言敏感度之診斷工具,用以評估人工智慧在低資源、形態複雜語言上的表現,並於「不可譯性」、語用意涵與語意保存等議題上,貢獻具體實證價值,對演算法設計、語言保存與語料本位評估等均具實質啟示意義。