arXiv · 2025 · Preprint

Adisai Na-Thalang et al. (SCB 10X (Typhoon Team)) · → Paper · Demo: ✗ · Code: ✗

Documents the design and release of the first open, natural (non-read) conversational speech corpus for Isan, the most widely spoken regional dialect in Thailand, together with a linguistically-derived spelling standard, transcription convention, and phonetic dictionary for a variety that had no prior standardized orthography.

Problem

Thai speech AI development has concentrated almost entirely on central Thai, leaving regional dialects such as Isan with little open data despite being widely spoken. The two existing Isan-related corpora have specific limitations that make them poor foundations for conversational speech modeling. Suwanbandit et al. (2023) collected read speech from Korat speakers, but Korat is contested as a genuine Isan variety rather than a dialect of Nakhon Ratchasima province specifically. Thatphithakkul et al. (2024, LOTUS-TRD) collected speech from a single province (Khon Kaen), which does not represent the lexical and phonological variation across the wider Isan-speaking region. Both corpora also share three deeper limitations: they elicit speech by having speakers read aloud text that was translated from central Thai, which (a) produces narrative/read-aloud speech rather than spontaneous conversational speech, missing phenomena such as disfluencies, hedges, and prosodic variation that occur only in natural talk; (b) carries translation artifacts, since the translation process pulls phrasing toward central Thai rather than authentic Isan usage; and (c) inherits inconsistent, ad hoc spelling, because Isan has no officially standardized orthography, so transcribers spell words inconsistently, in ways that conflict with Thai orthographic conventions, or in forms unfamiliar to native speakers. The paper sets out to address all of these gaps: define the target Isan variety on linguistic grounds, design elicitation that yields natural conversational speech, and establish consistent spelling and transcription standards usable across multiple annotators.

Method

The paper documents a four-part corpus-construction and standardization pipeline; it does not train or propose any speech model.

Dialect definition. Rather than treating “Isan” as a single homogeneous variety, the authors classify Isan sub-dialects using Gedney’s (1972) tone-box methodology, which cross-tabulates tone-mark and consonant class against surface tone realization. Applying this method to Isan varieties yields two broad groups: a six-tone system (spoken across most Isan provinces, including Khon Kaen, Udon Thani, Ubon Ratchathani, and others, and the variety most represented in media) and a non-six-tone system spoken in fewer areas. The corpus targets the six-tone variety as the standard, by analogy to how Bangkok speech is treated as standard central Thai despite the existence of other central Thai sub-dialects (§3.1, §3.2, Table 2).

Conversational speech elicitation. To avoid the read-speech and translation-artifact problems of prior corpora, data collection uses stimuli designed around two principles: naturalness (prompts phrased in Isan spelling and vocabulary, e.g. using “ซุมื้อ” for “every day” and “เจ้า” for “you,” so that speakers are primed to think and answer in Isan rather than central Thai) and diversity (open-ended prompt types, including text questions, keyword description, audio-based questions, image/video description, and dialogue completion, chosen to elicit varied, non-repetitive responses). Image-based stimuli are used specifically where text prompts risk cueing the central Thai term instead of the Isan one, illustrated with the example of eliciting the Isan term for a food dish that most Isan speakers also know by its central Thai name (§4.1). Speaker gender and geographic origin are recorded alongside each response, and all recordings are manually verified as genuine six-tone Isan speech before transcription. The paper does not report corpus size (total hours, speaker count, or utterance count).

Spelling standard. Because Isan has no official orthography, the paper defines a rule-based spelling standard covering three word categories: proper nouns (kept in their central Thai spelling to avoid confusion), loanwords (spelled per the Royal Institute’s central Thai conventions, since phonetic spelling of loanwords was judged too unfamiliar and rule-heavy to be usable), and native Thai-derived words undergoing systematic Isan sound correspondence (e.g. central ร → Isan ฮ, central อึ → Isan เออ). For the correspondence category, the authors ran a preference survey of Isan speakers to determine, word by word, whether the phonetic (Isan-pronunciation) spelling or the central-Thai-conforming spelling was more familiar, and split correspondence pairs into two rule sets accordingly (§4.2, Tables 5-6). Native Isan-only words without a Royal Institute dictionary entry are spelled using the tone-box method directly, assigning them one of six unnamed tone categories (labeled T1-T6, since Isan tones, unlike the five named central Thai tones, have no established naming convention) (§4.2.4, Table 7).

Transcription convention and phonetic dictionary. A companion transcription convention specifies rules for punctuation, spacing, abbreviations, and handling of loanwords/romanization, plus explicit treatment of conversational-speech phenomena: syllable elision, sound assimilation, pitch raising, vowel lengthening, and consonant reduction (transcribed at their “deliberate,” fully-articulated form rather than the reduced surface form), and dedicated spelling conventions for sentence-final particles, discourse markers, and hesitation fillers (§4.3). Separately, the paper defines a grapheme-to-phoneme transcription scheme for a phonetic dictionary, compatible with PyThaiNLP’s G2P conventions, specifying Isan’s phoneme inventory (including /ɲ/, which does not occur in central Thai, and the absence of /r/), syllable structure C(C)V(V)(C)T, and a tone inventory of six native tones plus a borrowed “tri” tone for central Thai loanwords, along with rules for free variation and homographs (§4.4). The dataset, phonetic dictionary, and all four guideline documents (dialect classification, spelling standard, transcription convention, phonetic transcription guidelines) are released openly on Hugging Face and GitHub (Appendix).

Key Results

The paper reports no quantitative results: no corpus statistics (hours, speakers, utterances), no downstream model training, and no MOS/WER/CER or other speech-quality or recognition metrics. Its output is entirely the released corpus, phonetic dictionary, and the four linguistic standardization documents described above, illustrated throughout with qualitative example tables rather than benchmark numbers.

Novelty Assessment

The contribution is data-curation and linguistic-standardization work rather than architectural or algorithmic. What is genuinely new relative to the two prior Isan-adjacent corpora is (a) using naturalistic, non-leading, non-translated stimuli specifically designed to elicit spontaneous conversational speech instead of read/translated narrative speech, (b) applying a systematic linguistic method (Gedney’s tone box) to both define the target dialect and derive a consistent orthography for a variety that has no official writing system, and (c) pairing the corpus with a full transcription convention and a G2P-compatible phonetic dictionary, giving downstream users a complete, documented pipeline rather than raw audio alone. No new modeling technique, architecture, or evaluation methodology is introduced.

Field Significance

moderate — this is a first-of-its-kind open resource for the most widely spoken regional dialect of Thailand, addressing a genuine data gap for an underrepresented language variety, and its orthography-design methodology (tone-box classification plus a speaker-preference survey for correspondence spelling) could be reused as a template for standardizing other unwritten Thai regional dialects. Its impact is limited to being foundational infrastructure rather than a demonstrated technique: the paper reports no corpus statistics and no downstream TTS, ASR, or dialogue system trained or evaluated on the released data, so its practical value for speech generation work is not yet validated within the paper itself.

Claims

  • supports: Naturalistic conversational-speech elicitation surfaces phonetic and discourse phenomena, such as spontaneous prosody, particle reduction, and hesitation markers, that read-aloud or translated-stimulus elicitation systematically misses, limiting the usefulness of read-speech corpora for building models of natural conversational speech.

    Evidence: The paper contrasts its approach with two prior Isan-related corpora that elicit speech by reading translated central-Thai text, and documents conversational-only phenomena such as the particle “ครับ” surfacing as “คับ,” “ขับ,” “คาบ,” or “คราบ” through tone shift or reduction, and frequent hesitation fillers (“อืม,” “เออ,” “อา”) absent from narrative-style recordings. (§2)

  • supports: For a spoken variety with no official orthography, a tone-box-based phonological classification combined with an empirical speaker-preference survey can be used to derive a spelling standard that is both linguistically systematic and acceptable to native speakers, without inventing a new script.

    Evidence: The spelling standard applies Gedney’s (1972) tone box to classify the target six-tone Isan variety and assign untabulated native words to one of six tone categories (T1-T6), while a preference survey among Isan speakers determines, for words with regular sound correspondence to central Thai, whether the phonetic or the central-Thai-conforming spelling is more familiar, yielding two distinct rule sets. (§4.2, Tables 5-7)

  • complicates: Elicitation stimuli must be carefully designed to avoid priming speakers toward a dominant contact language, since even direct questions in the target variety can trigger vocabulary borrowed from the higher-prestige contact language rather than the target variety’s own terms.

    Evidence: The authors report that asking Isan speakers directly for the Isan name of a food dish frequently surfaced the central-Thai-influenced term (“ส้มตำ”) rather than the Isan term (“ตำบักหุ่ง”), because the central Thai term is itself already familiar to most Isan speakers; this required switching to image-based (non-verbal) stimuli to elicit the target term. (§4.1)

  • complicates: Standardizing spelling and transcription for a spoken-only, non-codified language variety involves rule sets complex enough that even the standard’s own authors identify residual inconsistency and usability limitations for annotators unfamiliar with the underlying linguistics.

    Evidence: The paper’s own limitations section reports that the spelling guideline requires transcribers to hold many interacting rules in mind and risks misclassification by non-specialists, that some standardized spellings (e.g. “ขอย” with a falling tone mark) conflict with forms already common in casual online usage, and that superficially similar words can receive different spellings depending on word category (e.g. “ซาง” for a common noun vs. “ชาง” for a proper name with identical pronunciation). (§5.2)

Limitations and Open Questions

The paper is transparent about several open issues in its own standardization work. The tone-based dialect classification is only one possible classification criterion among several linguistic methods, and future work could apply alternative criteria (§5.1). The spelling standard’s rule complexity, occasional conflicts with spellings already familiar to online Isan users, and inconsistencies in edge cases are acknowledged directly, along with unresolved questions about whether loanword and proper-noun spelling conventions hold across all cases (§5.2). Transcription of ambiguous forms (e.g. sentence-final particles with more than one plausible reading) still requires transcriber interpretation, and the convention leaves open how English loanwords should be rendered in Thai versus Latin script (§5.3). The phonetic dictionary is limited to root/morpheme-level entries that follow the paper’s own spelling standard only, its phoneme inventory may not match how some speaker subgroups (e.g. younger speakers) actually pronounce certain sounds, and the criteria for choosing a primary versus secondary pronunciation variant are not yet rigorously validated (§5.4). Separately, and outside what the paper itself discusses: because no corpus statistics or downstream model results are reported, the practical utility of the released data for training conversational TTS, ASR, or dialogue systems in Isan remains unverified within this paper.

Wiki Connections

This is a data-curation and linguistic-standardization resource with no trained model and no in-corpus citations; it does not directly inform any existing concept page. It could serve as a future training resource for regional-dialect Thai speech systems, but that use is outside the scope of the paper itself.