The ǀxam Archive2 of 3

Restoring the Sound of a Sleeping Language

Why no tool can voice ǀxam clicks today, and a proposal for a community-led pronunciation layer built from the living sister language Nǀuu

Feasibility study and project proposalJune 202624 min read

First published at bleek-lloyd-rag.vercel.app.

Abstract

The Bleek–Lloyd archive preserves the words of ǀxam, a sleeping San language of the South African Karoo, but not its sound. The reading interface described in the companion paper renders the corpus aloud in English but cannot articulate the ǀxam click consonants (ǀ ǁ ǃ ǂ); the click characters are removed before synthesis, so the informant ǁkabbo is voiced "kabbo."

This paper examines that limitation. A survey of current speech technology and the primary literature indicates that no available system can produce ǀxam clicks, and that this follows from training data rather than configuration: neural models reproduce a speaker's timbre, not new phonemes, and none has been shown to generate click consonants absent from its pretraining. Two approaches could in principle address it. The first is a low-data formant-synthesis method, for which the acoustic parameters of Nǀuu clicks have already been measured (in part from Ouma Katrina Esau's own speech); the second is a higher-data neural method. Both would be bridged to ǀxam by linguistic reconstruction from the living sister language Nǀuu. On that basis the paper sets out a phased, consent-first, community-led proposal for an authentic pronunciation layer: the clicks and the corpus's recurring names voiced correctly, attested by living Tuu-family speakers and reconstructed by linguists. The realistic end state is accurate pronunciation, not a synthetic voice narrating in ǀxam, for which no data source exists. The broader claim is that modern speech synthesis, used on this foundation, offers endangered and sleeping languages a practical opportunity to recover a sound that no recording preserved, bounded by available data and grounded in reconstruction rather than invention.

1. The missing sound

ǀxam is a sleeping language: fully documented in the nineteenth-century notebooks of Wilhelm Bleek and Lucy Lloyd, but with no living mother-tongue speakers. A retrieval-augmented interface (described in the companion paper) makes that corpus readable, searchable, and, in English, audible, reading replies aloud with a word-level karaoke highlight.

The audio pre-processing chain, however, removes the click characters ǀ ǁ ǃ ǂ before sending text to the speech engine, because the English-trained voices mishandle them. The interface therefore reads every story aloud while mispronouncing the archive's names: ǁkabbo becomes "kabbo." Because the clicks are integral to the language, this is a substantive limitation, and it is the subject of this paper.

The clicks are phonemes, not optional ornamentation. A click-free rendering of ǁkabbo is not an accented version of the name but a different word, missing its initial consonant. Much of the corpus's distinctive vocabulary, and the names of its informants, are click-bearing; a system that cannot produce ǀ ǁ ǃ ǂ cannot say them.

2. ǀxam, Nǀuu, and the Tuu family

ǀxam belongs to the Tuu family (formerly "Southern Khoisan"), the name Güldemann (2005) proposed for the family, after the noun *tuu 'people'. The family has two branches, Taa and !Ui, and ǀxam sits in the !Ui branch, documented in the Strandberg, Katkop and Achterveld dialects recorded by Bleek and Lloyd.

The relevant point is an asymmetry. ǀxam is documented but silent; its closest living relative, Nǀuu (the Nǁng cluster, also !Ui), is endangered but still spoken by Ouma Katrina Esau, its last fluent speaker. Of the whole Tuu family, the only modern survivors are several Taa varieties in Botswana and Namibia and, in South Africa, the remnants of the Nǁng cluster. For this reason Nǀuu, rather than the geographically closer Bantu languages, is the relevant reference point for ǀxam's sound: the phonetics ǀxam lost can be approached through the phonetics Nǀuu retains. (Kora, sometimes compared to ǀxam for its click inventory, is a Khoekhoe language, not Tuu; that comparison is one of click typology across family lines, not common descent.)

3. Why no tool can voice ǀxam clicks today

This section is drawn from a survey of the primary literature and current product documentation (June 2026); sources are in the References.

3.1 Existing speech systems

A survey of current speech technology finds no dedicated text-to-speech or speech-recognition system for any Khoisan/San language.

A click-capable voice for any Khoisan language is thus an unaddressed gap rather than a solved problem.

3.2 Neural TTS: corpus requirements

Modern neural voices are trained on large corpora: roughly 27,000 hours across 16 languages for Coqui XTTS, ~50,000–60,000 hours for the VALL-E family, ~32 hours each across 1,100+ languages for MMS; a single-speaker English voice from scratch (the LJSpeech benchmark) is about 24 hours.

The apparent counter-example is that voices can now be cloned from seconds of audio: 3-second prompts for VALL-E and XTTS, under a minute for YourTTS, under two minutes for ElevenLabs' instant cloning. Such cloning transfers a speaker's timbre, conditioned on phonemes the base model already represents; it does not add a new phonetic inventory. No published model produces click consonants zero-shot from click-free pretraining. The one partial exception, phonological-feature representations that approximate unseen sounds (Staib et al., Interspeech 2020), is contested and has not been shown to yield a full click inventory. A natural neural Nǀuu/ǀxam voice would therefore require click-bearing, in-language audio; the realistic fine-tuning target is on the order of 5–10 hours of clean, transcribed, click-rich speech, collected through a recording programme.

3.3 Formant synthesis: a low-data alternative

The older alternative is better suited to this case. Formant / articulatory synthesis (Klatt-style) builds each phoneme from an acoustic-parameter table rather than a corpus, and is an established method for moribund languages with few or no speakers (Koffi & Petzold, Linguistic Portfolios 11, 2022, demonstrated on the moribund West African language Betine). It requires no corpus, only measured acoustic parameters.

For Nǀuu clicks, those parameters have already been measured. Exter (2011), "The Acoustic Modeling of Click Types" (ICPhS XVII), constructs a source–filter model of abruptly-released clicks and validates it against Nǀuu recordings from three speakers (Ouma Katrina Esau, Ouma Anna Kassie and Ouma Hanna Koper), reporting formant targets (apical click F1 ≈ 1284 Hz, F2 ≈ 4917 Hz; laminal F1 ≈ 1885 Hz, F2 ≈ 4425 Hz; apical roughly an order of magnitude more intense than laminal). The paper is an acoustic analysis rather than a synthesiser, but it provides the kind of parameter set such a synthesiser would draw on, derived in part from the last fluent Nǀuu speaker's own recordings. The related descriptive work (Miller et al., JIPA 39(2), 2009) classifies all 73 Nǀuu consonants, and further published work on click acoustics (Sands) provides additional grounding.

3.4 The Nǀuu → ǀxam bridge is reconstruction, not a second model

There is no ǀxam audio to train on, and there will not be. The path from Nǀuu to ǀxam is linguistic rather than computational: living Nǀuu phonetics (recorded, or read from Exter's tables) mapped onto a reconstructed ǀxam inventory and phonotactics (du Plessis 2018), and assembled into ǀxam words. A formant or concatenative approach makes this mapping relatively direct. A neural Nǀuu voice could in principle be driven with ǀxam phoneme sequences, but ǀxam prosody and coarticulation are unattested, so the suprasegmentals would be reconstructed rather than learned.

4. Precedents and comparable projects

The proposal does not stand alone. A small but maturing body of work has built speech technology for endangered, under-resourced and Indigenous languages, increasingly under community-defined governance. None addresses a click language, or a sleeping language reconstructed through a living sister tongue, so the proposal occupies a genuine gap; but each precedent settles questions of data scale, method, and governance. The full review, with figures and citations, is given in the supplement.

The most instructive precedent is Te Hiku Media's Papa Reo programme in Aotearoa New Zealand, which built a te reo Māori speech recogniser from over 300 hours of community-contributed speech, crowdsourced through its Kōrero Māori campaign (more than 2,500 contributors reading over 200,000 phrases in about ten days), reaching roughly 92 per cent reported accuracy (Te Hiku Media and NVIDIA, 2022). Two findings carry over: a usable low-resource model was built from a few hundred hours rather than the tens of thousands used for commercial English voices; and the data is released under the Kaitiakitanga Licence, which treats data as held in guardianship, with benefit returning to the source community. Te Hiku has publicly declined to surrender its corpus to external models, arguing that indiscriminate scraping of Indigenous-language data reproduces a colonial pattern. The Lauleo project carried this model directly to Hawaiian (ʻŌlelo Hawaiʻi), in collaboration with Te Hiku and the University of Hawaiʻi at Hilo, showing that the approach transfers to a second community under a derived guardianship licence.

On data scale and tooling, the National Research Council of Canada built Inuktut speech recognition from on the order of seventy-five hours of transcribed speech and released the roughly 1.3-million-pair Nunavut Hansard parallel corpus (Joanis et al., 2020), and Australia's Elpis (Foley et al., 2018) lets community language workers train their own recognisers without machine-learning expertise. Mozilla Common Voice confirms that the field treats tens of hours as a meaningful low-resource working point, and that it holds no Khoisan dataset, just as Meta's Massively Multilingual Speech has no Khoisan coverage. The low-resource neural-synthesis literature places intelligible new-language voices at a few hours of paired, transcribed audio, with roughly twenty to thirty minutes marginal, which is the basis for this proposal's five-to-ten-hour neural target.

The decisive negative finding is that no text-to-speech or speech-recognition system exists for any Tuu, Kxʼa or Khoe-Kwadi language, and no formant or articulatory synthesiser exists for any click language. Where mainstream systems produce clicks at all, it is for the Bantu click-borrowing languages (isiXhosa, isiZulu) and only because in-language training audio contained them; the clicks were learned, not generalised. The acoustic groundwork for the relevant language nonetheless exists, measured in part from Esau's own speech (Exter 2011; Miller et al. 2009): the science to build clicks exists in pieces but has not been assembled into a working voice. That is the space this proposal enters.

Taken together the precedents support the proposal's positions: a few hours suffices for fine-tuned synthesis, so the binding constraint is the session capacity of a single elderly speaker, not raw volume; the low-data formant route is better matched to a sleeping language than the data-hungry neural route; the field favours tools community members can operate; and community-defined governance (the Kaitiakitanga Licence, Indigenous data sovereignty, and in this jurisdiction the San Code of Research Ethics) is the established operating frame, not an optional overlay.

5. Proposal: an authentic ǀxam pronunciation layer

5.1 Vision and deliverable

This paper proposes a ǀxam pronunciation resource integrated into the archive interface, to be developed with the Nǀuu-speaking community and linguists specialising in ǀxam and the Tuu family. The realistic end state is accurate pronunciation: the clicks and the corpus's recurring names and terms voiced correctly, attested by living sister-language speakers and reconstructed by linguists. It is not a synthetic voice narrating in ǀxam, for which no data source exists. The deliverable comprises:

surfaced as audio against glossary terms and informant names, so that the reading interface also supports listening.

5.2 Data requirements

The data needed differs by approach (sources in the References; a fuller data and cost model is given in the supplement):

ApproachData required
Neural Nǀuu voice (fine-tuning a multilingual base)~5–10 h of clean, transcribed, click-rich speech, collected through a recording programme
Formant synthesis (clicks + key terms)No corpus; acoustic-parameter tables, already partly published (Exter 2011)
ǀxam reconstruction stepNo additional ǀxam audio; linguistic mapping from Nǀuu onto du Plessis's reconstruction

Existing Nǀuu audio is a useful starting point but not a ready TTS corpus: the ELAR archive holds approximately 22 hours across the last 10 speakers (deposit dk0089), and a Nǀuu revitalisation project holds recordings of Ouma Katrina Esau and her students, including time-aligned audio of the 160-page Nǀuu reader (deposit dk0505). This is multi-speaker documentation, valuable for acoustic reference and reconstruction but not a clean single-voice TTS corpus.

5.3 Phased plan

5.4 Roles

An indicative division of labour, by where the rights and expertise would sit. Nothing here is arranged: no individuals or institutions have been approached or have agreed to take part, and none are named for that reason. This is a hypothetical structure, not a commitment by anyone.

5.5 Ethics and data governance

Data governance is integral to this work, not an addendum. It involves an elderly last speaker, and the constraint is human rather than computational: a recording programme is limited by her voice and session capacity and must be paced accordingly. It should follow the CARE Principles for Indigenous Data Governance (Collective benefit, Authority to control, Responsibility, Ethics; Carroll et al., 2020): free, prior and informed consent, community ownership of the recordings, and benefit returning to the community. The governing principle is engagement rather than extraction, which is also why the reading system has so far deferred a synthesised voice pending a partnership of this kind.

Because the speaker and her community are San, the binding instrument is the San Code of Research Ethics (South African San Council, 2017), the first research-ethics code issued by an Indigenous African community, which requires prior approval from the San Council and commits researchers to fairness, respect, care and honesty (Schroeder et al. 2019; Callaway 2017). Beneath it sits free, prior and informed consent (UNDRIP, 2007, Articles 19 and 32), which here demands that the speaker understand, in Nǀuu, what a synthetic voice is and can be made to do. CARE governs the data lifecycle that follows; OCAP (First Nations Information Governance Centre) supplies a concrete possession model under which the community, not the institution, holds the master recordings and the trained voice; and any academic, cultural or commercial value generated should be the subject of an explicit benefit-sharing agreement, on the model of the San's own Hoodia precedent. The full treatment, with citations, is in the supplement.

5.6 Funding routes (pursued in parallel)

Relevant institutional context: the Bleek–Lloyd collection is a UNESCO Memory of the World inscription, and the Digital Bleek & Lloyd is an established multi-institution digitisation of it.

5.7 Risks and limits

The principal risks, graded as planning judgements (likelihood and impact low / medium / high) rather than measured probabilities:

#RiskLIMitigation
R1Data scarcity for clicks: no click TTS/ASR exists, and click-bearing audio is rare beyond the archived Nǀuu materialHighMedLead with the formant route, which needs measured parameters not a corpus, and for which Nǀuu click acoustics are published (Exter 2011; Miller et al. 2009); treat any neural corpus as a bonus, not a dependency
R2Speaker availability and health: the most authoritative recordings depend on one speaker in her nineties; the window is finiteHighHighRedundancy-first elicitation (most valuable material first); short, speaker-paced sessions; reinforce with other Nǀuu speakers and learners and archived audio; never let the schedule pressure the speaker
R3Neural clicks fail technically: no model has been shown to generate an unseen phoneme class zero-shotMed-HighMedKeep the formant route as the primary deliverable; run an early, cheap empirical test of click rendering before committing neural effort; report a negative result honestly
R4Community consent and trust, against a documented history of extractive research on San communitiesMedHighOperate under the San Code of Research Ethics from the outset; secure free, prior and informed consent in the speaker's language; build the right to pause and withdraw into the workflow
R5Governance and ownership disputes over recordings, model, and productMedMedSettle ownership and licensing before recording on a Kaitiakitanga-style guardianship model; community possesses and stewards masters and weights (after OCAP); document benefit-sharing
R6Sustainability and maintenance: model formats churn, hosting lapses, one-off grants do not fund the long tailMedMedPrefer the formant route's transparent, reproducible parameter tables; deposit recordings and parameters in a durable archive (ELAR; the Digital Bleek and Lloyd); keep the interface loosely coupled
R7Scope and expectation creep towards fluent generative ǀxam, which has no data sourceMedMedState the ceiling plainly to funders and partners: the deliverable is accurate pronunciation of clicks and key names, not generative ǀxam

6. Conclusion

The reading interface makes the words of a sleeping language legible but not its sound, because available speech tools cannot produce its clicks and the click characters are removed before synthesis. No current system can voice ǀxam clicks, and this reflects the absence of training data rather than a configuration that could be adjusted. The components of a solution do, however, exist separately: an acoustic model of Nǀuu clicks validated against Ouma Katrina Esau's speech (Exter 2011), a published phonological reconstruction of ǀxam (du Plessis 2018), and the living phonetics of Nǀuu. What is missing is a programme that combines them. This paper has set out such a programme: community-led, consent-first, and phased from a low-cost glossary control to a grant-scale pronunciation model, so that the archive might eventually be heard as well as read.

The wider point is that the synthesis tooling now exists to give endangered and sleeping languages back their sound. Neural and formant synthesis, of the kind that already renders this archive aloud in English, is no longer the preserve of high-resource languages; applied with the right linguistic and community foundation, it offers a genuine opportunity to restore a pronunciation that no recording captured and no living person can any longer produce unaided. That opportunity is bounded, and the bounds matter: the realistic end state is accurate pronunciation reconstructed from living sister languages, not a synthetic voice narrating freely in ǀxam, for which no data exists. Within those bounds the approach, living phonetics plus linguistic reconstruction plus modern synthesis, is replicable for other languages documented in writing but lost in sound.

Acknowledgements

This work, like the reading interface it extends, builds on Pippa Skotnes's recovery of the Bleek–Lloyd archive: the Digital Bleek & Lloyd and the scholarship around it (Sound from the Thinking Strings, 1991; Claim to the Country, 2007; When the World Was, 2025, among others), without which the corpus would not be legible or reachable.

References

Speech technology and synthesis

Click phonetics

ǀxam, Nǀuu, classification and reconstruction

The archive and its recovery (Pippa Skotnes)

Endangered-language speech-technology precedents (Section 4)

Governance

Appendix: Sourcing notes

Corrections and caveats established during research, reflected throughout:

  1. Exter (2011) is an acoustic analysis, not a synthesiser; it supplies validated parameter targets for two click types (apical and laminal), derived from a deliberately rough articulatory model: useful seed data, not a turnkey table for the full Nǀuu/ǀxam click set.
  2. The "seven click releases" sometimes attributed to ǀxam is a reconstruction inference (from comparison with Kora; du Plessis 2018), not a directly observed inventory.
  3. Several speech-technology figures (e.g. ElevenLabs cloning thresholds, the YourTTS hour total) are search-surfaced quotes of official sources rather than directly re-fetched; treat as high-confidence, not exact.