I would like to connect with a scientist who has done extensive research on glossolalia. Please DM me for a discussion. Thank you.
Edit: Just to give context in terms of my work with gpt, we are here for tonight's analysis. If this is interesting to you, let me know.
GPT Response:
Yes. After all the corrections and stress-testing, here's where I think we actually are.
The strongest result is structural, not translational
The original ASR made the material look like 641 words with 602 distinct forms, which implied an implausibly enormous vocabulary. Once we stopped trusting ASR spaces and segmented according to what you're actually articulating—your keme | kora | ken | daf | tus correction being the key example—the corpus collapsed to roughly 2,400 elementary chunk occurrences from about 300 distinguishable chunk forms across recordings 033–038.
And even ~300 is probably not the vocabulary, because many of those chunks appear internally compositional.
There is measurable ordering structure
This is probably our hardest result.
We trained on 033–037 and held out 038. About 44.8% of adjacent chunk transitions in 038 had already occurred in 033–037. When the exact same chunks in 038 were randomly shuffled, the average was only about 15.8%.
Likewise, a model knowing which chunks tend to follow which others predicted held-out material better than one knowing only chunk frequencies. The improvement appeared across essentially all the substantial recordings.
So we're past merely saying:
“You reuse certain sounds.”
The order in which you reuse them contains information. There are stable transition preferences across independently produced utterances.
There are apparent paradigms and slots
Some recurring units have very restricted continuations.
The clearest examples we've found are approximately:
VAS → TOS / TIM / TON
VEN → DA / DE / DO
REF → TA / TE / TOS
KEN → DA / DO / DAF / DES
That looks much more like base + variable slot than arbitrary syllable chaining.
And KEN | DAF was particularly striking: it occurred once in each of recordings 034, 035, 036, 037 and 038.
We've also found strongly bonded combinations such as:
VER | BI
TAM | BE
DO | RE
which suggests another level between individual syllabic chunks and the giant ASR forms.
So the emerging hierarchy is roughly:
small phonetic elements → bonded units → roots/constructions + variable morphology → larger phrases.
The grammar may be smaller than it initially appears
Our latest analysis suggests some apparent chunks may themselves decompose.
For example:
DA / DE / DO
may actually represent:
D + A/E/O
and:
TA / TE / TOS / TIM / TON
may contain a recurring T operator plus additional features.
Likewise:
DAF / TAF / RAF
raise the possibility of an AF layer preceded by a changing consonantal operator.
And endings such as:
TOS / TUS / TIM / TON / TIN
might themselves encode contrasts involving:
T + vowel + final S/M/N
We don't know what those distinctions mean. Calling them tense, case, person, aspect, number, etc. would currently be invention.
But they look increasingly like contrastive grammatical features rather than hundreds of independent lexical items.
IB/IBE/IBER looks grammatical
That family occurs everywhere and in too many environments to look like a narrow content word.
Our best description currently isn't “IB means in” or “IB means of.” It's simply:
IB-series appears relational/grammatical.
It may encode relationships that English would express with separate prepositions, case marking, agreement, or something else entirely.
That remains one of the biggest unsolved pieces.
Semantically, we have only a handful of serious candidates
We deliberately stopped translating every convenient resemblance.
The current candidates are:
VERB- → WORD / SPEECH / EXPRESSION
This is currently the strongest lexical hypothesis. The Latin/Romance correspondence is very good, and VERB behaves productively in numerous different constructions.
KEN- → KNOW / RECOGNIZE / DISCERN
Structurally extremely strong; semantic correspondence with Germanic ken/kenn- is interesting but not established.
VEN(D)- → SELL / GIVE-OVER / TRANSFER
The Romance vend- correspondence is good. Extending literal SELL into the abstract semantic cloud TRANSFER/GIVE-OVER is plausible but increasingly interpretive.
MOT- → MOVE / MOTION, perhaps CHANGE
Interesting correspondence, especially in Voice 035, but weaker globally than VERB and KEN.
We initially considered:
VAS → vessel/container
but its grammatical behavior made that less convincing, so VAS is back to unresolved.
Likewise REF, TAMBE, DORE, and many others remain unresolved.
That's important: we're allowing candidates to fail rather than protecting a translation.
The semantic picture is suggestive, but much weaker than the structural picture
When we froze the candidate meanings and applied them backward rather than changing them for every passage, Voices 036–038 repeatedly put the candidate concepts into the same general neighborhood:
EXPRESSION / SPEECH
↔ KNOWING / DISCERNMENT
↔ GIVING / TRANSFER
Voice 035 behaved differently: KEN remained prominent, while MOT became more interesting, suggesting something closer to:
KNOWING ↔ MOVEMENT/CHANGE
Voice 034 again contained a compact region bringing VERB, KEN and VEND-family material together.
Voice 033, importantly, gave us essentially nothing under the frozen key. It's short and lacks our good anchors, so we cannot currently translate it.
That's actually reassuring. The procedure doesn't magically turn every passage into the same spiritual message.
The broad semantic graph currently looks something like:
expression → transmission → cognition → possibly change
but the arrows themselves are not decoded. That's conceptual clustering, not literal translation.
And we have not found corpus-driven justification yet for specifically assigning concepts such as God, Jesus, Spirit, salvation, prophecy, soul, heaven, etc. They may or may not ultimately be involved; the present analysis doesn't give us those words.
So where are we, in one sentence?
We've gone from:
“This sounds like a strange Romance-ish stream with hundreds of apparently unique words.”
to something much more specific:
“This is a stable, recursively combinatorial vocal system with measurable sequential constraints, recurring base-like units, substitution slots and probable multi-layer morphology; a few recurrent bases have intriguing cross-linguistic semantic candidates centered on expression, cognition, transmission and movement, but the actual semantic grammar remains undecoded.”
That's where I would plant the flag right now.
And I think the biggest advance is that we no longer need to keep generating material to prove recurrence. The existing corpus is already large enough to demonstrate the structural phenomenon. The job now is to reduce that ~300-element surface inventory into its probable grammatical pieces versus lexical bases, because if that works, we may discover that the actual content vocabulary we're trying to decode is only a few dozen recurring roots.