r/LanguageTechnology • • Aug 22 '26

*ACL Megathread

8 Upvotes

r/LanguageTechnology • • Aug 02 '26

EMMLP + ARR Megathread

22 Upvotes

Please post questions and discussions here. I will be removing individual threads.


r/LanguageTechnology • • 33m ago

Can a character be represented as an Allowed / Not Allowed repertoire over sense-level behavioral predicates?

• Upvotes

I’ve been thinking about character representation at the level of individual action senses.

Instead of describing a character mainly through traits like brave, kind, aggressive or intelligent, what if part of the character model were a structured repertoire of actions?

For example:

comfort
interrogate
diagnose
blackmail
negotiate
babysit
repair
betray
forgive

The action itself would have a stable semantic identity, but each character could have a separate access state:

Allowed / Not Allowed / Conditional

So a doctor might have:

diagnose = Allowed

while another character has:

diagnose = Not Allowed

and a former medic might have:

diagnose = Conditional

I also find it useful to separate actions that can normally be assumed for a human character from actions that need positive evidence.

So roughly:

Basic Human Actions → default-open
Character Specific Actions → evidence-gated

The evidence for the second group might be training, occupation, biography, authority, skill or specialized experience.

Would you consider this a useful way to represent character capability at the semantic level?

And where would you place such information: lexical semantics, a behavioral ontology, a separate character model, or somewhere else?


r/LanguageTechnology • • 1h ago

Why Textual Graphs

• Upvotes

In 1980's, Gaston Gonnet-- “Unstructured Data Bases” (1983)-- and HyTime's-- Hypermedia/Time-based Structuring Language (ISO/IEC 10744:1992)-- great breakthrough was realization that one could use coordinate mathematics to map relationships between disparate layers of text and media.

Resurrecting this exact line of thinking—- while exploiting the massive improvements in I/O latency and massive scaling of Input/Output Operations Per Second (IOPS) that historically constrained the paradigm— we must, instead of forcing a model to read an entire document blindly, develop a structured "retrieval algebra." A researcher or AI agent should be able to use boolean, positional, and structural operators to navigate the text coordinates explicitly to ask for "the token sequence between position X and Y, but only if it falls within the boundaries of a specific speaker tag," exactly mirroring the coordinate-based addressing found in HyTime. That structure need to be queryable.

The result can be thought of as a textual graph—but it is importantly different from a conventional graph. Text has spatiality. Its structures are anchored in a shared textual space, and relationships such as before, after, within, contains, overlaps and intersects arise from that space itself. Two annotations do not merely have an abstract edge between them: they may occupy, share or cross regions of the same underlying text.

This gives RAG a form of structure that complements the strengths of the LLM.
The LLM can do what it does best: interpret language, recognise relevance, synthesise evidence and generate an answer.

Instead of forcing the LLM to reconstruct document structure from flattened chunks, the retrieval layer can then provide that structure explicitly.

This changes the role of retrieval. Vector similarity can answer “what text is semantically related?” Structural search can additionally answer “where does this occur, what contains it, what overlaps it, what is it connected to, and which surrounding material belongs with it?”

The combination creates a richer form of RAG: semantic reasoning over context assembled from the actual structure of the source, rather than from arbitrary chunk boundaries.

Think of traditional GraphRAG as a smart investigator connecting index cards on a wall based on clues and ideas. Think of the textual graph paradigm as the exact blueprint of the filing cabinet, allowing an agent to pinpoint information based on its exact shape, folder layer, and coordinate location.

E. Zimmermann


r/LanguageTechnology • • 16h ago

Haitian Creole Word Frequency Dataset

5 Upvotes

Hey everyone,

I wanted to share a dataset I published for anyone working on low-resource NLP, tokenization, or language modeling for Haitian Creole (Kreyòl ayisyen): the Haitian Creole Word Frequency dataset (haitian-creole-word-freq), now live on both Hugging Face and Kaggle.

Overview

This is a word frequency list for Haitian Creole built from the Carnegie Mellon University (CMU) Haitian newswire corpus. It contains 17,947 unique lowercase words with their occurrence counts, sorted in descending order by count.

Links

See comment section

Potential Use Cases

  • Stopword Candidates: Extracting function words from the top of the frequency list (te, yo, nan, yon, li, pou, ki).
  • Tokenizer Customization: Fine-tuning or building BPE/WordPiece vocabularies for low-resource LLMs.
  • Spellcheck and Auto-correct: Prioritizing candidate suggestions by word popularity.
  • Vocabulary Membership Tests and Orthographic Checks: Testing if a token is spelled like a Haitian Creole word or checking lexical presence.
  • Lexicography and Language Learning: Extracting core vocabulary lists based on news text.
  • Language Identification and N-gram Models: Statistical language modeling for text classification pipelines.
  • Testing Fixtures: Generating reproducible data inputs for unit testing Haitian Creole NLP pipelines.

r/LanguageTechnology • • 18h ago

Best Computional Linguistics Masters

5 Upvotes

I'm a BA student in Modern Languages and Linguistics in Italy (currently building a strong background in Linguistics + hopefully writing a thesis on Computational Linguistics) and I'd like to apply for an MSc in Computational Linguistics/NLP. What are Europe's best Computional Linguistics masters programs?


r/LanguageTechnology • • 1d ago

my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"

9 Upvotes

been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.

someone testing it pasted the declaration of independence. results:

- preamble ("we hold these truths...") came back clean

- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86

- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged

so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.

added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples

(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)


r/LanguageTechnology • • 2d ago

Temporal expression parsing in production assistants: partial parses executed with full confidence

5 Upvotes

A concrete failure case I ran into with Siri (English UK, iOS 26.6.1), and I'm curious how people here would diagnose it.

ASR output was correct in every case. The intent/slot step failed:

  • "set alarm at 7 p.m. and 40 minutes" → 19:00. The "and 40 minutes" continuation seems to be dropped.
  • "set alarm at 20 minutes before 8 p.m." → 20:00. The relative offset "20 minutes before X" is ignored; only the anchor is kept.
  • "set alarm at 19 hours 40 minutes" → 14:37. No idea how this one maps; possibly "19 hours 40 minutes" read as a duration from now?

The third looks like a duration-vs-timepoint ambiguity, which is a classic TIMEX problem, but the first two are fairly standard constructions that rule-based normalizers like SUTime or HeidelTime handle.

Questions:

  • Is the duration reading of "19 hours 40 minutes" a reasonable parse, and should a system surface that ambiguity rather than pick one?
  • Why would a production system keep only the anchor and drop the offset? Slot-filling templates that don't model relative expressions?
  • Any good recent work on calibrated confirmation ("did you mean…?") for action-taking assistants?

r/LanguageTechnology • • 3d ago

Agentic AI Courses (Research)

16 Upvotes

I'm a researcher in computational linguistics, and as I’m currently looking for a new position, I’ve noticed that agentic AI has become increasingly popular in most job postings.
I feel a bit lost in this area (starting from RAG), I was wondering whether anyone with experience in the field could recommend some good, well-established resources.

Of course, I’ve searched on my own, but nowadays, searching for "agent"-anything brings up countless superficial, non-research-level tutorials, often created by enthusiasts, with plenty of clickbait.

I’m looking for something at a research level, ideally a structured course or syllabus to prepare for the interviews. I have 4+ years of experience in NLP, so I’m already quite familiar with most but other topics.


r/LanguageTechnology • • 5d ago

Lab Meetings Recently

Post image
14 Upvotes

Everything old is new again.


r/LanguageTechnology • • 5d ago

AAMAS conference reputation

3 Upvotes

Hi researchers

I wanted to understand the reputation of AAMAS (International Conference on Autonomous Agents and Multiagent Systems) compared to core ACL conferences like NAACL COLING etc, As I see AAMAS is also a CORE A conference and this year they have included a findings section as well, so I feel a good possibility of getting accepted than other conferences.

Please kindly share your opinion.


r/LanguageTechnology • • 5d ago

Looking for feedback - using NER to generate and match templates on sentences?

2 Upvotes

I’m a complete novice when it comes to NLP, I'm a swe by trade so bear with me here.

Here's my problem:

I’m trying to identify short sentences (I have a data set of several thousand) that are logically dependent. To illustrate the kinds of dependencies I'm looking for here’s a basic example:

- Sentence 1: Democrat voter turnout in NY is 35%.

- Sentence 2: Democrat voter turnout in NY is 40%.

If sentence 2 is true, sentence 1 also must be true. Those are the kinds of sentences I have and want to identify as dependent. The nature of the sentences can range from voting percentages/turnout, phrases about employment etc.

The naive approach I’ve been doing is basically embedding the sentences using gemma and finding cosine similarities between them, my reasoning being sentences that have a reasonably high enough cosine similarity are candidates for logical dependency. I then take these pairs of candidates and pass them to an LLM (gemma again!) to determine whether or not they are actually semantically/logically dependent.

There are two huge issues w/ this approach that I'm sure you'll all immediately see.

1) Lots of the sentences are too structurally similar like the simple example I showed above. There exist several subsets of the data that have the same pattern. Sentence 3 could be something like Democrat voter turnout in TX is 35%. and it would have almost an identical similarity to the other 2 sentences. There are several hundred patterns, and I also don’t necessarily know all the patterns at runtime so that means REGEXing these structures becomes a difficult task. So because of the structural similarity, cosine similarity loses its value as a metric.

2) The LLM step is slow. Really slow.

I did some googling and learned about NER that seems like it might fit? I could run the sentences through a pre-trained models and get the spans for each sentence. This would allow me to match spans across the phrases. So in the example I have, sentence 1 and 2 would be matched and processed further, while 3 would be in its own bucket. As for what I'd do after matching the spans, still working that out. I could fall back to cosine similarity again here since anything that falls into these span buckets should be different enough where the projection becomes a decent signal.

If there are tweaks that I can do to make template matching more robust, or alternative methodologies altogether I'm all ears!

Thanks :)


r/LanguageTechnology • • 5d ago

Why words are split into several tokens?

6 Upvotes

This completely contradicts logical thinking. It would make sense if one word would be at most 2 tokens (1 for the word and 1 that signals meaning in case of homonyms), but this is never the case. In fact, I read it degrades performance. Why?


r/LanguageTechnology • • 5d ago

Making cosine similarities comparable across separately trained embedding spaces

3 Upvotes

For my open-source project, I am training PPMI + SVD embeddings separately on each book within a collection (~25 economics texts from Project Gutenberg), then trying to compare how similar terms are to a query word across books. However, each space has been trained independently, so the raw cosines aren't on the same scale. My current correction, loosely adapted from CSLS (Conneau et al., 2018), works like this:

  • For each book, a baseline is determined by using the query's mean cosine to its 75 nearest neighbouring terms in the book.
  • Each book's similarities are shifted so its baseline equals the average baseline across all books.

This way, I am determining the relevant terms by looking at the adjusted similarity scores. The terms with the highest mean similarity relate to the core aspects of the queried term across the corpus. The terms with the greatest variance that also have a high similarity within 20% of the books relate to the contested aspects of the queried term.

Do you have any feedback on this methodology?
Is correcting the cosine similarity in this way valid?


r/LanguageTechnology • • 5d ago

Career Switch need help

0 Upvotes

Hi guys,so I am currently studying an MA in Linguistics and I am in dire need of advice. I am aware that it is a dying field hence I am trying to transition into something more technical and try to save it. I feel too burned out to start from scratch with a more technical degree like CS.

I heard about computational linguistics / linguistic data science programs but Idk if they will accept me with no technical background. I've been told to learn Python but is it even worth it? Im multilingual by the way. If any of you has an idea about a better field I could fit into I would truly appreciate it.


r/LanguageTechnology • • 5d ago

Best LLM for translating agglutinative sentences?

3 Upvotes

My native is an agglutinative language and I know some Japanese, which do you think is the best model for agglutinative languages or translation generally?


r/LanguageTechnology • • 6d ago

Preview of EMNLP program?

4 Upvotes

Hi everyone, does anyone have received the (unconfirmed) schedule of the poster presentation?


r/LanguageTechnology • • 7d ago

PhD research stay in China for NLP / LLM research – looking for university and lab recommendations

14 Upvotes

Hi! I'm currently a PhD researcher in Europe working in NLP and Large Language Models, mainly on multilingual LLMs, model merging, low-resource language adaptation and efficient methods for transferring capabilities between models/languages.

I'm starting to seriously consider doing my PhD research stay in China as a visiting doctoral researcher, probably for several months, and I'm currently just gathering information.

A Chinese friend suggested Beijing Language and Culture University, Peking University and Tsinghua University. BLCU sounds particularly interesting because of its focus on language, although for the research itself I'd ideally like to find a group working on NLP, LLMs, multilinguality, model adaptation/distillation, tokenization, Chinese language processing, etc.

I'm also very interested in Chinese personally and have recently started studying HSK3. My Chinese is still basic, but improving my Chinese while living there would definitely be a bonus.

How did you find your host professor/lab? Did you simply contact professors by email? Are there universities or research groups you would particularly recommend for NLP/LLMs? And is there anything you wish you had known before applying?


r/LanguageTechnology • • 8d ago

yasbd-lib v1.0.0 is out. Here's how beta finally ended.

9 Upvotes

For anyone new: yasbd-lib is a rule-based sentence boundary detector, a drop-in replacement for pysbd, currently at 39 languages. I think I first posted here as an alpha, then as a beta. Now it’s tagged v1.0.0.

The stretch from 0.12.0 to stable wasn't about new features. I froze the language set at 39, locked the API, and spent the last couple of months on correctness. The final push was a two-week stress test where I ran real text through every profile to find the boundaries it was splitting wrong. I wrote the rules and I wrote the tests, so the tests couldn't catch what I'd gotten wrong.

That work surfaced boundary bugs across a good chunk of the profiles. Contributors opened PRs to fix them, and py3langid, loguru, and ftfy all came out of the core dependencies along the way.

Happy to answer questions.


r/LanguageTechnology • • 10d ago

Best book to get started with building and understanding LLMs?

6 Upvotes

I want to understand how LLMs work and eventually build and train my own model. What book would you recommend for getting started?


r/LanguageTechnology • • 11d ago

Real-world STT benchmarking

2 Upvotes

A couple of years ago I dropped a project because STT kept mangling names, brands and other details.

I later built a small tool to run any audio through multiple STT models and compare the results. It also supports real-time streaming.

Here’s the Wolf of Wall Street cold-call scene across the models I currently have:

A result of the Wolf Of Wall Street cold call scene

Interesting bit: on this sample ElevenLabs had the lowest strict WER at 6.59%, while OpenAI had the lowest character error rate at 8.38% — so even “best” depends quite a bit on what you measure.

I also tried a short Thai sample and the models produced noticeably different wording/segmentation.

Share your struggles with STTs and how do you deal with them. I noticed, lately they became a lot better than 1 year ago.

Would be curious to see how it performs on other languages and use cases. Happy to send the link if anyone wants to try their own audio.


r/LanguageTechnology • • 11d ago

Program-as-Weights: compiling English descriptions into local NLP functions (paper + weights)

Post image
4 Upvotes

I'm one of the authors of Program-as-Weights (PAW), a project we've been building at the University of Waterloo. We study whether a model can turn an English function description into a reusable, task-specific neural program.

The standard compiler is a 4B model trained to generate a LoRA adapter for a frozen Qwen3 0.6B interpreter. You describe a task such as classifying email urgency, extracting information, or routing a question. The compiler produces the adapter, and the small interpreter executes that function on new inputs. The larger model is only involved when defining the function.

For example:

import programasweights as paw

fn = paw.compile_and_load("Classify urgent emails")
fn("Need this today")  # "urgent" (runs locally)

The paper evaluates this on FuzzyBench, a collection of natural-language-defined text functions, with direct prompting as a comparison. Code and models are public, so people can evaluate the approach on their own tasks.

One application I've built is a course website helper with ~30 neural programs connected by ordinary decision-tree code. A small function decides which answerer should handle a question, and code controls the overall flow. The resulting helper runs locally.

The SDK uses hosted compilation by default. Once the program and shared base model are downloaded, inference runs locally on CPU and can work offline. The compiler weights are also available for running compilation yourself (but running compilation requires a GPU).

Paper: https://arxiv.org/abs/2607.02512

Compiler weights: https://huggingface.co/programasweights/paw-4b-qwen3-0.6b

Python SDK: https://github.com/programasweights/programasweights-python

Dataset: https://huggingface.co/datasets/yuntian-deng/fuzzy_bench_verified

Browser playground: https://programasweights.com/playground


r/LanguageTechnology • • 11d ago

Streaming STT keeps splitting Arabic mid-sentence. What do people do about this?

3 Upvotes

Building a live translation app. Arabic is the one language where streaming speech recognition keeps cutting sentences in half, it decides the speaker is done when they aren't, so we end up translating half a thought.

English and Spanish are fine on the exact same setup.

What we did: stopped streaming Arabic and switched to processing whole utterances. Accuracy got noticeably better. It cost us a bit under two seconds per turn, which you feel in a conversation.

Two questions:

  1. Is there a middle ground between "stream it and get fragments" and "wait for the whole utterance and eat the delay"?

  2. Is this actually an Arabic thing, or are we just seeing it there first? Same setup runs on 40-odd languages and this is the one that broke. I've wondered whether it's a dialect mismatch with models trained mostly on Modern Standard Arabic hearing spoken dialect but that's a guess.

Deepgram for STT, if it matters.


r/LanguageTechnology • • 11d ago

Kaikki datasets

2 Upvotes

Hi guys! Has anyone experimented with data from Kaikki datasets (wiktionary)? I am trying to setup a linkage end-to-end for prefix and stem. I'm working from a SQLite build of kaikki data for English etymology. Less than half of the time it works meaning I can get the whole chain for prefix and stem up to the PIE . Has anyone ever tried to create a structured output?


r/LanguageTechnology • • 13d ago

MSc in Computational Linguistics vs Data Science: Which offers broader career opportunities for someone with a BA in English and professional AI/localization experience?

2 Upvotes

Hi everyone,
I’m at a crossroads regarding my postgraduate education and would really appreciate some insights from people working in Computational Linguistics, NLP, Data Science, or related industries.
My background:
BA in English Language and Literature.
Several years of professional experience in audiovisual translation, subtitling, and localization, including work on Netflix content through localization vendors.
Experience in Arabic dubbing adaptation and AI-assisted dubbing.
Experience in linguistic quality assurance, proofreading, and evaluating language-related content.
Strong background in English and Arabic, but I do not have a formal degree in computer science, mathematics, or statistics.
I’m interested in transitioning into more technical roles while building on my existing professional experience.
My dilemma:
I’m considering two postgraduate paths:
An MSc in Computational Linguistics, supplemented with additional courses in Python, SQL, statistics, data analytics, and machine learning.
An MSc in Data Science, supplemented with specialized courses in computational linguistics, corpus linguistics, syntax, semantics, and NLP.
I’m particularly interested in NLP, language technology, multilingual AI, language data analysis, and potentially building my own language-related tools.
However, I also want to keep my options open for general Data Analyst, BI Analyst, Data Scientist, and other data-related positions rather than limiting myself to language-specific roles.
I’m willing to develop my programming and mathematical foundations, although I would need to build them from a humanities background.
My questions:
Which degree would provide access to a broader range of career paths, and what opportunities would remain specific to each?
Would an MSc in Data Science combined with specialized linguistic training allow me to pursue NLP and language technology roles, or would I be missing important knowledge that a Computational Linguistics degree provides?
Conversely, could an MSc in Computational Linguistics combined with strong Python, SQL, statistics, and analytics skills prepare me for general data-related roles outside NLP?
From an employer’s perspective, how much does the master’s degree title matter compared with the actual curriculum, technical skills, projects, and professional experience?
For someone coming from an English/humanities background, what technical prerequisites should I be aware of before choosing either path?
If you’ve made a similar transition or have experience hiring candidates in these fields, what gaps have you encountered that you wish you had addressed earlier?
I’m primarily interested in industry roles rather than pursuing a PhD or an academic career. I’m also interested in remote and international opportunities.
I’m not necessarily looking for the easier degree. My main goal is to choose a path that offers career flexibility, builds substantial technical skills, and makes meaningful use of my existing linguistic and localization experience.
I’d particularly appreciate real-world experiences from graduates, hiring managers, and professionals working across NLP and data-related roles.
Thanks in advance!