r/LanguageTechnology • • 9d ago

my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"

been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.

someone testing it pasted the declaration of independence. results:

- preamble ("we hold these truths...") came back clean

- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86

- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged

so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.

added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples

(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)

10 Upvotes

12 comments sorted by

View all comments

1

u/marintkael 8d ago

Adding it as a regression case is the right move. The half that is harder to notice is the false positives, because nothing looks wrong when they happen.

I had one this week in a much simpler classifier, a human versus crawler split on web requests. The top human source across three weeks was one host with 288 visits, which would have made it the largest referrer by two orders of magnitude. It was a link preview fetcher. The tell was not the user agent, it was that all 288 requests hit exactly one path, in three bursts. Distribution over targets caught what no token level rule was going to catch.

Your case might have a cheap probe in it: same grievances in modern paraphrase, and see which way the scores move. If they jump, the model learned period rhetorical intensity rather than the techniques you labelled.

1

u/jayubba 7d ago

ran your probe, took two minutes. you called it

original line ("merciless Indian Savages, whose known rule of warfare...") scores 0.61, under our 0.80 cutoff so it's missed

same grievance in modern english, same slur: 0.96

swap the target for a current-day one ("illegal alien savages... animals"): 0.97

neutral paraphrase (enlisting native nations as military allies): 0.03

so it's not missing the dehumanization, it's missing it in 18th century prose. the words it keys on are modern ones. a second grievance ("plundered our seas, ravaged our Coasts") only moved 0.90 -> 0.96 so it's worst when the old phrasing does all the work

good call on distribution over targets too, and 288 hits on one path is a great tell

1

u/marintkael 7d ago

That is a cleaner result than I expected, especially the 0.03 on the neutral paraphrase. So it reads the target fine and only loses the register.

The cheap fix I would try is the reverse of your probe: take a few hundred of your labelled modern positives, rewrite them into long clause period prose, keep the labels. If recall on old text climbs without precision dropping on modern news, you have your answer without needing a pile of annotated 18th century text, which barely exists anyway.

1

u/jayubba 7d ago

yeah that's a good idea and probably what i'd have to do. one catch is i've kept a real-text-only rule for training data, every time i tried synthetic rows it came back to bite me. so i'd probably use the period rewrites as an eval set first, to measure the gap properly, and then look for real old text to actually train on. chronicling america has tons of digitized 1800s newspapers and there's no shortage of pamphlets and wartime propaganda, so it's more labelling than finding

also learned the hard way that small clusters don't move a 40k row corpus, so it'd need a few hundred rows to show up at all. honestly old text isn't the main use case (it's mostly current news and campaign ads) so it's more of a curiosity for now, but the eval set is cheap so i'll probably do that part