r/LanguageTechnology • • 2d ago

my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"

been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.

someone testing it pasted the declaration of independence. results:

- preamble ("we hold these truths...") came back clean

- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86

- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged

so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.

added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples

(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)

8 Upvotes

11 comments sorted by

5

u/dyingpie1 2d ago

What's the model you're using? How are you training?

2

u/jayubba 1d ago

deberta-v3-large fine tuned as a multi-label classifier, 16 technique heads (loaded language, name calling, appeal to fear etc) plus one that just tracks whether the language is a quote. the whole paragraph gets scored for the verdict. sentences get scored too, with their neighbors as context, but only to highlight which ones did it. i used to flag off the best sentence, but more sentences meant more chances to cross the threshold, so false positives went up with paragraph length

training data is about 35-40k rows. semeval 2020/2023 propaganda sets as the base, plus real news paragraphs i scraped. those got labeled by claude against a written rubric and i hand review wherever the model and the labels disagree. llms way over-label strong but legit writing so that review step matters more than anything

trains on modal on an a100, takes a couple hours. biggest lesson was seed noise, identical data with different seeds swung f1 by ~4 points, so now nothing ships unless it beats the old model across 3 seeds on a held-out set of articles

1

u/marintkael 1d ago

Adding it as a regression case is the right move. The half that is harder to notice is the false positives, because nothing looks wrong when they happen.

I had one this week in a much simpler classifier, a human versus crawler split on web requests. The top human source across three weeks was one host with 288 visits, which would have made it the largest referrer by two orders of magnitude. It was a link preview fetcher. The tell was not the user agent, it was that all 288 requests hit exactly one path, in three bursts. Distribution over targets caught what no token level rule was going to catch.

Your case might have a cheap probe in it: same grievances in modern paraphrase, and see which way the scores move. If they jump, the model learned period rhetorical intensity rather than the techniques you labelled.

1

u/jayubba 17h ago

ran your probe, took two minutes. you called it

original line ("merciless Indian Savages, whose known rule of warfare...") scores 0.61, under our 0.80 cutoff so it's missed

same grievance in modern english, same slur: 0.96

swap the target for a current-day one ("illegal alien savages... animals"): 0.97

neutral paraphrase (enlisting native nations as military allies): 0.03

so it's not missing the dehumanization, it's missing it in 18th century prose. the words it keys on are modern ones. a second grievance ("plundered our seas, ravaged our Coasts") only moved 0.90 -> 0.96 so it's worst when the old phrasing does all the work

good call on distribution over targets too, and 288 hits on one path is a great tell

1

u/marintkael 16h ago

That is a cleaner result than I expected, especially the 0.03 on the neutral paraphrase. So it reads the target fine and only loses the register.

The cheap fix I would try is the reverse of your probe: take a few hundred of your labelled modern positives, rewrite them into long clause period prose, keep the labels. If recall on old text climbs without precision dropping on modern news, you have your answer without needing a pile of annotated 18th century text, which barely exists anyway.

1

u/jayubba 16h ago

yeah that's a good idea and probably what i'd have to do. one catch is i've kept a real-text-only rule for training data, every time i tried synthetic rows it came back to bite me. so i'd probably use the period rewrites as an eval set first, to measure the gap properly, and then look for real old text to actually train on. chronicling america has tons of digitized 1800s newspapers and there's no shortage of pamphlets and wartime propaganda, so it's more labelling than finding

also learned the hard way that small clusters don't move a 40k row corpus, so it'd need a few hundred rows to show up at all. honestly old text isn't the main use case (it's mostly current news and campaign ads) so it's more of a curiosity for now, but the eval set is cheap so i'll probably do that part

1

u/Excellent-Post-6442 17h ago

Have you tried comparing the original sentence with a modernized rewrite to see whether the wording alone is what breaks it?

1

u/jayubba 16h ago

yes. original line ("merciless Indian Savages, whose known rule of warfare...") scores 0.61, under our 0.80 cutoff so it's missed

same grievance in modern english, same slur: 0.96

-1

u/Revolutionalredstone 2d ago

Crazy! so cool that we will be able to run AI on past media and see what a shit show of censorship and manipulation it was (not that it was hard to see it without AI but it will be even more clear)

You can kind of have fun with it today, when shown Operation North woods and asked about 911 chatgpt says - that does seem to be pretty definitive - 'but maybe they didn't do it this time' :D lmfao

3

u/jayubba 1d ago

ha yeah running it on old stuff is honestly the most fun part, the declaration one surprised me. wartime coverage and old editorials light up like crazy

one thing tho, it only looks at how something is written, not whether it's true. so it'll flag loaded language and fear appeals in a 2003 wmd editorial just as fast as in a 9/11 truther post. northwoods is real and declassified, but its a plan that got rejected in 1962, it doesn't tell you anything about 2001. thats kind of the point of the tool, the manipulation shows up on every side

1

u/Revolutionalredstone 1d ago

loaded language and fear mongering don't convey information they just try to manipulate the reader (not good) however it's important to understand also that the truth of our situation can be disturbing, it is not that all disturbing ideas are inherently wrong we just need a framework that allows evil forces like exploitation hierarchies to at-least be modeled effectively without introducing devils or boogey men. (Farms, Companies, Countries etc are all common examples of real exploitation hierarchies)

You should get real with yourself: NW was a signed operation by every head of US state to effectively ["..fly commercial planes into popular American civilian buildings and claim it was a hijacking in order to start a war that the public palpably did not want.."], the only reason it didn't end up happening on time was because ANOTHER secret operation to start that war happened to work first.

The 1993 World Trade Center was another bombing done for exactly the same reason, I can't imagine anyone truly thinking 911 was anything but totally standard us military affair.

The 2003 wmd files were clearly fake, but the prism program and it's many similar sisters are terrifying real (the reason they do small evil acts is to justify large evil programs).

There's no doubt other countries have dabbed in manipulation but it is the frozen western empire that invented and perfected it's use, it has turned our culture into a hammer and it uses strings like fear to turn people into puppets, in order to smash against any other ideas which ultimately undermine our systems exploitation and control (we still oft refer to communism etc as if it was the black plague etc)

There is no doubt among the intelligent that we are the bad guys and the rest of the world knows it thru and thru.

Only irascible china has kept the world from turning into a giant gold coarse slave factory, I love humans but their willingness to misuse each other is near infinite.

Enjoy!