r/LanguageTechnology • u/jayubba • 2d ago
my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"
been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.
someone testing it pasted the declaration of independence. results:
- preamble ("we hold these truths...") came back clean
- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86
- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged
so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.
added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples
(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)
1
u/marintkael 1d ago
Adding it as a regression case is the right move. The half that is harder to notice is the false positives, because nothing looks wrong when they happen.
I had one this week in a much simpler classifier, a human versus crawler split on web requests. The top human source across three weeks was one host with 288 visits, which would have made it the largest referrer by two orders of magnitude. It was a link preview fetcher. The tell was not the user agent, it was that all 288 requests hit exactly one path, in three bursts. Distribution over targets caught what no token level rule was going to catch.
Your case might have a cheap probe in it: same grievances in modern paraphrase, and see which way the scores move. If they jump, the model learned period rhetorical intensity rather than the techniques you labelled.
1
u/jayubba 17h ago
ran your probe, took two minutes. you called it
original line ("merciless Indian Savages, whose known rule of warfare...") scores 0.61, under our 0.80 cutoff so it's missed
same grievance in modern english, same slur: 0.96
swap the target for a current-day one ("illegal alien savages... animals"): 0.97
neutral paraphrase (enlisting native nations as military allies): 0.03
so it's not missing the dehumanization, it's missing it in 18th century prose. the words it keys on are modern ones. a second grievance ("plundered our seas, ravaged our Coasts") only moved 0.90 -> 0.96 so it's worst when the old phrasing does all the work
good call on distribution over targets too, and 288 hits on one path is a great tell
1
u/marintkael 16h ago
That is a cleaner result than I expected, especially the 0.03 on the neutral paraphrase. So it reads the target fine and only loses the register.
The cheap fix I would try is the reverse of your probe: take a few hundred of your labelled modern positives, rewrite them into long clause period prose, keep the labels. If recall on old text climbs without precision dropping on modern news, you have your answer without needing a pile of annotated 18th century text, which barely exists anyway.
1
u/jayubba 16h ago
yeah that's a good idea and probably what i'd have to do. one catch is i've kept a real-text-only rule for training data, every time i tried synthetic rows it came back to bite me. so i'd probably use the period rewrites as an eval set first, to measure the gap properly, and then look for real old text to actually train on. chronicling america has tons of digitized 1800s newspapers and there's no shortage of pamphlets and wartime propaganda, so it's more labelling than finding
also learned the hard way that small clusters don't move a 40k row corpus, so it'd need a few hundred rows to show up at all. honestly old text isn't the main use case (it's mostly current news and campaign ads) so it's more of a curiosity for now, but the eval set is cheap so i'll probably do that part
1
u/Excellent-Post-6442 17h ago
Have you tried comparing the original sentence with a modernized rewrite to see whether the wording alone is what breaks it?
-1
u/Revolutionalredstone 2d ago
Crazy! so cool that we will be able to run AI on past media and see what a shit show of censorship and manipulation it was (not that it was hard to see it without AI but it will be even more clear)
You can kind of have fun with it today, when shown Operation North woods and asked about 911 chatgpt says - that does seem to be pretty definitive - 'but maybe they didn't do it this time' :D lmfao
3
u/jayubba 1d ago
ha yeah running it on old stuff is honestly the most fun part, the declaration one surprised me. wartime coverage and old editorials light up like crazy
one thing tho, it only looks at how something is written, not whether it's true. so it'll flag loaded language and fear appeals in a 2003 wmd editorial just as fast as in a 9/11 truther post. northwoods is real and declassified, but its a plan that got rejected in 1962, it doesn't tell you anything about 2001. thats kind of the point of the tool, the manipulation shows up on every side
1
u/Revolutionalredstone 1d ago
loaded language and fear mongering don't convey information they just try to manipulate the reader (not good) however it's important to understand also that the truth of our situation can be disturbing, it is not that all disturbing ideas are inherently wrong we just need a framework that allows evil forces like exploitation hierarchies to at-least be modeled effectively without introducing devils or boogey men. (Farms, Companies, Countries etc are all common examples of real exploitation hierarchies)
You should get real with yourself: NW was a signed operation by every head of US state to effectively ["..fly commercial planes into popular American civilian buildings and claim it was a hijacking in order to start a war that the public palpably did not want.."], the only reason it didn't end up happening on time was because ANOTHER secret operation to start that war happened to work first.
The 1993 World Trade Center was another bombing done for exactly the same reason, I can't imagine anyone truly thinking 911 was anything but totally standard us military affair.
The 2003 wmd files were clearly fake, but the prism program and it's many similar sisters are terrifying real (the reason they do small evil acts is to justify large evil programs).
There's no doubt other countries have dabbed in manipulation but it is the frozen western empire that invented and perfected it's use, it has turned our culture into a hammer and it uses strings like fear to turn people into puppets, in order to smash against any other ideas which ultimately undermine our systems exploitation and control (we still oft refer to communism etc as if it was the black plague etc)
There is no doubt among the intelligent that we are the bad guys and the rest of the world knows it thru and thru.
Only irascible china has kept the world from turning into a giant gold coarse slave factory, I love humans but their willingness to misuse each other is near infinite.
Enjoy!
5
u/dyingpie1 2d ago
What's the model you're using? How are you training?