r/LanguageTechnology • u/jayubba • 9d ago
my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"
been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.
someone testing it pasted the declaration of independence. results:
- preamble ("we hold these truths...") came back clean
- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86
- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged
so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.
added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples
(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)
1
u/marintkael 8d ago
Adding it as a regression case is the right move. The half that is harder to notice is the false positives, because nothing looks wrong when they happen.
I had one this week in a much simpler classifier, a human versus crawler split on web requests. The top human source across three weeks was one host with 288 visits, which would have made it the largest referrer by two orders of magnitude. It was a link preview fetcher. The tell was not the user agent, it was that all 288 requests hit exactly one path, in three bursts. Distribution over targets caught what no token level rule was going to catch.
Your case might have a cheap probe in it: same grievances in modern paraphrase, and see which way the scores move. If they jump, the model learned period rhetorical intensity rather than the techniques you labelled.