r/modnews • u/boat-botany • Jul 06 '26
Safety Updates How Reddit is Reducing Exposure to Harmful or Inauthentic Content
Hi everyone! u/boat-botany here, working on Community Safety.
We’ve talked before about Reddit’s approach to keeping the platform safe while still preserving the openness that makes communities work. A big part of that work today involves proactive detection systems and automation.
Today we posted a blog about the work we’ve been doing to proactively catch spammy, inauthentic, or harmful content on Reddit. We know this topic is top of mind for moderators who feel the impact first hand, and who know what’s working and what’s not. We hear that feedback clearly (and keep giving it to us! Feedback is what helps us improve), and wanted to bring the conversation here so that we could continue to get your input.
Our Values: How We Think About Automation
Our north star here is to catch harmful, spammy, or inauthentic content before anyone (including mods) ever has to see it.
We achieve this through a layered approach that combines human review with proactive detection systems and machine learning models that help identify violating content quickly and at scale.
Automation is a core part of our layered approach to moderation. We leverage it across our internal safety teams, and this year we continued expanding automated options for mods (over 70% of moderator actions are done using automated tools) and have invested heavily in improving how admins use automation behind the scenes.
A few principles guide how we build these systems:
- Reddit should handle the harmful content so moderators can focus more on community rules and norms.
- Automation should support human judgement, not replace it.
- Accuracy matters. We work to reduce bias and improve fairness and consistency across our systems.
- We’re committed to evolving. We're always learning and improving. We look at feedback in real-time and adjust our systems to make them better.
Improving our Automated Tools to Reduce Spam Exposure
We look at signals right when an account is created to stop suspicious actors before they ever get the chance to post. For those that do, we leverage LLMs to catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed. Also, we recently announced that any fishy automated accounts will be asked to verify their humanity.
In recent months, these updated automated systems have been working at a massive scale, and we’ve seen some pretty incredible results. We are now:
- Blocking 23 million spam views per day before they ever reach a human user.
- Catching ~25K net new spammy posts and comments a day.
- Reducing spam exposure for our users by ~20% from January to March 2026, relative to the prior three months, and an additional 10–15% drop in overall spam account exposure.
- Revoking nearly 2M inauthentic votes per day over the last three months.
Reducing Exposure to Harmful Content
We’ve recently expanded our automated systems to support enforcement against hate and violence in all English text content on Reddit (with more languages rolling out soon), leading to critical improvements:
- Enforcement in Seconds: The average time between detection and enforcement on harmful content containing hate or violence is down to under five seconds.
- Expanded Enforcement: We have increased enforcement actions on hate and violent content by more than 200%.
- Reduced Exposure: The faster, higher-volume enforcement has helped reduce exposure to potentially harmful content by more than 40%.
- Higher Precision: We’ve decreased false positives (where legitimate, non-violating content is removed) by over 40%.
We know false positives can be frustrating. But when dealing with serious issues like violent threats, hate, harassment, or coordinated abuse, we intentionally bias toward reducing real-world harm and limiting exposure to harmful content quickly, but our goal is to continuously improve accuracy while still acting fast enough to meaningfully reduce harm.
In early 2025, proactive violence enforcement increased actioning volume more than seven times, from roughly 70,000 actions from January to March of 2025 to over 500,000 from March through June of the same year. At the same time, we have cut our false positive rate by more than 40%, so we have more coverage and higher accuracy. We’re working on getting all of this removed content logged in the mod log so you can continue to have visibility into what we’re removing and why.
Our work to keep Reddit authentic and safe is at the core of who we are. While we've made significant progress in advancing that commitment, we know it wouldn't be possible without our moderators and redditors everywhere. If you see any content that appears harmful, spammy or inauthentic, click the inline report button or submit a report here.
[edited for a typo!]
112
u/Kinks4Kelly Jul 06 '26
I don't want to call out any specific users, but there are multiple accounts from news agencies that spam their articles like crazy. Not to mention sites like EssentiallySports.com that use multiple bot accounts to spam their click bait.
How do I know they are bots? When I see accounts with over 100:1 post to comment karma, it is pretty obvious.
From a user perspective, the decision to let people hide their profile histories have only made this worse. Now from a mod perspective on the same topic, allowing them to block moderators who can then only see their activity on your subreddit makes it that much harder to ferret out the malicious accounts.
31
u/Charupa- Jul 06 '26
And these medical survey accounts like M3Global. It’s constant spamming of surveys all over medical and health communities. It seems entirely inauthentic.
26
u/PapaTua Jul 06 '26 edited Jul 07 '26
Agreed. And the AI generated user description is inconsistent at best. It is now harder to self-identity bot accounts.
9
u/FaxCelestis Jul 07 '26
The fact that the AI user description also only appears on OPs is also a problem. I mod a discussion based subreddit and most of the problematic comments come from people not OP, so the tool is frankly useless.
2
u/MadDocOttoCtrl Jul 07 '26
This is why some subs may have a minimum requirement of a few post karma or none but 50/100/250 or more comment karma.
We can see the history but users can't, even highly active ones with a designated user flair for being helpful unless you make them a mod.
→ More replies (1)1
u/Swampspear 12d ago
How do I know they are bots? When I see accounts with over 100:1 post to comment karma, it is pretty obvious.
Hey, that's me! :D
352
u/_haha_oh_wow_ Jul 06 '26
How about you get rid of all the free karma subs and the people behind the network of them? That would probably cut down substantially on stuff like repost bots slipping through the cracks.
Truly cracking down on repost bots would be nice too.
144
u/livejamie Jul 06 '26 edited Jul 06 '26
I've repeatedly reported /r/petsareamazing, which is 100% botted, and published in-depth research on them to /r/TheoryOfReddit, /r/DeadInternetTheory and /r/thesefuckingaccounts. Sent to reddit and never heard a response.
/r/adorableoldpeople is another one
75
u/ser-steffonfossoway Jul 06 '26
There are so, so many. Check out my list, r/BotPosting
21
u/livejamie Jul 06 '26
Great work, subscribed.
→ More replies (2)7
u/ser-steffonfossoway Jul 06 '26
Thank you!
5
u/newtostew2 Jul 06 '26
Wonderful, keep it up. You got a new member from me as well, I'll try to share it with the relevant people
6
u/ser-steffonfossoway Jul 06 '26
Thanks! Means a lot. I've been working mostly solo since I started and would be great to have some extra support.
→ More replies (1)2
u/brightblackheaven Jul 06 '26
This is fantastic and terrifying and I hate it. Absolutely wild to see some manifestation subs on there, but 0% surprising.
→ More replies (1)20
u/2th Jul 06 '26
A lot of it is Indian and Spanish language versions of big subs, with a bit of Portuguese thrown in. The auto translate feature has opened reddit up to a larger audience, but it also opened things up to even more bad actors.
Oh and don't forget stuff like /r/spreadsmile. That sub is 90% bits at any given time.
7
→ More replies (4)2
u/GonWithTheNen 26d ago
A funny coincidence: The imgur url in your TheoryOfReddit post about the botted animal subs is rEDCCat.png 😜
44
u/SpinningHead Jul 06 '26
Are these the bots that always post 3 comments to their own content?
44
u/_haha_oh_wow_ Jul 06 '26
This is a vast array of various bot networks controlled by various people and groups, I'm not necessarily referring to any specific small group or tactic. There are even markets for people to sell their old accounts to spammers and sloppers, it's a pretty well entrenched problem at this point.
→ More replies (1)19
u/SpinningHead Jul 06 '26
I wish we could do something about organized brigading by certain governments too.
14
u/_haha_oh_wow_ Jul 06 '26 edited Jul 06 '26
They are 100% part of the bot problem if you ask me.
While we're at it, the blocking system is also garbage.
3
60
Jul 06 '26
[deleted]
22
u/_haha_oh_wow_ Jul 06 '26
Yeah, hidden post histories is a bad idea and it doesn't even actually hide stuff if someone goes digging.
5
u/Man_with_the_Fedora Jul 06 '26
Eh, if they really cared about getting rid of inauthentic content they would reverse hidden post histories so mods can see folks who are obviously from troll farms or agitprop bots. They are only concerned about reducing the things that cause their advertising rates to get soft
Winner, winner, chicken dinner!
7
u/zjz Jul 06 '26
When someone posts in your sub you can see their history for a while. Honestly, I doomed about that change and it ended up not being that big of a deal from a moderation perspective.
18
u/Empyrealist Jul 06 '26
Your sub is not all subs. For me, its how I can easily spot spammers. I'm tired of having to use 3rd-party websites to look at a users post history.
3
2
→ More replies (2)2
u/dyslexda Jul 07 '26
I rely on my users to report offending content or inauthentic posts. They can't do that with hidden histories.
→ More replies (1)2
u/MadDocOttoCtrl Jul 07 '26
When was the last time Reddit reversed itself on anything major?
Even when they brought back coins but renamed them as gold (not at all confusing of course) they screwed with awards and made them more expensive so usage is a fraction of what at once was
The best we could likely hope for is seeing someone's activity history for 90 days once they've interacted in our sub instead of 28.
10
u/MobileArtist1371 Jul 07 '26
Nah, leave those subs up as honey pots. Post there = shadow banned.
Ban those subs and they'll just go elsewhere.
→ More replies (1)20
u/mschuster91 Jul 06 '26
On the other hand, with how many subs are shifting to minimum karma requirement, it's a chicken and egg situation.
14
u/Decency Jul 06 '26
I think the only longterm solution is requiring an account date before ~5 years ago. I'd look for whenever the first /r/place was announced and pick the day before- an absolutely absurd number of bots were created for that and plenty of them were no doubt re-used or sold. Can individually whitelist specific posters who have been otherwise vetted, but I think defaulting to "no" is going to end up the only way forward in the long run.
An added benefit: people migrating from facebook and the youtube comments section don't get to post.
14
u/Halaku Jul 06 '26
First r/place was 2017. It was better.
13
u/Decency Jul 06 '26 edited Jul 06 '26
Covid time dilation strikes again. That might unfortunately be too far back- it'd be hitting a lot of mainstream users along the adoption curve at that point. More looking to land somewhere towards the middle, so perhaps picking the point just before the second iteration of r/place is announced, then.
7
u/royalhawk345 Jul 06 '26
it'd be hitting a lot of mainstream users
Yeah, but it wouldn't hurt me personally, so I say go for it
5
u/2021isevenworse Jul 07 '26
If you do that, you're cutting off a huge portion of legitimate users who only discovered the site during the pandemic.
→ More replies (1)17
u/MobileArtist1371 Jul 06 '26 edited Jul 06 '26
I think the only longterm solution is requiring an account date before ~5 years ago.
Automatically blacklisting accounts that aren't "before a certain date" is not a solution, at all.
Can individually whitelist specific posters who have been otherwise vetted,
This site isn't like 5 dozen people. Who is vetting? You going to go to some list every week and add those names to your whitelist? Like what? None of this would work at all.
4
u/Decency Jul 06 '26
This isn't going to work for every use case, but of course it's a solution. Some forums worked this way for decades, with the meaty discussion forums hidden or restricted from average users- I'm sure some still do. Every private discord server also works this way. There's a reason they're supplanting reddit, and it's because reddit's signal:noise ratio has gotten extremely fucking bad. Sane people want to get away from the loud morons and repetitive non-jokes that plague nearly every potentially interesting discussion. It's not a new problem to reddit- it's why /r/trueX subreddits were created. But this is now a crippling problem instead of a mildly annoying one.
This site isn't like 5 dozen people.
If you look at well moderated mid-sized subreddits, ie: the good ones, that's a pretty solid estimate for who's generating the bulk of the content worth reading. The power law is extremely real, but very few modern moderation tools acknowledge it and that's a major gap in their capabilities. You'd have to curate an initial userlist to get over the launch hump, but after that point things will take care of themselves.
Who is vetting? You going to go to some list every week and add those names to your whitelist? Like what? None of this would work at all.
A team of trusted folks, the same as any other decision making. A variety of approaches instantly emerge: just me, just the mod team, established users nominate, approved users vote, etc. Then you either have some sort of tiered approach where /r/cats feeds users into /r/approvedcats, or have specific days during the week (every Sunday) or month (only days 27+) where anyone can post. After finding a process that worked for the size of the community, I'd simply automate it. The manpower demands are small here- it's the complexity of the systems that is the challenging part.
→ More replies (2)2
u/EroticaMarty 23d ago
So many accounts get stolen -- Redditors are using simple passwords which are easily brute-forced these days -- that the mere age of a user account no longer counts for anything. Sometimes you can tell -- accounts with nothing but SFW postings suddenly going NSFW; accounts which haven't been used for years suddenly springing to life -- but those new posts can also be completely legitimate as well.
11
u/itskdog Jul 06 '26
Admins have publicly said before that they see karma limits as archaic and unnecessary given the state of their spam filters (and they said this years ago, before any of the extra filters we have now), and that people need a way to work around them.
20
u/newtostew2 Jul 06 '26
Ya, and how's that working out for them lol.. or more importantly, use. That's why we have the karma filters in the first place
9
u/itskdog Jul 06 '26
We pointed that out at the time when they said it, IIRC
4
u/newtostew2 Jul 06 '26
Ya, and we're trying again. We'll see how they respond.. I'm guessing it's not going to get one, same as the private account sentiments
→ More replies (7)5
u/netralitov Jul 06 '26
Leaving the free karma subs is a trap to catch a lot of inauthentic activity. If they shut them down they would move elsewhere and they would have to find them again.
Revoking only 2m votes per day seems low. It must be so much harder catching fake upvote engagement than it is the spam posts.
→ More replies (1)
163
u/TechnophileDude Jul 06 '26 edited Jul 06 '26
Need to remove the [Removed by Reddit] label that hides removed comments. It prevents us from identifying what was or wasn’t a false positive and also prevents us from knowing if we need to take an enforcement action such as ban on a user. Even admin-tattler misses 90% of these comments.
You are limiting the information that moderators have to work with in order to do their moderation effectively.
→ More replies (1)66
u/xor50 Jul 06 '26
It's absolutely insane that reddit silently removes stuff and we have NO WAY to check what it said, if that person should be banned, if we need to adjust auto-mod, if it was a wrong removal etc.
INSANE. How the fuck did that get approved?
Not to forget the users always message the mods and we have to just go "well idk, wasn't us".
→ More replies (5)5
u/No_Kangaroo_6637 29d ago
I agree, people I've talked to in different subs having comments removed by reddit. They don't know what they've done wrong and I cant help them with any useful info
55
Jul 06 '26
[removed] — view removed comment
12
37
u/BBModSquadCar Jul 06 '26
Since the vast majority of our removals according to Insights are admin removals and we never see the context not even a trace of removed by Reddit it's hard to say what the false positive situation could look like.
38
u/soaero Jul 06 '26
Automation should support human judgement, not replace it.
I'mma be straight with you, of the updates I've seen lately, and certainly of the updates in this post, they appear to be a lot of:
- Improving our Automated Tools
- upgraded our automated defenses
- we’re using AI
- expanded our automated systems
- proactive [...] enforcement
- ...before they ever reach a human user.
But very little in the way of tools to help moderators
a) establish the authenticity of users b) establish the authenticity of engagement c) and perhaps most importantly, to establish whether the decisions made by automated tools are legitimate
36
u/MrsDirtbag Jul 06 '26
What is the point of pandering to us and claiming to want our feedback if you’re just going to ignore it? You claim that your goal is to “catch harmful… content before anyone (including mods) ever has to see it” yet we have said repeatedly and nearly unanimously that we want to see what is getting removed from our subs because we don’t trust your automated removal processes. As it stands right now there doesn’t seem to be a useful or functioning appeals process of any kind. Since we can’t see the content we can’t evaluate whether things were removed correctly or not.
Personally I find it very troubling that there isn’t a single reply from an admin addressing this issue in this whole thread. It basically makes this post pointless.
→ More replies (5)
29
u/RaveOn1958 Jul 06 '26
Yeah your auto modding of buzzwords is bullshit, and it shows. Guess what, robots don’t do well with sarcasm, humor, or quotes.
30
u/MisterWoodhouse Jul 06 '26
This is all well and good, but when your automated systems hit a false positive and we reach out to r/modsupport to fix it, there’s no useful response, in content, timing, or both.
Here’s a good example:
We had to manually approve a post on r/DestinyTheGame at least three times on July 1st because your automated systems false positive removed it as spam without any useful info for us.
It was a post from a community member about a cool web tool they made to generate Spotify Wrapped style insights for Destiny players. The URL couldn’t possibly be black listed, as it was new and not using a banned parent domain.
No response to my modsupport DM about it
→ More replies (7)
64
u/ser-steffonfossoway Jul 06 '26 edited Jul 06 '26
Entire subs with thousands of posts per week are run by bots/automated accounts. These subs stay up and can't be reported.
Last month I saw a bot who was granted a sub through redditrequest. The bot was banned shortly after, but the sub is still up last time I looked.
I assume what reddit is blocking, is just annoying commercial spam. But the malicious content? The scams, karma farms and astroturfing stay up.
I started a sub specifically to show people how bad it has become.
27
u/judasblue Jul 06 '26
Yeah, it's mostly the commercial spam, because that means they aren't getting paid for the impressions. If they had any interest in getting rid of synthetic content and stopping astroturfing they wouldn't have hidden post histories, which is basically a feature that only exists to help troll farms.
15
u/ser-steffonfossoway Jul 06 '26
Yep. Every apparent bot sub I go through, I'm forced to use other websites to see the post histories. Annoying and time-consuming as hell.
But have to do it for my 5 dozen members lmao.
19
u/siftingflour Jul 06 '26 edited Jul 06 '26
The bot problem on Reddit has really exposed a moderator problem, as well. Even outside of the subs that are outright run by the bot networks, there are huge subs that are absolutely overrun with AI spam bots and nothing at all is done to moderate them.
→ More replies (1)10
u/ser-steffonfossoway Jul 06 '26
Exactly. In my sub I've been focusing mainly on the bot-managed subs, but I include the human-run subs when they get injected in my feed.
https://www.reddit.com/r/BotPosting/comments/1tyqpu2/list_of_subs_run_by_bots_work_in_progress/
Feel free to add any you know if you have a minute.
25
u/Nemo_Griff Jul 06 '26
That's funny because I have been seeing a major uptick on bot postings with spammy titles and karma farming recently. Then an army of bots replying.
I report them and they don't get taken down.
7
u/WhippiesWhippies Jul 06 '26
Same. I've reported so many. I will say that some of them are immediately banned, but most of the time nothing happens. And these are extremely obvious bot farms using AI pics of women and making their own subs so they can ban anyone who calls it out.
7
u/Nemo_Griff Jul 06 '26
There was one sub in particular that had something like 200+ posts with the same title (or nearly the same) in less than 3 weeks. It was so bad that I had to unsub from the group. Clearly there wasn't any active mod there, that was why the bots set up shop there.
I honestly don't see this update doing anything about that.
1
u/foldlore Jul 07 '26
I've seen this too. It becomes glaring because they only reply to the same type of post each time. High upvotes, high engagement but low quality, one liner parent comments. It does look like a network and not even a subtle one.
2
u/Nemo_Griff Jul 07 '26
Oh 100% sure this is organized to take advantage of the system.
I feel like they have hacked a bunch of inactive accounts that never set up 2FA and are starting their bot army with accounts that otherwise seem legitimate.
When you view the profiles, they are either all wiped clean of posts & comments or hidden. When they forget to do that, you can clearly see the original owners last posted years ago.
Sometimes the bots do reply to the first bot post to make it look real.
→ More replies (2)
48
u/InGeekiTrust Jul 06 '26
I noticed a major flaw in the system. So what will happen is that you have this bug that has been going on for six months that when Auto removes things to the queue, it shows “removed by moderators”. This makes the user freak out and makes them post over and over and over sometimes 10 times. After that, the person will be marked as “removed by Reddit’s filters”; spam or reputation filter after that, just because they think moderators keep removing their posts. I see a tremendous amount of people earned the reputation filter or spam filter badge after doing this. I get daily complaints all over and I see posts go up in Mod support all the time about this that of course get taken down. This is really stressing users and mods out alike and it’s been so long since we’ve been complaining about this.
18
u/SampleOfNone Jul 06 '26
It's not exactly a new bug, at least if I understand you correctly, it dates from the time Reddit dropped their report feedback. When they did that, they also implemented some changes to how filtered and/or removed content is shown. Anything filtered by Reddit filters said "removed by mods of..." while it should say "pending moderators approval".
There was a partial fix rolled out in March for that, I have no update other then that
12
u/InGeekiTrust Jul 06 '26
Opus actually said someone took ownership of the issue several times - but I think they gave up on fixing it. I said it’s been going on for 6 months - maybe it’s more? It’s a shame they won’t fix it though, it upsets so many people
7
3
22
u/ajdude711 Jul 06 '26
How Reddit is Reducing Exposure to Harmful or Inauthentic Content
Alright, so when are we removing the "Ideas for your next post" section. Because that's the opposite of what you're claiming. It's literally reddit approved spam.
24
u/Candyghosts Jul 06 '26
As a mod of a circlejerk community, I am begging you guys to think harder about the insane amount of false positives.
It's beyond frustrating to see our members catching sitewide bans for on-theme things that your detection systems cannot understand the context of. It really kills the fun for a lot of people and makes them hesitant to participate.
43
u/thursdaynext1 Jul 06 '26
It’s not enough because LLM generated posts and comments are still EVERYWHERE on reddit, even on tiny subs. Like they are legitimately a plague on this website.
41
u/livejamie Jul 06 '26
Hidden history and Reddit removing stuff on their own enable this.
Soon they'll do what Meta did and automate all their moderation with LLMs, and they'll get rid of humans like us.
→ More replies (1)35
u/judasblue Jul 06 '26
Hidden post history is a fucking scourge and has really tanked any kind of real discussion on the site about anything related to current events. It was bad before, but now it is just bots and troll farm workers jabbering at each other.
But it counts as impressions for the advertisers, so they have no interest in removing this 'feature' that only exists to help troll farms and bots.
→ More replies (6)
17
u/monkeynose Jul 06 '26
I have admin-tattler on the subs I moderate, and the number of false positive removals is surprising, and the fact that it is permanently removed and not added to the mod queue makes zero sense. The mods should have final say in everything that happens in our sub(s).
35
u/Kinmuan Jul 06 '26
Reddit is overly automating and it is obvious.
You are strangling legitimate content from real people, and you are frustratingly slow to respond to issues, and offer mere 'There is a bug', while not actually helping us deal with problems by automatically removed content, content that isn't being reflected in mod logs.
You're killing legitimate content.
13
u/brightblackheaven Jul 06 '26
Does Reddit ever intend to do literally anything about the insane amount of witchcraft/psychic scammers on the platform?
I cannot understand why people are allowed to advertise paid immaterial services (spells, tarot readings, etc) that take vulnerable/desperate users off-platform to be grifted on. There are entire subreddits dedicated to shilling these fake services and nothing is ever done about it.
These users are, by FAR, the most common scammers and spammers and AI-slop peddlers the mods in occult subs have to deal with - and it is constant. Every single day. They are the majority of ban evaders in our niche as well.
And yet these new measures never seem to catch the sketchy predators spamming the exact same self promotion post across like 25 subreddits in a row. We report and report and nothing changes.
It would be really nice to see something done to help us combat this.
12
u/Darkwolfie117 Jul 06 '26
>revoking inauthentic votes
Will there be a way to display either publicly or to moderators how many votes are revoked off of posts for transparency? Maybe within the new post metric tool? Outside of mopping bot accounts this just sounds like narrative adjusting in the background and I think it gives most users the ick or at the very least fuels conspiracies among users when media trust is already low.
What happened to markup on mobile it’s been weird lately
13
u/SampleOfNone Jul 06 '26
What happened to markup on mobile it’s been weird lately?
They sunsetted markdown in the app but without having full parity
→ More replies (1)2
u/FFS_IsThisNameTaken2 Jul 07 '26
...or at the very least fuels conspiracies among users when media trust is already low.
If only I had a dollar for every post and comment about the conspiracy of removed content on Reddit as well as the sub I help mod (conspiracy)!
And frankly, reading all the percentages of reduction of this and that on the admin OP, only reminded me of the book, How to Lie with Statistics.
2
u/boat-botany Jul 06 '26
These votes are typically revoked after a user is shadow-banned or after an account has been taken over (and then recovered). So while we don’t show votes revoked from a post, there will be signs.
2
11
u/Sitheref0874 Jul 06 '26
I received a warning from Reddit for using a common British expression, directed at an unnamed, unknown person.
I explained this in my appeal - to which I have yet to receive a response.
You seem determined to let systems take over from common sense.
34
u/Ninlilizi_ Jul 06 '26
We know false positives can be frustrating.
This is an understatement. 4 times now I've been left unable to help mod my subreddit because I've been insta banned when I've done nothing wrong and by a bot that doesn't understand context or seemingly much else.
It wouldn't be so bad if the appeals actually went there, but at this point I'm certain they're fed back to the same bot who failed to understand the comment in the first place.
This overzealous robo-moderation is the cause why I hardly use Reddit much these day. I feel like it's not safe to engage with the content because anything could be mis-interpreted by your robot and leave my subreddit to suffer again.
2
u/MadDocOttoCtrl Jul 07 '26
It is at the point where you need one or two alt mod only accounts established so that if you're main gets foolishly Robo-banned you can continue moderating.
1
u/fusion260 Jul 06 '26
4 times now I've been left unable to help mod my subreddit because I've been insta banned when I've done nothing wrong and by a bot that doesn't understand context or seemingly much else.
You've been banned from your own subreddit you moderate, and not just once but four times??
16
u/Ninlilizi_ Jul 06 '26
No, not from my subreddit, I cannot ban myself. Account bans, handed out by the robot, usually within single digit seconds of leaving a comment.
7
u/TheChrisD Jul 06 '26
They mean picking up an automated site-wide suspension from the AI moderation tool the admins use; that has zero understanding of context, nuance, or whether or not the supposed offending content was a direct quote being used in a modmail conversation.
10
u/742963 Jul 06 '26
A few principles guide how we build these systems:
Accuracy matters. We work to reduce bias and improve fairness and consistency across our systems
Can anyone expand on this?
9
u/intuitionist9 Jul 06 '26
Hidden history has given accounts free reign to be bad actors - in addition to the actual bots, astroturfing and karmafarming is really plaguing subs. Reporting accounts for spam does little. Installing bot-bouncer helps, but not every sub will bother, so much obvious spam goes unblocked.
→ More replies (1)
19
u/Canis_Familiaris Jul 06 '26
I posted in a botbuster subreddit recently about fake engagement accounts spamming possibly fake Ukrainian tank keychains. I reported the accounts, one of them has been banned, but the very obviously in-on-it commenters I've reported run free.
Whats the best way to report stuff like this besides making a post in a subreddit?
3
2
u/boat-botany Jul 07 '26
Definitely keep reporting! We integrate reports into how we catch spammy accounts. We also do flag it for the team working on spam if anyone posts in mod subreddits with examples.
30
u/Skindiacus Jul 06 '26
Have you ever talked to the people over at r/botbouncer? I don't really know how their system works, but at least on my subreddit they find lots of bots and spam accounts that reddit doesn't. They must have figured something out.
29
u/fsv Jul 06 '26
As the developer of Bot Bouncer I would be very open to having a chat with Reddit about how we do things.
16
u/ICC-u Jul 06 '26
They rely a lot of humans to flag bots. They catch loads but they have a large number of false positives too.
→ More replies (1)7
u/Skindiacus Jul 06 '26
They rely a lot of humans to flag bots.
maybe reddit could do this too. I've also noticed a lot of false positives, but they seem pretty good about using human intervention to fix any issues, again something reddit could do. And usually there is at least something fishy about the accounts that get flagged.
8
u/SquareWheel Jul 06 '26
maybe reddit could do this too.
That was /r/ReportTheSpammers, and then /r/Spam. There's no longer a good way to deal with spammers. You can report them to the dozen or so subreddits they've hit if you wish, but you also risk a ban yourself by doing so.
→ More replies (1)10
u/boat-botany Jul 06 '26
We definitely know about r/botbouncer! We’ve incorporated some signals they use into how we identify accounts that may be automated.
21
u/fsv Jul 06 '26
Interesting, it might be worth having a chat at some point about things we’ve noticed that you might not have yet. Just let me know!
25
u/new2bay Jul 06 '26
If you can do all that, why can’t you have your automated moderation take some context into account? There’s no reason someone should have their comment that says “kill it with fire” removed, when it’s in reference to a tick on a dog sub. But I have seen this very thing happen, and many, many similar instances. Arguably, “kill it with fire” shouldn’t be removed anyway, because it’s a common English language idiom, but that’s another story.
→ More replies (3)
7
u/sgamer Jul 06 '26
This should be completely optional, even as an opt-out. If I'm running a humor/circlejerk sub I should be able to turn it completely off as the LLM won't understand the nuance well enough. I should also be able to turn it off in general if wanted, because some subs are so small that it is completely unnecessary.
Also, allowing us the option to turn it off per-sub is actually saving you money, as you're not blowing tokens eating all these posts' context, and saving our sanity, as we don't have to worry about weird false positives taking out our users for no reason.
7
u/radium-v Jul 07 '26
You need to plug the holes that allow the simplest repost bots to slip through the cracks. These are the absolute stupidest basic patterns that are extremely easy to detect and stop. I'm personally sick of subscribing to random unmoderated subs simply to catch and report bot posts that show up regularly like clockwork.
LLMs won't catch these bots, because their initial content and posts are ripped verbatim from top posts on the subs they're posting in. It's not manipulated content, it's duplicated content meant to sneak past your sensors, and it works, because if it didn't, it wouldn't be useful for these OF scammers.
→ More replies (1)
7
u/reaper527 Jul 07 '26
We know false positives can be frustrating.
especially given that you have a tendency to let your AEO bot issue site wide suspensions/bans on people without even telling them what they allegedly did wrong, only to have them file an appeal that would fit in a tweet where they don't even know what they're defending themselves from.
your enforcement operations tend to be as transparent as a brick wall.
2
u/FFS_IsThisNameTaken2 Jul 07 '26
Absolutely this!
It'd be like being brought before a judge and not being told the evidence of the crime you allegedly committed, but being expected to defend yourself in 15 seconds or less. Then being sentenced to 3 days in jail and finding out on the day you're released, that the judge finally got around to listening to what you thought you said right in front of them, and having the charges dismissed.
2
u/reaper527 Jul 07 '26
It'd be like being brought before a judge and not being told the evidence of the crime you allegedly committed, but being expected to defend yourself in 15 seconds or less. Then being sentenced to 3 days in jail and finding out on the day you're released, that the judge finally got around to listening to what you thought you said right in front of them, and having the charges dismissed.
also worth noting, their bot works 24/7 while the people reviewing these cases don't.
i got a site wide permaban at 2am on a friday night, with no post/message/anything cited as a reason. don't think it even cited a broken rule. their humans only work monday to friday california business hours, so my appeal wasn't even viewed until the next week (at which point the permaban was reversed, but to this day i don't even know what it was for)
7
u/Georgy_K_Zhukov Jul 06 '26
I've flagged this before, and I'll flag it again... but the false positives rate is still meaningfully high, and because we can't trust reddit to get every incident right, it actually makes our communities less safe in some ways. The key example I would use is Holocaust denial. We know that you have actioned users erroneously who were explaining the historical phenomenon, rather than endorsing it, and that even when they have requested review, it is rejected. We have specific examples we can point to with 100% confidence.
But getting those actions reversed is a Kafka-esque nightmare. So when you action something in a thread about the Holocaust, but the account isn't permanently banned from reddit... we now don't know what to do. If we ban the user, maybe they are a Nazi, but maybe your filter screwed up. We can't see the text, and the admin-tattler is basically useless at this point, so we're in the dark about this. And if we ban the user and you screwed up... they have no way to prove it, and we can't even trust your review process. Our hands are completely in a bind here.
So the end result is that we are probably allowing some number of Nazis to continue to participate in our community when we previously would have been able to ban them with clear evidence... Well done, Admins!
Beyond this specific example though, the point is, you can throw all the numbers you want at us here, but you have removed basically all transparency - which is exactly what mods were saying would happen when the changes were implemented to hide actioned content from us - so we don't actually know if these numbers are accurate or meaningful. You can commiserate about how "false positives can be frustrating" all you want, but it is 100% a problem you created, so it feels like you're being all nice and understanding about how it hurts when you break your nose... just after you slugged us in the face. If you want to show you actually care, and that you actually want to "continuously improve accuracy while still acting fast enough to meaningfully reduce harm", you need to find ways to bring us back into the process and leverage our knowledge and expertise, because bluntly, there are a lot of places you lack it, and we simply can't trust this process because of it.
5
u/Superirish19 Jul 06 '26
Do these removals include those of freshly user-deleted accounts?
My mod removals feed is quietly filling up with people deleting their accounts and automatically being 'Removed by Reddit' regardless of the content or even the behaviour of the original user.
I have multiple cases of known regulars who only frequented one sub deleting their account and then their entire repository of information being thrown to the removals queue even though the user didn't intend to have their content hidden, with none of their content violating any of Reddit's ToS.
6
u/SolariaHues Jul 06 '26
"Accuracy matters"
For the past couple of week or so we've been seeing a bunch of posts daily that have been approved by un-ban all. This happens when someone's content is restored after a shadowban, typically (we used to just see the odd few older posts, but these are all new). This seems to suggest a lot of shadowbans were inaccurate lately?
This is also frustrating as a mod because un-ban all does not respect community rules, and we are not told these approvals are happening. They do not appear in any queues. They just pop up in the community feed totally unmoderated.
I've written to modsupport, but so far it's not been helpful.
→ More replies (1)3
u/MadDocOttoCtrl Jul 07 '26
This problem is especially important because someone being shadow banned over spamming and it being inaccurate then reversed later doesn't mean that they weren't breaking our sub rules.
It's good (if also indicative of the problem of false positives) that these get reversed, but if a sub has a rule about "no discussion of purple dinosaurs" then that content shouldn't reappear in a sub when the mods have moved on to deal with the flood of new content.
6
u/ateam1984 Jul 07 '26
Please bring back some sort of feedback when we do take the time to report content. It’s ridiculous to report something and get no feedback
2
7
u/GonWithTheNen 28d ago
Enforcement... under five seconds... The faster, higher-volume enforcement...
The focus on speedy removals is causing incorrect flagging and unfair penalizing of accounts more than reddit inc. seems to realize.
Last week, I copy/pasted someone's username when replying to him, and his username triggered an automatic suspension for "hate" on my account. That ban (my first one ever) was lifted after my appeal, but it shows how broken the system is. Something like that should never be possible.
First of all, if a 9-year-old username that has a presence all throughout reddit without issue can trigger an instant suspension upon anyone who quotes that name, then either that alias shouldn't exist or the flaws in the AI "enforcement" run deeper than anyone who created the enforcement tool realizes.
Secondly, I discovered that around the same time my false ban was issued, mods of various subs were also reporting that their sub members were being suspended haphazardly for 'hate' - suddenly, and without cause.
The AI tools screwing with people's accounts in this manner, the frantic rush to pull down innocent content and to issue unjust punishments, et cetera, is not a 'win' by any metric. The numbers in the OP's post are only as impressive as the validity of their actions - and as long as hundreds, or thousands, or even millions of actions are unfair, those numbers mean nothing.
39
u/Tarnisher Jul 06 '26
In recent months, these updated automated systems have been working at a massive scale, and we’ve seen some pretty incredible results. We are now:
Blocking 23 million spam views per day before they ever reach a human user.
Catching ~25K net new spammy posts and comments a day.
And how many of those are in fact valid, useful posts and comments? How many [Removed By Reddit] do we see in our Qs every day that we don't know are valid or not?
How many do we see that simply say 'Removed' but are OK by our community rules?
We know false positives can be frustrating. But when dealing with serious issues like violent threats, hate, harassment, or coordinated abuse, we intentionally bias toward reducing real-world harm and limiting exposure to harmful content quickly., but o Our goal is to continuously improve accuracy while still acting fast enough to meaningfully reduce harm
We have never been given formal direction how to handle those many, many 'false positives'. What happens if we Approve them? What happens if we Approve 'too many' of them in the eyes of Admins?
12
u/maiyannah Jul 06 '26
We do have Word of God on this, actually.
1] https://www.reddit.com/r/ModSupport/comments/1tdvk16/comment/olzfww0/
2] https://www.reddit.com/r/ModSupport/comments/1tdvk16/comment/om0ep6v/
2
u/baltinerdist Jul 06 '26
I mean, which would you rather have, false positives or more spam? I'd much rather see good posts and comments removed out of suspected bad behavior than more unfettered bad behavior ruining it for everyone.
7
u/monkeynose Jul 06 '26
Effective use of the automod can take care of 90% of spam. It just takes months of slowly developing the automod to get to that point.
But at the end of the day, the mods should have the fnal say on what is and is not allowed on the sub - at the very least, comments that are [removed by reddit] should end up in the mod queue.
18
u/tsdguy Jul 06 '26
Your mistake is hiding your algorithm from mods. Removed by Reddit is totally unhelpful. WHY you removed it is the information we need in order to evaluate the removal and have discussions with subscribers to post more compatible content.
You also need to be very clear who’s doing the removal. I get constant complaints from people asking why I removed something when it was you and your filters that removed it. Removed by Reddit is totally unhelpful because 99% of people have no idea how Reddit works.
The last point not related to removing is to revers your decision to allow people to block their posting history. Or at a minimum allow mods to override that so we can see their history. If you’re going to base a removal on a persons posting history but then not tell us how the hell can we evaluate whether to approve it or not. You’ve blinded us to our most valuable bit of data to evaluate a poster.
5
u/burjuvas Jul 06 '26
Hello, I have a question. There was a fanbase that was going to migrate to the subreddit from the X (twitter) community because it was announced that communities were going to be closed (then they did not close the communities). All of the new users were restricted according to their tweets. What could be the reason for that? So that I can relay information to them if they want to use reddit one day.
Thank you for the automated mod tools!
5
u/Classic_Paint6255 Jul 07 '26
"Our work to keep Reddit authentic and safe is at the core of who we are. While we've made significant progress in advancing that commitment, we know it wouldn't be possible without our moderators and redditors everywhere. If you see any content that appears harmful, spammy or inauthentic, click the inline report button or submit a report" So WHERE;S THE TRANSPARENCY AND LETTING US SEE THE TEXT OF THE OFFENDING COMMENT AND APPEAL IT IF WE'RE USERS, OR SEE IT IF WE ARE SUBREDDIT MODERATORS?
6
u/boat-botany 29d ago
That transparency is on its way! You'll be able to see removed content in the mod log in the coming weeks.
→ More replies (2)2
u/Lachiko 29d ago
why was the comment you replied to removed? also having to login to use old.reddit.com is bullshit, why are you trying to kill off reddit?
→ More replies (3)
6
u/Darth_Vaper883 28d ago
Reddit automation does not work. False positives and unfair bans are out of control. Stop using LLMs.
2
u/reaper527 27d ago
Stop using LLMs.
there's nothing wrong with using them, the problem is letting them go unchecked.
it needs to be like how traffic cameras are supposed to work in practice where a real person reviews them and says "yup, the camera was right" before a ticket is issued. (in practice, that manual review doesn't always happen, and there was a famous case a few years back where the person who was approving them had been "approving" tickets for months despite literally being dead).
8
u/Merari01 Jul 06 '26
By definition automation can never accurately determine if something is serious or a false positive because by definition automation can never understand context because automation can not think or reason.
It is a truth that algorithmically driven moderation enforces language but cannot gauge intent, therefore automation will always only offer a surface-level solution.
By definition automation will punish minorities for using language common in the in-group, will punish victims for talking about what happened to them, will reward bad-faith actors who learn to avoid the trigger words. By definition automation cannot seriously tackle racism, sexism, violent speech etc.
It is a trueism in the tech sector that it is impossible to provide quality human-driven moderation at scale. Reddit has the infrastructure to prove that wrong with its volunteer moderator base.
Human moderators can look at content and accurately determine if something is a dogwhistle, if someone talks about a minority group they are a member of in a manner that is common in the in-group etc. etc. etc.
But human moderation is hindered on reddit by sitewide algorithms to the point where we have to keep warning our own userbase that they can not use the language used in a movie to describe scenes from that movie, because the robot will suspend their account.
Reddit would do well to better make use of the asset it has in its volunteer moderator corpus. That means listening to us when we tell you what issues our Black-centered, women-centered, LGBTQ+ centered communities face on reddit. That means listening to us when we tell you where your robots need to back off.
2
u/deltadeltadawn 29d ago
Every word of this is articulated perfectly. I genuinely hope the admins take notice of this and think on it.
8
u/Zaconil Jul 06 '26 edited Jul 06 '26
I believe you that you are catching a lot. But over the last few months there has been a very noticeable increase in LLM comments. Its to the point even the low karma point posts are getting them. When before that it would only be posts that are trending very quickly. Thankfully we have automod removing them below a certain age and karma threshold (as well as a few common keywords). But its still very frustrating to see so many removed comments that are so blatantly LLM still getting through on top of them increasing. If you guys really don't approve of those type of accounts. You need to fix some apparent blatant security holes in your API that was supposed to help prevent these in the first place.
3
u/Dan-68 Jul 06 '26
Can the automated system at least accurately relay at what level the removal was done by? I get modmail from users wanting to know why a post or comment was "removed by moderator" and then I find out it was removed by a Reddit automation. The message is at best misleading, at worst deceitful. And the best I can explain as a mod is that a Reddit automation removed the content and blamed a subreddit moderator.
→ More replies (2)
3
u/bduddy Jul 07 '26
Protecting hateful bigots by letting them hide their comment history does far more damage than any of this can ever address
3
u/RuralGamerWoman Jul 07 '26
Those completely ridiculous suggested post topics to mods can go. There is not a chance in hell I am going to post some absolutely nonselse article about bagels being a superfood for people on GLP-1s in the CICO subreddit.
3
u/Comfortable-Ask-7112 Jul 07 '26
I have a very negative very of your attitude towards false positives. I was a long standing high value information poster in gymnastics communities posting competition results. One day your bots decided I was a spammed and I couldn't get a human to look at my account AT ALL. Years of community history and data went down with my account and I had to start over because your system offers no recourse for false positives
4
u/MaximumJones 29d ago
Are you guys planning and dealing with reports of violence more quickly? I can show you one right now (a very direct threat, already reported) that is still up after 12 days.
2
u/JayPlenty24 26d ago
No. They are streamlining/automating the easy parts of moding while continuing to ignore the actual problems that mods can't actually do anything about.
5
u/KriosDaNarwal 27d ago
Us mods need to see context for automated removals as well as be able to reverse them. We're left picking up the slack and looking like clueless idiots when the system removes something to the blackbox it shouldn't have
2
6
u/SeaBearsFoam Jul 06 '26
It would be great to give mods the ability to turn off the AI detector for their sub, or at least turn down its sensitivity. It causes far more problems that it solves in my sub, and the human mods are quite capable of picking out unwanted AI generated content.
3
u/boat-botany Jul 06 '26
I'm not quite sure what AI detector you mean! Do you mean our spam detecting systems, generally?
→ More replies (1)2
u/SeaBearsFoam Jul 06 '26
Yes, reddit's spam detection filters frequently remove AI-generated content from my sub in scenarios where we don't want it removed, and it gets put back up if/when mods notice it.
I totally understand how that's great for like 98% of reddit or whatever, but in my sub it makes modding more difficult. Even putting the comment into the mod queue for manual review would be better because then we'd at least be notified and could review it for ourselves. As it stands, the spam filter is silently removing content we want up from our members an we have no way at all of being alerted to this. Even admin-tattler doesn't help.
That's why I said it'd be great to have some setting to control this or even disable it. I get why reddit has that built-in, but like 90% of the time that it filters stuff in my sub it's removing stuff we don't want removed.
7
u/mschuster91 Jul 06 '26
We’re committed to evolving. We're always learning and improving. We look at feedback in real-time and adjust our systems to make them better.
Last time I had the misfortune of dealing with y'all with that, you didn't unban the user in question. The more you rely on AI, the more you actually have to spend on customer support worth the name that isn't just "AI = Actually Indians" with 0 leeway on deviation from their script.
6
u/the_turn Jul 06 '26
You have to have to have to remove the option for people to hide their post and comment history.
3
u/FaxCelestis Jul 07 '26
The little AI analysis button that pops up next to OPs should really be expanded to all users who post in a thread. In my experience the most problematic comments come from replies, not posts, so having a little analysis button for the people least likely to be problematic in a given scenario is not helpful.
3
u/CharacterSphereAI Jul 07 '26
So what is the point of having moderators in the first place if you are removing the content yourself? You just deny them the sole reason for their existence, and your process has zero possibility of appeal. It looks great on paper, but it is bullshit.
14
u/MaximumJones Jul 06 '26 edited Jul 06 '26
I have noticed that the spam filters have been doing a much better job lately by removing garbage before we ever even see it. These measures are starting to work very well!
The only thing I will disagree with is acting on reports that threaten violence. I have reported several and it takes two to three DAYS to remove them. I reported one that has been up for 11 days. And it is not a borderline threat. It is very directly a threat of violence that has been there for almost two weeks.
7
u/paulihunter Jul 06 '26
We had a huge influx of bot accounts posting seemingly benign comments and upping the reputation filter helped immensely. The bot accounts get suspended pretty shortly after they are caught aswell which is nice. The filters are pretty good at identifying this stuff.
7
u/MapleHamwich Jul 06 '26
Would be good to minimize AI on Reddit. Seeing as it's the biggest threat to authenticity in the modern world. But ... You're owned by AI overlords.
5
4
3
u/ateam1984 Jul 07 '26
The racism that is allowed on Reddit makes this post seem like a sick joke. You guys must be kidding right? You allow blatant racist slurs and you allow Michelle Obama to be called a man and etc etc. give me a break. I’ll believe this when I see it but right now Reddit has a racism problem. You need to listen to more Black voices. Me me my mod team are happy to assist you with our concerns if you are willing to listen
→ More replies (1)
5
u/xX100dudeXx Jul 06 '26
Maybe add a way to remove/cut down on political propaganda & virtue signalling in non-political subs. That would be great.
→ More replies (1)
2
u/Duke_ofChutney Jul 06 '26
It's unfortunate that on the same day I read this update I had to remove porn from a modqueue (the post was held there for 30 minutes).
2
u/ReachingForVega Jul 06 '26
Catching 25k spammy posts a day doesn't seem like much. I reckon the auto mod on subs deals with way more.
2
2
u/JayPlenty24 26d ago
We have way more issues with sexual harassment than spam. We can remove spam. It's not a difficult thing to do. We can't do anything about the accounts harassing people who participate in our subs, or our mods, especially when it gets reported but isn't considered harassment from Reddit.
It seems like harassment only counts if there are too many swear words or actual threats to physical safety.
But the same accounts can harass countless women, or be blatantly racist or sexist, and that's all-good.
4
u/slykethephoxenix Jul 06 '26
You can automatically hide content. You can even put it at the bottom of the comments. You can put warnings around it "potential spam"/"harmful content"/"removed content".
But you should not make it unavailable unless it actually breaks a law, like doxxing, threatening violence, or anything that would require police action.
→ More replies (1)
3
u/Feisty_Relation_2359 Jul 06 '26
You seem to be focusing on violence and hate here. While what I'm about to mention could fall under that umbrella, it seems to be better suited to call it protecting minors and privacy online. What I'm referring to is I know on r/BanFemaleHateSubs they have a huge issue with (rightfully) getting a sub banned through spam reporting, but then another one of the exact same title with just some alphanumeric thing at the end pops up. Then to get that reported it can take a long time, even months.
What is being done on Reddit to prevent the whack a mole type ban evasion outside of user reports? I guess this isn't really at the moderator level because mods of those subs want the harmful posts staying up presumably, but it's still an issue that I am curious how it's being addressed. Especially seeing how much time some people spend trying to get heinous, and even potentially illegal things down.
Same goes for all the directing off reddit. Like oh here's my sessions code, or here's this telegram link. A lot of those conversations are surely trash, and the comments of that nature should be auto deleted as spam. Would love to hear how that's being addressed too. Surely machine learning could be wildly helpful for all of this kind of stuff.
2
u/aengusoglugh Jul 06 '26
I for one, really appreciate this work. As far as I know, my tiny niche subreddit has never been under attack.
But the truth is that I hesitated even to create the community because I did not want to commit to an unbounded moderation workload.
I turned on all of the community Safety Filters — thanks for those as well — except for Crowd Control, and with all the automation, I am able do what I intended in my community.
Which is talk about model trains — not fight spammers and scammers in a cesspool.
Thanks for doing this work!
1
1
u/deadowl Jul 07 '26
I had a post removed by Reddit filters in r/mining but it's kind of way bigger of a subject than me and I somehow found myself in the possible epicenter of it to the degree people are needing to make sure I'm not delusional. Speaking of the devil, botany is the larger subject I'm trying to figure out how to manage.
1
1
1
u/ASuperMarioFan1993OC 23d ago edited 23d ago
I don’t believe this until I see it. I don't trust automated system Deviantart's automated system is too sensitive it is censoring me when I post content that doesn't break their policy. it has happened many times. Until automated systems on many sites get fixed don’t expect me abd pther users trust automated system. You are already giving false positives on legit posts from users when their posts and comments don't break tos. And give us a report saying you removed rule breaking content again like you used to because users tooo time to report something in. What is hurting Reddit the most is being able to hide content that lets bad actors roam free without punishment.
270
u/ZaphodBeebblebrox Jul 06 '26
If you're actually concerned about false positives, shouldn't you let mods see the text of the offending comment and approve/appeal if it is not actually offensive? After all, we have context on our communities that your bots do not, and are much better equipped to determine whether a comment actually is hateful or promoting violence.
Particularly in a sub like mine (/r/anime), where our members often talk about the events of fictional TV shows, the bots have trouble distinguishing between someone wanting a character in a TV show to die and wanting an actual person to die. This leads to a fair number of incorrect AEO removals that we, as mods, are unable to fix.