The problem with moderating after someone already read it
A reactive moderation system has a structural flaw that rarely gets said out loud: to work at all, it needs a victim to exist first. Someone has to see the harmful content, feel uncomfortable, offended, or affected, and only then report it, so that a human or automated team can review it and, eventually, take it down.
By that point, the harm already happened. The person already read the insult, already saw the malicious link, was already exposed to the exact content the system was, in theory, designed to prevent. Removing the message afterward is better than never removing it, but it's still a late response to something that already occurred.
In a live conversation space, where several people might be reading the same message almost the instant it's posted, that gap between publishing and review becomes even more of a problem. There's no way to 'report it before' most people have already seen it.
Why we flipped the order
On this site, every message goes through an automated content review before it ever appears in the conversation. That review doesn't happen after other people already read it: it happens in the interval, nearly imperceptible to whoever wrote it, between the moment a message is sent and the moment it shows up on screen.
When that first filter detects content that likely violates our rules (spam, explicit sexual content, hate speech, illegal content, phishing attempts, among other categories), the message is held back or discarded before it's shown. Nobody else gets to read it. Nobody needs to be affected first for the system to act.
When the first filter doesn't have enough certainty to decide on its own, the message goes to an additional review instead of being published by default. We'd rather have a legitimate message take an extra instant to appear than have a problematic one show up while we figure out what to do with it.
No automated moderation is perfect, and we don't pretend it is
It would be dishonest to promise that an automated filter never gets it wrong. Sometimes a system like this can be too strict and block something harmless; other times it can let through something it shouldn't have. No system, anywhere on the internet, has a perfect track record at this task, and anyone who claims otherwise probably isn't being fully transparent.
What we can honestly say is that the architecture is designed to err on the more conservative side: when in doubt, hold it back before publishing, rather than publish it and regret it later. That asymmetry is intentional.
Why this makes us responsible too, not just the filter
An automated filter, however good, isn't an excuse to stop paying attention to the space. We keep reviewing usage patterns, adjusting rules when we find new ways people try to slip past moderation (misspellings, symbols swapped in for letters, combinations meant to sneak past the filter), and updating blocked categories when needed.
This isn't a process that gets solved once and stays solved forever. It's ongoing work, because the ways people try to evade a filter also keep changing over time. Treating it as something that's 'already handled' would, again, be a lack of honesty on our part.
What this means for whoever participates
In practice, participating in this site's conversation means writing in a space where your message passes through a filter before another person sees it, not after. That doesn't remove the possibility of a mistake in either direction, but it does change the point at which we try to correct it: before harm happens, not after someone reports it.
It feels like the only coherent stance for a place that asks, at the same time, for people to speak freely and for that space to remain safe for whoever decides to join the conversation.