Stern Capital

Keeping Adult Content Away From Minors

In briefAdult or harmful content reaches minors not because of one broken control but because of a chain of weak links. Soft age signals, evasive uploading and labeling, imperfect classifiers, and recommendation systems that amplify without understanding what they push. A bad actor only has to beat one link. The defense is depth, so beating one link is not enough. Strengthen age assurance, layer the classifier with other signals, make amplification safety aware, and design every control assuming a determined evader.

The failure nobody wants to own

Of all the ways a platform can fail its users, one stands above the rest. Adult or harmful content reaching underage users. It is the failure that ends careers, brings regulators, and does real damage to real kids. And it keeps happening on platforms that spend heavily on safety.

It keeps happening because it is not one broken control. It is a chain of weak links, and a determined bad actor only has to get through one of them. This piece is about that chain, and about where platforms can actually strengthen it. It is written for the defensive side. It describes the problem at the level a safety team needs to fix it, and it contains nothing that would help anyone cause the harm.

Think of it as four links. Each one is a place where the system can fail.

Age signals. Most platforms know far less about a user's real age than they pretend to. Self declared ages are trivially false. Weak age signals are the foundation the whole safety system stands on, and when the foundation is soft, everything above it is soft too.

Upload and labeling. Bad actors who want harmful content to spread know exactly how the platform classifies content, and they frame, label and edit to avoid the categories that would restrict it. The content gets in wearing a costume the classifier does not recognize.

Classification. Automated content understanding is good and it is not perfect. It misses novel framings, it struggles with context, and it can be probed until a version gets through. Any classifier that is the only line of defense will eventually be beaten by someone testing against it.

Amplification. This is the link people forget. A recommendation system optimizes for engagement, and it does not inherently understand what it is pushing or to whom. Harmful content that gets past the first three links does not just sit there. It gets recommended, sometimes straight to the users who should never see it, because the system measured engagement and did its job blindly.

The damage happens when all four links fail together. Weak age signals, evasive uploading, an imperfect classifier, and an amplifier that does not know any better.

Where the real leverage is

You cannot make any single link perfect. The leverage is in defense in depth, so that beating one link is not enough to cause harm.

  • Strengthen age signals with more than a birthday field. The details are platform specific, but the principle is that age assurance should draw on signals that are hard to simply lie about, weighed against real privacy costs. This is the foundation, and it is worth the most investment.
  • Do not let the classifier be the only wall. Layer it with behavioral signals, uploader reputation, and friction on the pathways that abuse actually uses. A single classifier is a single point of failure by definition.
  • Make amplification safety aware. The recommendation system is not a neutral pipe. It is part of the safety system whether the team treats it that way or not. Content near a safety boundary should not be eligible for the same blind amplification as everything else, especially toward younger users.
  • Assume evasion and design for it. The people pushing this content are trying hard and adapting fast. A control that was never designed to face a determined evader will not survive one. Build for the adversary you actually have.

Why the attacker view helps the defender here

Understanding how bad actors evade these controls is exactly what lets a defender close the gaps. Not to replicate the evasion, but to anticipate it. When you know which link an attacker attacks first, and how they frame content to beat a classifier, and how they exploit blind amplification, you stop building safety that only works against the naive case and start building safety that holds against someone trying to beat it.

That is the entire reason the insider perspective matters in this work. The teams that only see the defensive side keep getting surprised by the same evasions. The point of understanding the attack is to stop being surprised.

The bottom line

Keeping harmful content away from the users it can hurt most is not a filter you buy. It is a layered system where age signals, content understanding, uploader accountability and amplification safety all reinforce each other, designed by people who assume the other side is trying hard. Get any one link right and it is not enough. Get the chain right and you actually protect the people you are responsible for.

If you run a platform and this is your fight, this is the work we do, from the defensive side, seriously, and with an understanding of how the other side actually operates.

Questions we hear

Why does harmful content reach underage users on platforms that invest in safety?

Because it is a chain of weak links, not one broken control. Age signals are easy to fake, bad actors label content to dodge classifiers, classifiers miss novel framings, and recommendation systems amplify content without understanding it. The harm happens when all four fail together, and a determined attacker only needs to get through one link.

What is the most important link to strengthen?

Age signals are the foundation. Most platforms know far less about a user's real age than their systems assume, and a self declared birthday is trivially false. Age assurance that draws on signals which are hard to simply lie about, balanced against real privacy costs, is the foundation everything else stands on and usually deserves the most investment.

Why is the recommendation system part of the safety problem?

Because it is not a neutral pipe. It optimizes for engagement and does not inherently understand what it pushes or to whom. Harmful content that gets past the first defenses does not just sit there. It gets amplified, sometimes straight to the users who should never see it. Content near a safety boundary should not be eligible for the same blind amplification as everything else.

Does this article help anyone reach minors with harmful content?

No. It is written for defenders and describes the problem only at the level a safety team needs to fix it. It contains no method for evading controls or reaching an audience. Understanding how evasion works is what lets a defender anticipate and close the gaps, which is standard security practice and the entire point of the piece.