Skip to main content

Making AI safety evaluations smarter

Announcing a new collaboration with Harvard's Berkman Klein Center to build a youth safety evaluation ecosystem the field can trust.
Post-it notes related to youth AI safety on a wall

By Robbie Torney

Head of AI & Digital Assessments, Youth AI Safety Institute

 

We want to answer a deceptively simple question: Are AI products safe for the kids and teens using them? And we want answers that are two things at once: fast enough to keep pace with products that change from one week to the next, and rigorous enough that a parent, school, or developer can act on them.

 

We're building toward a youth safety evaluation ecosystem the field can trust, with transparent standards for what to measure and validated methods for measuring it. That’s a bigger job than any one organization can take on alone, which is why we're excited to announce a new collaboration with the Berkman Klein Center for Internet & Society at Harvard University.

 

Harvard faculty, scientists, fellows, and students in Berkman Klein’s new AI and Youth Safety Research Initiative will work alongside Common Sense Media's Youth AI Safety Institute to identify which evaluation and benchmarking methods hold up—and build the ones the field is missing.

 

"The institutions, methods, and benchmarks needed to evaluate AI are still catching up to the technology itself," said Berkman Klein Faculty Director Jonathan Zittrain. "Rigorous evaluation is essential to understanding their effects on young people and informing decisions by parents, educators, policymakers, and developers."

 

So what are we setting out to solve together?

 

Evaluating AI systems for youth safety is still a young field, and its most basic methodological questions are still unresolved—not for any one lab or benchmark, but for everyone doing this work, us included.

 

What needs measuring, and says who? Someone has to set the standard for how an AI system should handle a given topic. But what is that standard validated against, and who checks it? A rubric can look thorough and still be one team's judgment calls, untested against anyone else's, and it needs constant updating as products ship new features and teens find new slang and new ways to push at the limits.

 

How good is the stand-in for a real teen? Some kind of proxy—a synthetic persona scripted over an API, a role-played conversation, even a test account built and operated by researchers—stands in for how a kid or teen would use the product. How do we know it behaves the way a real teen would?

 

Are we even testing the real product? A model evaluated on its own isn't the same as the product a teen actually opens. Most evaluations call the underlying model directly through an API, which is cheap and easy to run at scale. Deployed products are full of friction: accounts, onboarding, and anti-bot rate limits that make automated testing against the real thing hard. That’s why system prompts, safety classifiers, crisis-intervention flows, memory, and age-assurance and parental-control settings—the pieces that make the deployed product the product—often don’t show up in evaluations at all. If you leave them out, an evaluation can end up scoring a version of the product no teen has ever used.

 

Can the judgment be trusted? Every evaluation ends with someone or something rendering a verdict: a second model, a human reviewer, a rubric applied by hand. Whatever it is, how do we know its read on ‘harmful’ matches what actually harms a kid—and that it isn't blind in exactly the places the system it's judging is blind?

 

Once there's a score, then what? A number on a leaderboard doesn't, by itself, tell a parent or a school whether a product is safe enough to use—or what to do next.

 

We run into every one of these ourselves. Today, the Institute's risk assessments answer them one way: rather than scripting a synthetic persona over an API, we build and operate our own test accounts, configured to a specific age, and have our team and child-safety experts hold the conversations and review what comes back. That gets us closer to the product a teen would encounter—though we're clear-eyed that a researcher operating a test account still isn't a teen living their life alongside the product.

 

That rigor has a cost: speed. A full assessment (building the accounts, running the conversations, manually reviewing what comes back, calibrating and validating the ratings) can take weeks. Models change continuously, so by the time an assessment publishes, the product it tested may already look different.

 

These aren't tradeoffs between a flawed approach and a better one. They're a hard problem for the whole field. Fast, automated methods can run at scale but struggle to validate what they're actually measuring or to test the real product. Slower, human-intensive methods like ours get closer to the real product but can't keep pace with how quickly it changes. Neither, on its own, is a foundation the field can build on.

 

We start from a premise: standards and science aren't two separate steps. How you test should follow from what you're trying to measure and why. The reverse is an easy trap: letting whatever methodology is cheap and available decide what gets measured, so anything it can't capture goes untested. It's easy to hear "evaluation" and picture a single test with a pass/fail score. The harder and more important work is agreeing, first, on what needs measuring.

 

The Institute's core work is setting standards: what needs measuring, and why. The Berkman Klein Center can help catalyze the field around putting those standards into practice—a common vocabulary, validated methods, and clear judgment about which tool to use when.

 

Neither half works without the other. That's what's driving our collaboration: standards and methods that are both fast enough to matter and rigorous enough to trust.