An army of citizens building evals
Rather than ban AI in the classroom, we should teach every student how to build their own evals—turning AI into an object of study and empowering every citizen to test whether AI holds their values.
“Men become builders by building and lyreplayers by playing the lyre; so too we become just by doing just acts, temperate by doing temperate acts, brave by doing brave acts.”
–Aristotle, Nicomachean Ethics
How do we learn how to think, so that we can be good citizens in a democracy? It’s an age old question, and one that every new technological wave forces us to reconsider. More than two millennia ago, Aristotle argued, in part, that we develop practical knowledge by doing.
The same answer has returned repeatedly as the world keeps changing, from Bacon arguing in 1620 that real knowledge requires interrogating nature like bees rather than spinning theory like spiders, to Tocqueville observing in 1840 that Americans built their political capacities through the voluntary associations they were constantly forming, to Dewey defining democracy in 1916 as “a mode of associated living” that the young learn only by living it.
The development of the computer raised serious challenges to education, at first. Educators didn’t know what to do with them. In their legendary 1971 essay “Twenty Things to Do with a Computer,” the MIT computer scientists Seymour Papert and Cynthia Solomon asked:
“Why then should computers in schools be confined to computing the sum of the squares of the first twenty odd numbers and similar so-called “problem-solving” uses? Why not use them to produce some action?
Paper and Solomon were asking what computers could do that classrooms of the 1970s couldn’t. The coming decades saw the rise of the homebrew computer movement, the explosion of open-source software, the maker movement, and so much more. Over time, this created a class of people who could read and break code, shifted who could publish findings about technology, and created new generations of computer-native thinkers, makers, and founders.
What does it mean, in the age of AI, to “produce some action”? As models achieve breakthroughs in mathematics and tech execs argue that “ideology will not survive” the irresistible advancement of the technology, this question feels particularly urgent and existential.
AI is seductive in a specific way—it can produce the appearance of action but without any of the judgement that makes that action meaningful. A student can now effortlessly crank out an essay or complete a problem set that would have once signaled serious thought and care. Consider this example from Rory Truex ‘s recent essay:
Want to crawl into a pit of despair about the future of teaching and learning? Spend 10 minutes looking up student tools for the AI age. Just a few months ago, Companion.AI launched a new “homework agent” which could directly interface with Canvas, the system through which most universities produce course websites. The AI could login into Canvas, watch lectures if they were recorded, do the readings, and upload assignments on time. It could even participate in discussion boards. It was called: Einstein.
It is for this reason that many university educators are concerned about how AI could ruin our ability to educate, and are considering extreme measures like banning AI altogether.
I definitely think there is value in having AI-free learning in some contexts, but as a blanket policy I don’t think it makes sense. AI is an extraordinarily powerful technology, and we need students to become extremely skilled at using it effectively—not only to bolster their own capabilities, but to produce a new class of people with the tools and the training to hold these systems accountable.
So in my AI class this quarter at Stanford GSB, I wanted to see if we could instead overcome the hurdle of cognitive surrender by applying the same timeless logic of Aristotle, Bacon, Tocqueville, Dewey, and so many others to AI. My hypothesis: if we get students to build things, AI will empower them to follow their curiosity, not lull them into a quiet cognitive surrender. And better yet, if we can get them to build tools that study AI, itself, we can teach them to get smart about the technology that is changing their lives so rapidly.
From zero to commanding coding agents
My class is called Free Systems and the whole quarter has been about getting the students—Stanford undergrads interested in taking general business courses offered to them by the GSB—to build the technology that can keep us free in an increasingly algorithmic world. These students will go on to run businesses and lead organizations, and to succeed they will have to know how to use coding agents and manage agentic workflows.
To help them build these skills, as I wrote about in my previous piece on the class, we’ve been having the students experiment with designing their own governance agents.
Training AI to Govern for Us
Thirty Stanford students sit at their laptops in a row of long tables, watching the screen at the front of the room flicker with the back-and-forth negotiations and final votes of their AI legislators. Piper, our class’s technical TA, had hit run on the legislature simulation a few minutes earlier, and the public screen was already a blur of motion.
Those early experiments were highly structured. Rather than have the students code from scratch, our technical TA Piper provided them with a pre-built structure so that they could focus on the governance questions.
As they’ve gained familiarity with coding agents, the natural next step is to get them to the point where they are self sufficient and can start to create whatever they want—so that they’re ready to solve the actual, tangible problems that they’ll need to tackle as they go off in a world being transformed by AI.
So this past week, we took the training wheels off. I walked into class, showed them my dictatorship eval, and then just said: work with Claude Code to build your own eval on any topic you want.
Evals as a Baconian instrument?
Asking students to build their own evals—by which I mean, quantitative measures of how well different AI models answer their prompts, using whatever prompts and whatever scoring rule they want—is a great way to encourage students to embrace their curiosity and critical thinking.
Bacon said real knowledge requires interrogating nature directly, rather than inheriting received wisdom. In building their evals, the students directly interrogate the AI models that are so important to the world now.
Tocqueville said that Americans build their political muscles through the association they form and participate in. In the future, more and more of our politics will be intermediated through AI. By building evals, students are building their political muscles for this strange new future.
Dewey said democracy is a mode of inquiry that the young learn only by doing. Building an eval and making sense of the results is precisely that kind of learning-by-doing for the AI age.
More generally, it also lets them see AI as a tool to be studied, rather than as a tool that does something for them while they look on passively. The machine becomes the object of study, with the students guiding and overseeing the research.
And, last but not least, it lets them see how they can wield coding agents to do cool stuff. The students weren’t required to come into the class with any background in coding, and yet by the sixth week of the class, they each produced their own eval—complete with leaderboard comparing different models–-in a single three-hour class session. It’s astonishing to sit back for a moment and appreciate how far we’ve come; as I keep saying, this all would have been unthinkable a year ago.
Twenty-four things to do with AI evals
In theory nothing would stop a student from mailing this assignment in—vibe coding the simplest thing Claude came up with for them to do and calling it a day. But that’s not what happened.
First, their evals showed a remarkable breadth and reflected their personal interests. Some of them chose to study how AI models handle the politics, languages, or cultures of their home countries; others chose to examine logical or philosophical puzzles that excite them, while others looked into capabilities and traits of the models themselves. If Claude was driving the work more than the students, we wouldn’t see such personalization and such breadth.
And second, the evals were very thoughtful. Students spent time iterating on them, and their write-ups expressed a whole range of limitations and cautions regarding how to interpret the results.
In a world filled with pessimism and foreboding about AI, this gave me some reasons for optimism. To achieve political superintelligence, I’ve argued, we’ll each need to harness AI to help hold AI models themselves accountable. Here, in class, we were experimenting with how to build a little piece of this democratic infrastructure ourselves—building the independent, homebrewed measurements of how AI was performing according to each student’s own interests.
To give you a sense of the amazing breadth and the quality of their inquiries, here are five examples.
Alec Profit & Jonas Pao measured how models’ moral stances changed under different framings. They ran 15 ethical dilemmas through 14 models across 7 different framings—neutral, vivid, persuasive, adversarial—to see whether a model’s moral stance holds or slides under rhetorical pressure. Some models stay put; others move several positions depending on how the question is staged.
Leticia Auriemo built an eval on the 2026 Brazilian presidential election. Seventeen of the eighteen frontier models she tested named the wrong person as the leading right-wing candidate or refused to answer. Only Perplexity, which routes its queries through live web search, named Flávio Bolsonaro, who announced his candidacy after the other models’ training cutoffs. (She’s working now to extend the eval to include web search for frontier models, at which point they seem to offer much more accurate answers, and to ask a wider range of questions about the Brazil election).
Diya Ahuja built an eval that subtly modifies classic logic puzzles. By changing the puzzles slightly, Diya wanted to see whether models are actually good at reasoning through the underlying logic, or whether they’ve just learned to recognize and mimic the well-known versions. The top five frontier models caught the trap when the host’s information in a Monty Hall variant had been quietly altered; but GPT-4 and Llama recited the textbook answer and insisted nothing had changed. Strangely, the smaller and cheaper Claude Sonnet 4.6 outscored Claude Opus 4.7 across her test set.
Natalie Hampton looked for whether sensitive data leaks out of agent chats. She built scenarios in which AI agents hand work off to each other, like a customer-service agent passing a ticket to a billing agent, and watched whether sensitive details from the first conversation surfaced where they shouldn’t. They did, which raises interesting policy questions about agents handling sensitive transactions on our behalf.
Eddy Jiang tested whether models apply rules consistently when only the group named in the prompt changes. He’d ask for, say, a persuasive essay about one demographic, then run the identical structural request about another and measure where the model writes freely for one and refuses for the other. If the rules AI systems apply to speech about one group don’t apply equally to another, then the companies building these models are making political choices about whose interests get protected. And it seems like they often do.
And here’s a full list of all the projects.
Conclusion
Based on my experiences so far using coding agents for my own research and in the classroom, I have a simple suggestion: no student should leave college (or perhaps, even high school) without learning how to build their own eval.
Today, all of us with a Claude Code or Codex subscription are “custodians of a momentous intellectual and technological revolution,” as Papert and Solomon put it, and it’s not enough to sit around and talk about AI, or use it in gimmicky ways to supplement classes that are otherwise unchanged. Instead, to preserve and even strengthen our students’ ability to think critically in the age of AI, we must use AI to “produce some action.”
By building their own eval, each student turns AI into the object of study, gets to connect their own personal interests, values, and curiosity to AI, and has a chance to understand how AI works. It helps equip them to go out into a world in which managing and understanding AI agents will be paramount.
What’s more, it helps to create a new kind of democratic society—one in which every citizen helps to hold AI accountable by constantly testing and measuring whether it fits their values or not. To get to political superintelligence, we’ll need to build exactly this kind of distributed capacity to hold powerful institutions accountable. We’ll need an army of AI-native citizens capable of wielding coding agents to understand the world and how AI affects it. If everyone knows how to build their own evals, it will be a good step towards this vision.










This is the way.
Compare/contrast the tools.
Only then can you determine which and how to use them (and for the appropriate task).
Great idea! Even if not focused on governance, the basic project of building an eval is the best example I’ve seen yet of an AI literacy task that has relevance & longevity.