Date: Friday, July 31, 2026
Hi everyone! I’m Sudheer Kumar Gandu, and I’m new to program evaluation, though not new to evaluation itself. For the past several years I have been a Lead Artificial Intelligence/Machine Learning Engineer at a global financial services company, where a big part of my work is evaluating AI systems before they go into production: how accurate they are, where they fail, how they hold up for different users and regions, whether they clear the safety and compliance bar.
What I’m newer to the daily practice of evaluation: evaluating programs as social interventions, with the methodological care and equity lens this community brings. Joining the GSNE TIG opened up something I didn’t expect. The most important questions to ask about AI systems turn out not to be technical. They’re the questions evaluators already know how to ask, pointed at a new kind of thing.
Here are three lessons I’ve come to rely on, and I think you’ll find they fit with the instincts you already bring.
Any team can show off an AI system on a good day. What evaluation cares about is the average day, with a real user. Ask the program team to walk you through cases where the system didn’t perform the way it was supposed to. When a team can show you their failure log openly, that’s a great sign. It means they’ve been paying attention. And if the failure log doesn’t exist yet, you’ve just identified the most useful first step the program can take.
AI systems are built on data, and that data reflects whoever was in it. I’ve seen models that perform well on average but noticeably worse in non-English languages, or for users in regions that were underrepresented in the training data. The same logic applies to whatever subgroup matters most for your program language, geography, age, income level, race, gender. The question to ask is the same in every case: has the program team measured performance separately for the groups your program is most trying to serve, and what did they do when the numbers came back uneven? If the answer is “we haven’t measured it that way yet,” that’s often where the most valuable evaluation conversation begins.
This is the question I have come to find most important. AI systems can fail quietly. Unlike a broken form that throws a clear error, an AI might give a confident, fluent answer that just happens to be wrong, and nobody notices.So ask: is there human review for high-stakes outputs? Are errors getting logged anywhere? Is there a path for a user who got bad information to flag it? Programs that have thought through these questions are the ones positioned to learn from their AI system over time, instead of being surprised by it.
Next time you’re evaluating something AI-powered, before any meeting, ask the program team to share their failure log. Not a deck. Not a success summary. The actual record of cases where the system got it wrong, and what was done about it. The conversation you have with them about that document will set the tone for the entire evaluation and will surface insights you couldn’t get any other way.
If you’re working at this intersection, AEA’s Integrating Technology into Evaluation (ITE) TIG is where these conversations are happening inside AEA. It has been an energizing place to bring my own questions and learn from evaluators who have been thinking about technology in practice for far longer than I have.
The thing I want to leave new evaluators with is this. You don’t need to learn how to build AI systems in order to evaluate them well. The instincts you already bring to your work are exactly the instincts AI evaluation needs right now, asking the question behind the demo, looking for who isn’t in the data, centering the people a program is meant to serve.
AEA is hosting GSNE Week with our colleagues in the Graduate Student and New Evaluators AEA Topical Interest Group. The contributions all this week to AEA365 come from our GSNE TIG members. Do you have questions, concerns, kudos, or content to extend this AEA365 contribution? Please add them in the comments section for this post on the AEA365 webpage so that we may enrich our community of practice. Would you like to submit an AEA365 Tip? Please send a note of interest to AEA365@eval.org. AEA365 is sponsored by the American Evaluation Association and provides a Tip-a-Day by and for evaluators. The views and opinions expressed on the AEA365 blog are solely those of the original authors and other contributors. These views and opinions do not necessarily represent those of the American Evaluation Association, and/or any/all contributors to this site.