Date: Friday, October 2, 2026
Hi. We’re Jeffrey Niblack and Karen Byrnes. When we set out to build Measuring The Intersections of Arts, Health, & Well-being, an evaluation toolkit designed to work across arts forms and across health and wellbeing outcomes, we knew the goal was ambitious.
Our shared purpose was to help to help arts organizations with limited evaluation knowledge and capacity effectively understand how their programs impact health and well-being outcomes in an informed, systematic, and authentic way.
We started our process with a shared understanding that “arts and health” is not one field – it’s dozens of fields wearing a shared name. A choral program for older adults with dementia, a freerange arts center addressing social connection, and a clinically-based music therapy initiative might all fall under “arts and health,” but they rarely share a theory of change, a funding structure, or a workforce. Our job was to build something flexible enough to serve all of them without becoming so generic it served none of them well.
Sixteen months later, the toolkit is done and some of the most useful things we produced were the lessons we learned building it. We want to share three of them here, because we suspect they’ll resonate with anyone doing evaluation work in the arts and health sector.
The same artistic discipline, delivered in two different settings, could plausibly be evaluated against multiple constructs of wellbeing and several self-reported health outcomes, or none of the above, depending entirely on the goals of the organization, the expectations of funders, and the needs articulated by participants themselves. A singing group for people with Parkinson’s might be evaluated for its impact on vocal strength and swallowing function in a clinical setting, and for its impact on isolation and identity in a community setting. In each case, it’s the same core activity with fundamentally different outcomes logic.
From the start we knew we had to avoid a fixed – or prescriptive – health and well-being outcomes menu that’s based on the established literature alone. We were working towards a toolkit structured around outcome mapping guidance rather than a prescriptive outcomes list. We emphasized a process to help organizations articulate their own plausible outcomes that is based on a combination of internal organizational knowledge about the impacts of their programs combined with published research and evaluations. The lesson for us, and maybe for the field, is that portability in an arts and health toolkit can’t mean standardizing outcomes. It has to mean standardizing the process by which organizations figure out what outcomes are honest for their specific context.
We created this toolkit after working with nine different organizations ranging from well-resourced institutions with dedicated evaluation staff to small, volunteer-run arts organizations for whom “evaluation” had previously meant a satisfaction survey handed out after a workshop. All groups needed evaluation support. None of the groups needed the same approach.
We addressed this by building the toolkit in tiers, letting organizations essentially self-select an entry point based on their existing capacity, with clear off-ramps to more or less guided content. This added real time and complexity to development but we hope it improves how the toolkit performs in the field. If your organization builds evaluation resources for a sector as varied as arts and health, we’d argue capacity-tiering isn’t a nice-to-have. It’s close to a prerequisite for a similar resource actually getting used.
This last lesson probably shouldn’t have surprised us, and in a sense it didn’t. But it’s worth naming plainly, because it’s easy to lose sight of amid all the talk of logic models and instrument selection: the arts matter, and they matter in ways that are hard to fully capture.
Across our open-ended survey questions, interviews, and focus group discussions, the same pattern kept surfacing. No matter the art form or the setting, participants described a wide and often unexpected array of ways the arts participation shapes their wellbeing, a sense of purpose, a way back into their own body, a connection to people they’d otherwise never have met, a language for grief or joy they hadn’t had before. Research is starting to catch up, and building a tangible evidence base for the arts’ effectiveness in controlled and measured settings, and that work matters enormously for funding, policy, and legitimacy. But it was the qualitative data that kept reminding us how much humanity comes alive when people participate in the arts and it is something that resists being fully reduced to an effect size, however important that effect size is.
We didn’t set out to prove this. It emerged, again and again, in the words people chose when we simply asked them to tell us what happened. It’s a good reminder for evaluators in this space: our job is to measure rigorously, but not to mistake the measurement for the whole story. We feel genuinely honored to do evaluation work in a field where the thing being measured is, so often, someone’s experience of being more fully human.
None of these three lessons were surprising in the abstract. Anyone with evaluation experience in community-based settings knows that context or capacity matters. What surprised us was how much they shaped the architecture of the toolkit itself, not just its content. A tool meant to travel across arts forms and health outcomes has to be built for variability from the ground up – flexible outcomes-mapping instead of a fixed menu, tiered entry points instead of one skill level, and adaptive data collection instead of a single instrument.
Measuring The Intersections of Arts, Health, & Well-being is currently available on the Performance Hypothesis website and Arts Board is strategizing around a broader community rollout. We’d genuinely welcome hearing from others in the AEA Arts and Museum TIG community who are doing similar work, especially, those of you who are grappling with these same tensions and landing somewhere different than we did. Evaluation in this space is still a young, fast-moving conversation, and the more we compare notes across organizations, the better the next generation of tools will be.
The American Evaluation Association is hosting Arts, Culture, and Museums (ACM) TIG Week. The contributions all week come from ACA TIG members. Do you have questions, concerns, kudos, or content to extend this AEA365 contribution? Please add them in the comments section for this post on the AEA365 webpage so that we may enrich our community of practice. Would you like to submit an AEA365 Tip? Please do so using the form on our website. AEA365 is sponsored by the American Evaluation Association and provides a Tip-a-Day by and for evaluators. The views and opinions expressed on the AEA365 blog are solely those of the original authors and other contributors. These views and opinions do not necessarily represent those of the American Evaluation Association, and/or any/all contributors to this site.