Demos, Not Memos: How Colorado Is Rethinking AI Evaluation 

Published by the Partnership for Public Service AI Center for Government with support from Civic Design Collaborative

Lead Contributor: Amanda Starling Gould
Contributors: Nadine Foik and Sean Baker

The Colorado Department of Revenue set a clear goal: reduce customer wait times to under 20 minutes at the state’s Taxpayer Services call center. To help meet that challenge, the division spoke to the public and to state employees to better understand their challenges. The agency enlisted the Colorado Digital Service, a team within the Governor’s Office of Information Technology that partners with state agencies to develop and deploy technology solutions, to help design and implement the solution. 

The Colorado Digital Service is not a typical government IT shop. Embedded within state government and staffed by researchers, designers, product managers, engineers and procurement specialists, CDS approaches technology problems with a multidisciplinary lens.

Alexander Schneider and Rose Barcklow at the Colorado Digital Service (CDS) shared what that work has actually looked like, including the practices that are succeeding, the decisions that shaped them and the questions that remain unresolved.

This is part of a series of interviews we are conducting with leaders across government to better understand how agencies are measuring the public impact of AI implementations.

Fit to Serve

When met with a challenge like this, CDS is charged with ensuring the proposed AI concept solves the agency’s problem and works well for the Coloradans it serves. CDS implements three overlapping practices to ensure their tools are fit to serve: research before deployment, continuous testing during rollout and an insistence that the agency, not CDS, has the tools it needs to measure effectiveness and outcomes.  

CDS is part of the state’s larger IT organization, the Colorado Governor’s Office of Information Technology, which recently hit the reset button, shedding its traditional project-based operating model for a new product-based model at scale. 

User research as the foundation

Before any AI system is built or selected, CDS insists on user research. This is not optional. “You don’t get approved for AI unless you talk to people to understand what the true pain points are and then determine if AI is actually the right solution,” Barcklow said.  

That research serves multiple purposes: it establishes a baseline; it surfaces the pain points that AI might address and those it cannot; and it creates data that can later be used to evaluate whether the system is working. 

In the tax project, this meant conducting call center interviews, building customer journey maps and pulling qualitative transcripts from the IVR system (the automated phone system taxpayers encounter when they call). Barcklow has been using those transcripts to inform decisions about what self-serve AI options to build. “I’m looking through transcripts and putting those qualitative transcripts in front of the tax team as we make decisions about the next iteration,” she explained.

Meet the AI Innovators

Alexander Schneider,  a digital service expert with CDS, spent years in software and consulting working with federal agencies on machine learning and language processing projects. Measurement has been a through-line. “Understanding how those work out over time has been a core part of my work for the last several years,” he said. He joined CDS in March 2025.

Rose Barcklow, a customer experience and design lead with CDS, has been in Colorado public service for 27 years, spending the first part of her career designing health services for adolescents at Denver Public Schools and Denver Health. About seven years ago, she moved into technology consulting and built a public-sector practice before joining CDS in September 2025.

Continuous monitoring, not one-time checkpoints 

The second practice is ongoing. Every Friday, CDS delivers a working prototype to its stakeholders at the tax agency. This is not a slide deck or a memo but a real, functional digital product.  

Those weekly demos are paired with what Barcklow calls “golden datasets”: structured sets of test questions that the tax team uses to evaluate the AI’s accuracy, usability and completeness.  

Beyond the golden datasets, CDS tracks what Barcklow calls “small M” metrics, feature-level indicators that trace the impact of each component. Call summaries used to take five minutes. With the new system, it is expected to take about one minute. The tool that helps call center agents find policy answers quickly is expected to cut lookup time from 10 minutes to two. “We’re able to break down the success of each of these products because we’re measuring specific impacts with each,” Barcklow said. 

On the technical side, Schneider described a layered approach to evaluating impact. They perform technical audits to ensure the tools they build or buy are working as intended and expected. “You can encode facts (like rules, policy and laws) into machine readable formats and test them,” he said. This helps CDS test the accuracy of a chatbot’s outputs. For more qualitative evaluation, including tone, plain language and conversational quality, CDS uses a method where a large language model from one AI “family” evaluates the output of a model from a different family. “If you’re in the same family and you know you’re judging the work of a family member, you’re going to be kind of biased. That’s the same for this technology,” Schneider said. 

For the highest-risk scenarios, the 20% that involves the most sensitive policy, the most vulnerable populations or legally required accuracy, human subject matter experts do the testing.  

Handing the evaluation to the agency 

The third practice may be the most important for long-term impact. CDS is a team of term-limited employees. It will not be around forever. The only way evaluation sticks is if the agency takes ownership of it. They are also the ones who know their employees and customers best. 

“It’s not enough to talk to people at the beginning and then be like, ‘See you later, we’ll hand you a finished thing,'” Barcklow said. “The people involved in designing the system are also involved in measuring the performance of the system. That has been huge.” 

Early in the project, CDS was carrying most of the product thinking. Post deployment, CDS is now helping the tax team own 80% of the product work, with the state’s central IT office owning the remaining 20%. “Taxation is involved in designing, in measuring the performance of these systems, and [they] can own them,” Barcklow said. 

CDS’s process connects evaluation findings to the people who can act on them and to the tax subject matter experts. For the tax agency project, CDS partnered with tax policy experts to create the first suite of ‘contextual evaluations’ – customized assessments tailored to test an AI model’s performance in a specific context – to make sure the tests are well designed to capture in situ impact. The tests help the team build better products. 

As CDS has broken the headline metrics into feature-level outcomes, the tax team has been able to see specific connections between what they built and what changed. That specificity, Barcklow argues, is what makes evaluation of the product’s performance in the real world actionable rather than superficial.

A Choice: What Not to Automate

Threaded through all of this is a deliberate decision that shaped the project from the start: CDS would not automate the parts of the call center job that matter most. 

When Barcklow was conducting employee experience research, she and Schneider were watching what was happening in the broader call center market, where AI tools were taking over more of the agents’ work. They made an explicit choice to stop short. “We’re going to stop at augmentation and just find ways to support the employee experience,“ Schneider said. “The dignified part of their work, which is connecting with taxpayers, is not ours to automate.” 

This showed up directly in the metrics they chose to optimize. CDS focused on average speed to answer and the time it takes for a taxpayer to reach a person rather than average handle time, which is the time spent on each call. “We want someone to spend as long as they need to help the person who is on the other side of the phone,” Schneider said. “That interaction is protected and sacred.” 

The same logic drove the design of the internal chatbot for tax supervisors. The tool handles rote policy lookups and questions with clear factual answers. The judgment calls, the interpretations, the conversations that require context and the overall experience impact the people who call for help. “We made an explicit distinction between the policy cruft, which we can automate and make simpler, and the tacit knowledge and judgment that should remain the realm of expert humans,” Schneider noted. 

What Other Governments Can Learn 

Neither Schneider nor Barcklow positions themselves as having all the answers. But the practices they have built over seven months of difficult, trust-intensive work point toward something other governments can use. 

  • Start with users. The initial focus should not be the AI tool or the vendor’s pitch, but the people who will use the system and the people it will affect. “Have they talked to users?” Barcklow said. “You don’t get approved for AI unless you talk to people.” 
  • Expect the iceberg. Foundational work, data governance, tech debt and organizational structure are not distractions from the AI project. It is the AI project. “You’ve got to solve all these underlying issues and then we can make AI work,” Schneider said. 
  • Deliver prototypes, not presentations. “Demos, not memos” is how the team describes its approach. Showing stakeholders a working product, even an imperfect one, every week does more to build shared understanding than any slide deck ever could. 
  • Hand evaluation to the people who will live with it. Support teams like Schneider and Barcklow will not be around forever and they are not tax/call center specialists. Evaluation practices that live only in the digital service team are fragile. The goal is an agency that can measure its own systems. 
  • Use procurement as a lever. “Demanding quality software, embedding standards into our contracts and using evaluations as a way to shape the demand signals back to industry,” Schneider said. “We have a huge responsibility to shape the way technology comes to us, and we can do that through procurement.” 
  • Understand that AI is an investment. The cost of using AI is not just the licensing fees, but the data engineering, the testing, the ongoing maintenance and the staff time for evaluation. “AI doesn’t save money,” Barcklow said. “AI is very expensive and that needs to be really clear.” 

The call center is still running. The agent assist tool is being tested. The self-serve IVR improvements are being iterated. The metrics are moving slowly. And Barklow and Schneider are meeting every Friday with their partners at the tax agency, opening a prototype and asking: Does this [still] work for people? 

That question, asked honestly, week after week, is what evaluation actually looks like in practice. 

Getting Started

Part of the discovery work CDS led involved a question that sounds obvious but rarely gets asked in practice: is AI actually the right solution here? 

“It is definitely not black or white,” Barcklow said. “There was some debate (regarding whether) it is the organizational structure that is causing the problem? Were there some organizational change that could have had the same impact on call center minutes as an AI solution? You also have to look at the willingness of the organization to change.” 

Schneider describes the dynamic that often plays out when a new technology arrives: “As a kind of disruptive moment, it’s opened up the ability for us to talk about policy, data governance, whether we’re protecting ourselves or getting in our own way. Those are hard conversations to have,” he said. 

Barcklow and Schneider use an iceberg image to explain this to agencies. The AI strategy sits at the top, visible and exciting. Below the waterline is everything else: tech debt, weak data governance, unclear roles and responsibilities. “Whether they wanted to approach that first or second, when you have an AI solution, those things have to be addressed,” Barcklow said

Questions That Remain

For all the progress, Schneider and Barcklow are candid about what remains unresolved. 

The biggest one: What is an appropriate accuracy threshold? When an AI system gets a tax question right 85% of the time, is that good enough? What about 90%? “Nobody seems to be able to say, ‘We should expect these models to be 90% accurate or 80% accurate,'” Barcklow said. “It makes me very nervous, especially in the public sector, when you need to be accurate around some things.” The tax leadership keeps asking the question. Neither CDS nor their vendors have a satisfying answer. 

Schneider frames the same problem in terms of what he calls “observability”: Are agencies giving up their ability to observe and/or respond?  It is not enough to evaluate an AI system at the output level. You need to see inside it to track what happens between one process and the next and to have the authority and capacity to diagnose problems and make changes. “The state should retain the authority to actually see what happens in between API calls,” Schneider said. “That’s something we need to ask for.” But seeing it and acting on what you see are two different things. “We still have a fundamental capacity gap,” he said. 

CDS can manage a handful of chatbots with careful, hands-on evaluation, but what happens at scale? Colorado has a backlog of nearly 300 AI projects across state agencies. And that number is only going to grow. “If we have 10,000, 100,000 agents, we’re way over our heads,” Schneider said. “That mismatch is something that I don’t know how we solve, but I would love to pose that question to other people you guys talk to: how do we do that?” 

Finally, how do we track and manage long-term public impact? CDS has good data on how long it takes to look up a policy answer. It does not have good data on whether AI is nudging Coloradans toward tax compliance, reducing administrative burden or changing the experience of interacting with government over time. “We don’t have a lot of great taxation longitudinal studies that look at the impact of this stuff over time,” Schneider said. “We probably need a cohort of all kinds of different actors, including fellow states, to get together to do that.” 

Explore other case stories

Explore more below or see all case stories.

Demos,
Not Memos

How Colorado Is Rethinking AI Evaluation

Read more

A Source of Authenticity

How the Library of Congress Approaches AI 

Read more

Building the Conditions for Responsible AI 

How Maryland is Approaching AI 

Read more