Building the Conditions for Responsible AI: Inside Maryland’s AI Innovation Lab 

Published by the Partnership for Public Service AI Center for Government with support from Civic Design Collaborative.

Lead Contributor: Amanda Starling Gould
Contributors: Nadine Foik, Sean Baker, Arianne Miller

Maryland’s AI Innovation Lab was launched in May 2026, a month before this conversation and its first three pilots got their secure virtual testing ‘sandbox’ environments set up the same week. That timing gave us an opportunity to capture in real time how a government team builds the infrastructure to decide, before any tool reaches a resident or a caseworker, what must be true for that tool to be trusted and how the state is starting to define what actually counts as helping the people it serves. Lauren Maffeo, who runs the Lab’s AI policies, pilots and infrastructure, shared highlights from that early work: the tests a tool has to pass before anyone sees it and the emerging framework for measuring whether it’s actually working for the public. 

This is part of a series of interviews we are conducting with leaders across government to better understand how agencies are measuring the public impact of AI implementations.

Introducing Maryland’s AI Innovation Lab 

The AI Innovation Lab, an initiative of Maryland’s Department of Information Technology, streamlines the department’s AI services to help state agencies build AI tools and scale their AI adoption. “We want to be the place agencies come when they have a specific AI use case to solve a key problem their agency has,” said Maffeo.  

The Lab is deliberately designed to enable smart experimentation and safe deployment: Agencies bring a specific workflow or service delivery problem to the Lab, and in return, it provides subject-matter expertise, a secure technical sandbox and structured testing to evaluate whether AI can actually solve it. Solutions that work are built as reusable templates so that other agencies facing a similar problem do not have to start from scratch. 

With the Lab, Maryland is creating a hub and standards for statewide AI products, services and evaluation.  

Meet the AI Innovator

Lauren Maffeo is senior program manager of AI/ML at the State of Maryland where she leads the State’s AI Innovation Lab and Community of Practice. She co-wrote the first statewide guidance to use AI coding assistants and led the launch of AI.Maryland.gov.

She previously led service design efforts in the Coast Guard’s first Chief Data and AI Office, and co-led product on the founding Digital Service team at Maryland’s Department of Labor.  

Choosing Use Cases  

Vetting use cases before a pilot is approved is a way the Lab adds an early layer of product evaluation. “The first thing that we require, one of the two most important things that determine if an agency is ready to run a pilot with the Lab, is a clearly scoped problem that they believe AI can help them solve,” Maffeo said. “We want people to have a very specific agency problem, whether their staff or constituents face it, that we can use AI to help solve for them.” When the agency understands the problem that they need to solve, the Lab can help build and test potential solutions and meaningfully evaluate if they are solving the defined problem. 

“The other thing,” said Maffeo, “is that each use case to be piloted needs an executive sponsor from the outset. If they don’t have sign-off from a C-level person at their agency, a director of IT, someone who is committed to helping this pilot transition into production, we can’t run a pilot with them.” The goal, she said, “is to make each pilot successful enough that agencies invest in them for the long haul. To achieve that goal, we need executive buy-in from agency leadership at the beginning.” 

These requirements keep the Lab focused on finding workable, testable solutions to critical problems. It also helps save staff time and public money from being spent on a project with no identified path to lasting ownership or clear accountability if something goes wrong once it is live.

The First Use Cases 

The Lab’s first three use cases were selected because each addresses a specific problem that is common across state government. One agency wants to leverage AI to move data from an outdated database into a more user-friendly one. A second agency is building an employee-facing tool to sort and draft responses to public comments with the aim of getting residents faster, more consistent replies. A third is testing whether a Model Context Protocol server can connect to three old but essential systems so that case workers can find what they need without manually digging through multiple databases. The Lab helps to build these tools and then helps agencies implement pilot tests to gauge each tool’s accuracy, efficacy and fit for purpose.  

Alongside these first agency partnerships, the Lab is building a portfolio of AI products available to all state agencies. Their team has already shipped a few smaller AI tools, and made them available to Maryland state agencies:​​ “We have an inbox analyzer tool that our AI/ML product director built and a YouTube-to-Docs tool that auto-creates written content from summaries of public hearings and legislative meetings,” Maffeo said. “Juggling longer-term pilots that mature into ongoing products and achieving a high success rate of moving to production, while innovating on low-lift, high-impact, scalable solutions, is what we want to see in a good balance.” That smart balance may also lead to productive innovation, with agency-led projects becoming trusted templates and Lab-built services inspiring agency-side problem-solving.  

The goal for each pilot, Maffeo said, is to produce a successful product that agencies are equipped to own, continuously evaluate and maintain. She is also looking for repeatable solutions and patterns likely to be useful across state agencies. “I think using MCP servers as gateways to legacy systems that provide easier access to data is going to be something that’s repeatable. Also, the fact that we have a productionized tech stack for chatbots will be repeatable. And then the third agency partnership [mentioned above] is also a good example because extracting information from legacy systems is really not fun work, but it is critical, and I think the need for that is not going away.”  

Building things that get reused, not just things that work once, is central to how Maffeo defines success for the Lab as a whole. A replicable solution, like the MCP connector or a validated chatbot stack, means other agencies and the people they serve do not have to wait for the state to solve the same problem twice. And it means a responsible solution is scaled.  

The Lab’s current focus on building tested, adaptable AI tools to create internal efficiencies reflects what we’re seeing across AI-savvy government units: when agencies work better, service and mission delivery to the public improve. And when agencies collaborate, solutions scale and cost and time savings compound.

Testing Before Deployment, Evaluating for Readiness 

Every AI tool built in the Lab goes through comprehensive user and security testing before it can leave the closed virtual sandbox environment. For agency partnerships, department-side project managers work with Lab experts to define evaluation metrics for each use case, and then pilot testing is designed to ensure that each tool works as expected. During the building stage, Maryland’s AI Lab experts conduct testing with the people who will be using each tool. Those users are then consulted again during pilot testing. A tool doesn’t move to production if it doesn’t work as intended or for the intended user.  

For the internal tool built to sort and draft responses to public comments, user research and testing were conducted during the build stage and focused on gathering enough feedback to anticipate the questions staff will ask the tool so they can draft and program viable answers to those questions into the tool. Before the tool is deployed, as with all others moving through the Lab, it will need to pass security scans and testing.  

The Lab offers AI red-teaming exercises as part of its service agreement with agencies. This type of exercise, defined by the National Institute of Standards and Technology as “an exercise, reflecting real-world conditions that is conducted as a simulated adversarial attempt to compromise organizational missions or business processes and to provide a comprehensive assessment of the security capabilities of an organization and its systems,” helps find security weaknesses before a tool is deployed. A pilot isn’t ready to leave the Lab until it has passed these tests. 

These stages and forms of testing evaluate a tool’s fitness for deployment. 

Measuring What Reaches the Public 

Though Maffeo sees most of the Lab’s work currently focused on building internal-facing tools for staff, she shared a live example of what evaluated, public-facing AI can look like. Last year, the Department of Human Services launched a chatbot that answers residents’ questions about the state’s SunBucks food benefits program. “Testing has found that the bot gives an accurate answer 97% of the time,” Maffeo said. It’s a small but concrete data point that a resident-facing AI tool in Maryland has been tested against a real accuracy standard.  

For the Lab, Maffeo described a broader evaluation framework organized around three goals: how widely AI is adopted across state government, how it improves operations for an enabled workforce and, most directly relevant to residents, its effect on constituents. “Another core bucket that we plan to focus on for KPIs is constituent impact, where we provide public value using AI as a tool,” she said. The specific measures she named include reductions in error rates, the number of residents who get an answer without needing to contact a call center, and backlog reductions in high-volume areas like permitting and benefits processing. 

These are emerging standards, and an evaluation toolkit is in the works. “None of these are set in stone just yet,” Maffeo said, but they preview how the Lab is thinking about measurement as it launches with the expectation that the framework will mature considerably within a year. 

The Lab’s work overall is governed by Maryland’s responsible AI use policy that was published last year. “If people are ever in doubt about what’s allowed or not by way of AI, what types of pilots we pursue, which instances of usage we define as out of bounds, I always reference our responsible AI policy,” Maffeo said. 

Enabling Maryland staff 

To supplement and inspire work in the Lab, Maffeo and a colleague run weekly AI office hours for state employees that serve as an informal way for agencies solving similar problems to connect. “We just had AI office hours with an agency where they showed us their use case and asked: ‘Has anybody else in the state done something like this?'” she said. “We connected them to another agency that did something similar.” 

The Office of Enterprise Data is currently leading a complementary effort to train staff on data and AI literacy through the launch of its Data Academy. OED is designing asynchronous courses for staff, Maffeo said, “and state staff will have dedicated pathways they can follow based on where they are with AI.” Maffeo told us staff can also take skills-based assessments through their agencies, which gauge readiness for AI. “There’s a lot of broad anxiety around using AI and the most effective thing you can do is show people how to use it in their own context to improve their lives while also stressing that they need to be judicious about when it adds value and when it doesn’t,” Maffeo said.  

Enabling staff allows them to use and produce better-fit tools responsibly and smartly. “AI is most successful when staff apply their unique subject-matter expertise to use it in very specific ways that benefit their roles,” Maffeo said. It is the experts, enabled by tech support, who produce the tools they need most.  

When done responsibly, this layer of education and empowerment is a critical component of evaluation – staff can catch critical errors and use tools safely and securely.  

What Other Governments Can Learn 

Gate before you build. A clear problem and a committed sponsor form a checkpoint before a pilot starts, not after something has already reached the public. 

Treat tone as a risk decision, not just a style choice. Punitive-sounding guidance makes people feel fear and anxiety around the use of AI tools. Responsible, coordinated AI use depends on open communication and collaborative engagement. 

Build things that get reused. A working pattern, like a data connector or a chatbot template, helps keep costs down and improves successful delivery timelines for every agency and resident who would otherwise have to wait for the same problem to be solved (and paid for) twice. 

Say what you can’t do yet. Naming your limits manages expectations, keeps your standards and the reasons for them clear and helps set teams up for sustainable success.  

Lean on peer networks. Office hours, communities of practice and cross-continental fellowships gave Maffeo a quick way to compare Maryland’s approach with peers doing similar work, especially in a field where a single agency may have just a handful of people doing it at all. “If you’re doing any type of civic AI work, the one thing I’d advise a person to do is to find a community of peers doing the same work,” Maffeo emphasized. 

Questions That Remain

Maffeo is candid about what the Lab hasn’t yet figured out. 

How much centralization is the right amount? The balance between centralized security and agency-level access is being actively renegotiated, not settled. “It’s certainly a tricky balance. It’s going to be a work in progress,” she said. The Lab’s current approach to balancing technical control and access keeps the governance and security layer centralized at DoIT since the department is the provider and needs a baseline level of visibility and oversight. At the same time, she’s giving agency sponsors more direct access to user roles and permissions than they have had before.  

Can governance keep up with autonomous AI agents? Maffeo flags a growing risk: autonomous AI agent traffic to the state’s own website is already substantial and is increasing. Documentation now has to serve AI agents and bots as well as people. Left unmanaged, that growth ‘could become untenable,’ she said.

How do we manage capacity and budget, and set expectations, as a statewide service-delivery lab? The Lab has just two full-time staff and doesn’t charge agencies to run pilots, so it absorbs the cost of every pilot itself. Agencies sometimes hear “a free sandbox for six months” and push to expand the scope beyond what was first proposed, asking to add more databases or cover more programs than originally planned. “Juggling agency expectations is a big challenge,” Maffeo said. “We literally cannot afford to work on initiatives that don’t have these standards up front because we don’t have the personnel.” She added that the team is expected to grow substantially in the near future. Until then, the Lab’s commitments to agencies and the people they serve must be realistic.

Explore other case stories

Hear directly from other government practitioners and leaders doing this work.

Demos,
Not Memos

How Colorado Is Rethinking AI Evaluation

Read more

A Source of Authenticity

How the Library of Congress Approaches AI 

Read more

Building the Conditions for Responsible AI 

How Maryland is Approaching AI 

Read more