From Curiosity to Confidence: A Practical, Human-Led Approach to Understanding AI Agent Success
At Blackbaud, our research helps connect what AI can do with what social impact professionals and the people they serve actually need. That means starting with the work people are trying to accomplish, the relationships they are building, and the challenges they are navigating, then asking whether AI is genuinely helping.
We have been applying that approach to Agents for Good™, and to a question that sounds simple but quickly opens up wider considerations around trust, donor experience, organisational readiness, and accountability:
How do you know whether an AI agent is genuinely helping your organisation?
The answer can’t be reduced to how many tasks an agent completes. Success depends on whether the work contributes to meaningful outcomes, whether it supports the people involved, and whether the organisation can understand and trust how that work is being done.
Free Resource
From Curiosity to Confidence: A Practical, Human-Led Approach to Understanding Agent Success
Across Development Agent customers, where the agent is assigned, our data signals that donors are engaging more, with a 3.6% reply rate compared with an industry email response rate of around 0.05%. And across more than 70,000 unique conversations initiated, these are a sample of the return-on-investment signals we keep tracking so we can iterate and improve the experience.
These aggregate figures are encouraging, but they are a starting point rather than a verdict. A single headline number can’t be looked at in isolation. Each organisation should decide which signals are meaningful in its own context. The rest of this post looks at how customers are defining what success means to them, and why careful, human-led interpretation of the evidence is what helps build trust responsibly over time.
What Development Agent supports
Development Agent is an AI agent from Blackbaud designed to help fundraising teams engage supporters more effectively. It identifies outreach opportunities, builds engagement plans, drafts personalised communications, supports follow-up, and involves a human colleague when needed. It also provides teams with information about performance.
The starting point for this work is a practical challenge familiar to many fundraising teams: relationships need timely and thoughtful attention, while staff are often managing large portfolios, competing priorities, and limited time.
Our Early Adopter Program customers described that challenge clearly:
“We’re trying to get more people back to engage with the university… generally just to reach more people, have more touch points with people that we can’t necessarily get to or who are just getting general communication. This way, they’re getting something more regular and it feels more personalized.” –Development Agent customer
“For us it’s about allowing the agent to have that interaction that I cannot find the time to do. Not because there isn’t a want, but because we’re juggling so many different things and I simply don’t have the resources or manpower to get it all done. So really, that’s it at the end of the day.” – Development Agent customer
These customers were interested in whether an agent could help with work they already valued, at a level of consistency they could not always reach.
Customers have data. The challenge is knowing what it means
To understand how customers would evaluate Development Agent, we conducted in-depth prioritisation sessions with Early Adopter Program customers. The sessions explored which metrics mattered, why they mattered, and how they might be presented in a useful dashboard.
We expected that limited access to data might be one of the main barriers.
The research pointed to a different challenge. Customers were often already generating plenty of data. The more difficult task was identifying which signals were meaningful, how those signals should be interpreted, and which decisions they could support.
This distinction is important for responsible AI. More measurement does not automatically create greater understanding. Teams need a focused set of indicators that helps them understand whether an AI-enabled service is contributing to positive outcomes, identify possible concerns and provide enough evidence to continue, refine, or change course.
The wider sector context reinforces this need. Blackbaud Institute research found that 85% of social impact professionals surveyed were using AI at work, while only 30% reported having a formal AI policy. Only around one-third believed their organisation was using AI very effectively.
As adoption grows, organisations need ways to move from experimentation towards informed and accountable use.
An additional benefit: using agent data to learn
Customers were also considering how information generated through Development Agent could help them understand their wider fundraising practices.

“Am I, as Director, able to make smarter choices that inform all of our fundraising?” – Development Agent customer
“One of the things that we’re hoping to learn from the Development Agent is what is effective and how can we then make that relate to what our Gift Officers are doing?” – Development Agent customer
For some organisations in the research, the agent was creating a reason to examine information they already had and start new conversations about communication, engagement and team practices. Customers wanted to understand what was working through the agent and consider how that learning might inform their wider fundraising strategy.
That creates value beyond reporting. Measurement can give teams another set of data points to support learning and improving decision-making, provided the data is interpreted carefully by people, and in context.
Three ways customers think about agent success
When we analysed the metrics customers prioritised, they clustered into three connected ways of understanding value.
Together, these lenses ask:
- Is the agent contributing to meaningful outcomes?
- Is supporter engagement improving?
- Can we trust how the work is being carried out?
1. Fundraising outcomes
These are the outcomes likely to matter to leadership and boards:
- Funds raised through agent-supported outreach
- Donor retention year over year
- Growth in recurring donors
- New or reactivated donors
Customers recognised that these outcomes take time to develop. Relationship-building can’t always be assessed through immediate financial return. However, teams still need enough early evidence to understand whether their approach is moving in a useful direction and whether continued investment is justified.
“The ROI and time to value in all reality… I feel like all of this and anything that we deploy, whether it’s a Development Agent or it’s a new donor touchpoint plan, that there always has to be time for it to bake.” – Development Agent customer
A practical starting question is:
What outcome would help our leadership or board decide that this is worth continuing to explore?
2. Donor engagement
Engagement metrics provide earlier feedback while longer-term fundraising outcomes are still developing. Customers highlighted:
- Email open and click-through rates
- Donor reply rates
- Engagement by audience segment, campaign or message type
- Unsubscribe and opt-out rates
Customers were particularly interested in what these signals could tell them about their next communication. A reply rate may help a team understand whether outreach is prompting meaningful interaction. Engagement by segment can show where an approach may need adapting. Unsubscribe rates can act as a guardrail for fatigue, poor timing or lack of relevance.
“I want to look at this in the sense of the engagement part… that’s what is most important to me. How can we not have people slip through the cracks?” – Development Agent customer
These measures should also be considered alongside direct feedback from supporters. People may notice issues with purpose, relevance, tone, or transparency that a dashboard cannot show.
A useful question is:
Which signal would help us improve the next communication, rather than only report on the last one?
3. Operational efficiency and quality
This lens considers whether the work is being carried out in a way the organisation considers appropriate and trustworthy. Customers discussed:
- Content accuracy and relevance
- Alignment with organisational voice and values
- Actions completed
- Escalation to a human colleague when appropriate
“I want to make sure that content is accurately represented. These people are our neighbours, they are our friends, anyone is a moment away from needing the support of the services that we offer.”- Development Agent customer
Actions completed can indicate capacity, but activity alone says little about whether a relationship was supported or a useful outcome followed. Customers therefore cautioned against allowing this to become a vanity metric.
Escalation also needs context. A higher escalation rate doesn’t automatically indicate poor performance. It may show that the agent is appropriately identifying situations where a person should become involved. The quality, timing and relevance of the hand-off are more informative than the number alone.
A practical question is:
What would need to be true for our team to trust how this work is being carried out?
Donor trust belongs in the measurement framework
Agent success also needs to be examined from the supporter’s perspective.
Blackbaud Institute research found that 76% of donors surveyed considered it important for organisations to disclose when and how AI is used, while only 26% of organisations reported doing so. The research found disclosure had a positive effect on trust across generations.
For agent-enabled communications, this means considering questions such as:
- Is the purpose of the communication clear?
- Does the message reflect the organisation’s voice?
- Is the use of AI communicated in a clear and proportionate way?
- Can the supporter easily opt out?
- Is there a clear route to a person when one is needed?
- What are supporters telling us about their experience?
Transparency should help people understand where AI is supporting the interaction, where accountability sits, and what choices are available to them.
Metrics to handle with care
Customers deprioritised several measures, particularly during the earlier stages of adoption:
- Mid-level to major donor conversion: a longer-term and sensitive area involving high-touch relationships
- Average gift size and large gift counts: potentially striking numbers that may reveal little about overall programme effectiveness
- ROI and time to value: important over time, but difficult to interpret before sufficient evidence has developed
- Touchpoints per donor: often shaped by a predetermined cadence in set-up of the agent
- Agent response time: generally assumed to be fast unless a problem occurs
These metrics may become useful later. Their meaning depends on timing, context and the decision an organisation is trying to make.
A simple check can help:
Are we measuring this because it will inform a decision, or because it is easy to count?
Questions to take back to your organisation
If your organisation is evaluating an agent, or another AI-enabled service, or still unsure, these questions can help structure the conversation around responsible AI practice and value:
- What problem are we asking AI to help with?
- What should improve for supporters?
- What should improve for staff?
- What would success look like to leadership or the board?
- Which data do we already have but not fully use?
- How accurate, complete, and well-governed is the data the AI will use?
- Where should human judgement and accountability remain visible?
- What would make us pause, refine the approach, or change course?
- Which early signals and longer-term outcomes should we examine together?
For each metric you’re considering tracking, ask three further questions:
- What decision will this help us make?
- Who needs to see it?
- What context will help prevent us from misreading it?
Building confidence through evidence
Agent success is best understood through a set of connected signals. Outcomes show whether the work is creating value over time. Engagement data provides earlier feedback. Quality signals highlight whether the work is being carried out in a way the organization can trust over time. Human judgement runs through all of these—the data shows what is happening, but people decide what it means and what to do next.
Start with a real problem. Select a small number of meaningful measures. Use the evidence to learn and adapt. Keep donor trust and human accountability visible throughout.
That is how organizations can be intentional when moving from curiosity towards greater confidence in AI agents: by asking better questions, interpreting evidence carefully, and remaining open to what they still need to learn.
