
AI in NGOs: from experimentation to a robust pilot
More impact, not fewer people — from quiet individual use to accountable teamwork in a Swiss NGO or NPO.
Situation assessment: clarify use, data and what happens to freed-up timeAfter three months, you have
- A situation assessment with questionnaire, micro-diary and process reconstruction
- A documented leadership decision with a genuine option to say no
- A pilot task with input, output, excluded data, and acceptance criteria
- A built pilot, measured against a baseline of the same task
- Evidence and a decision to scale up, adjust, pause or stop
Bringing AI into an NGO or NPO starts with a situation assessment and a time-limited, reversible pilot on a task you can name. Management and the governing body decide on use and data. The workflow is built with the focus on reviewing, followed by a comparison with the baseline. What is promised is not a productivity figure, not a guarantee that oversight works and not a general increase in impact — but a traceable decision on whether AI fits the task.
What is shifting right now
Work is shifting from doing to reviewing and approving.
People who used to write a text, sort a list or triage a request now increasingly describe what should happen and then judge the result. That sounds like relief, and it is — but it is a different activity from before, with different demands.
We are not selling you the idea that this turns everyone into a manager. For many people that would not be good news, just more responsibility without more pay. The sober version matters more: reviewing is a skill, and it does not develop by itself.
Why this is the core
Studies on the oversight of automated systems reveal an uncomfortable pattern: the more reliable a system appears in everyday use, the more uncritically its errors are accepted. A systematic review of 74 studies found a significantly increased risk of incorrect decisions when systems provided flawed suggestions. The statement “a human still checks it” is therefore not a safety measure, but an assumption that must be verified.
Why it still stalls
Not because people don’t want to. The causes are almost always organizational — and therefore changeable.
No one gave permission
Without explicit permission, anyone in doubt decides against using it — or uses it secretly. This is the most common cause and the easiest to fix.
No time to practice
Nobody learns something new between two meetings. Without protected time it does not happen, however good the training was.
Fear of appearing incompetent
In some places, people who use AI fear being seen as less capable. You cannot talk that away — it only helps when leaders show their own use, including the attempts that failed.
Uncertainty about quality
A fair concern. The answer is not an appeal to trust but an acceptance criterion: how do we tell that this result is usable?
Objections on the merits
Energy consumption, working conditions, copyright, credibility with donors. These objections are not a knowledge gap and cannot be resolved with productivity arguments.
Past projects
Many organizations remember a failed digital project. Ignore that and you repeat it.
Assess yourself first — 6 minutes
Thirteen questions in two parts: barriers in the team and whether your task suits a pilot. You get an evaluation right away — structured, no overall score, and your data is not stored.
- Duration: approx. 6 minutes
- No storage, no transmission — runs entirely in the browser.
What leadership must decide
This is not a formality, but a prerequisite. Until it is clear what happens with the time gained, the staff will hear the obvious answer — and then nothing moves.
- Whether to do it at all. With a real option to say no, not as a rhetorical question.
- What happens to freed-up capacity. In writing, before the pilot starts.
- No individual usage monitoring, no performance evaluation based on log data.
- Which data may leave the organization at which processing level.
- What remains excluded — particularly automated decisions about individuals.
- What ends the pilot, and who may call it off. Staff included.
We are not asking for an AI strategy. We are asking for permission to test a single workflow, time-limited and reversible. That is a decision even a volunteer board can take responsibility for.

What the built pilot covers — and what it does not
The term “built pilot” decides whether you buy. That is why we define the scope here, not only in the contract.
Included
- A specific, recurring task identified by the team itself
- A configured test environment with the tool selected for the pilot
- A baseline of the same task, measured before the start
- Written acceptance criteria before the first run
- A record of the decision on scaling up, adjusting, pausing, or ending
Not included
- Integrations into existing line-of-business systems without a separate mandate from your organization
- Operation, hosting, and ongoing security of the tools used
- Custom software development outside the pilot scope
- A rollout or training for all staff
- General AI training without a specific pilot task
- A guarantee of time, cost, or productivity gains
Organization’s responsibility
Data, access, operation, hosting, and ongoing security remain the responsibility of the organization. The pilot clarifies and documents these points; it does not assume them permanently.
Who it suits — and who it does not
A sober assessment before investing time and attention.
A good fit if
- there is a nameable, recurring task the team wants relief from.
- management wants to decide, before starting, what happens to freed-up capacity.
- the organization is willing to measure before the pilot — even if the result may be negative.
Not a fit if
- you are looking for general AI training.
- you want to buy a ready-made tool without clarifying the task and the decision path.
- the aim is to justify job cuts; this mandate explicitly does not do that.
The path
Six steps over about three months. The training course deliberately does not come first.
- 1
Situation assessment
Three sources instead of one survey: a short questionnaire on permission, concerns and wishes; a micro-diary over three to five working days; and the reconstruction of two real processes. Self-reports alone do not capture work reliably.
- 2
Leadership decision
A separate meeting for the board and management with three options: do not use, use within narrow limits, pilot with a path to scale. The result is a reasoned decision, not an ethics paper.
- 3
Task workshop
The team picks one task. On a single page: input, expected result, excluded data, acceptance criterion — and next to it the non-brief, i.e. what the system must not do. If you cannot state a result that can be checked, you do not have a pilot task yet.
- 4
Built pilot
We actually build the workflow, not a sketch of it. The baseline is measured first; otherwise there is nothing to compare afterwards.
- 5
Course: review before trust
Does not start with prompts. Participants get three results for their own case — a good one, a plausibly wrong one and an impermissible one — and decide with reasons: accept, revise, escalate, discard. Only then does it turn to technology.
- 6
Support and evidence
Office hours in weeks 2 and 6, each on a real case. After three months, a comparison against the baseline — and an explicit decision: scale up, adjust, pause or stop.
How we measure impact
Not with the question of how much time someone feels they have saved.
Self-reported time savings differ significantly from measured values, in either direction. Therefore, we establish a baseline for the same task before the pilot — cycle time, rework, errors, quality according to defined criteria — and compare the same metrics afterward. If the sample size allows, we add a control team or staggered start.
The outcome is not guaranteed to be positive. In the first weeks the net gain is often zero or negative, because learning time and review effort come first. We tell you this in advance so you do not read it as failure.
Honest limits
- No percentage. The available productivity studies come from other industries and show very different effects depending on the task and prior experience. Anyone who quotes you an average is transferring it without checking.
- No command line for everyone. It is useful for technical roles; as a mandatory step for all staff there is no evidence for it, and it mainly creates resistance.
- No substitute for professional expertise. Understanding the case, relationship work and responsibility stay where they are.
- Not every task should be delegated. Some take a lot of time and are still the core of a role. That belongs in the situation assessment — as a question, not as resistance.
Frequently Asked Questions
Will staff lose their jobs as a result?
Not if you decide otherwise beforehand — which is exactly why our process includes a decision on how freed-up capacity is used. In an NGO, needs exceed capacity anyway: what is freed up can flow into programs, client support or impact work. But that is your management’s decision, not a consequence of the technology, and it should be in writing before the first pilot starts.
Our staff are not particularly enthusiastic about technology.
This is usually not a question of enthusiasm. In comparable organizations, three out of four already use AI — just without rules and without anyone talking about it. What is missing is rarely curiosity, but permission, time, and a case worth pursuing.
We have concerns about energy consumption and working conditions.
These are legitimate objections, not misunderstandings. In the leadership decision we treat them as an option of their own, with a genuine choice to refrain, and we document the trade-off instead of overriding it with efficiency arguments. A reasoned decision against using AI is a valid outcome of our mandate.
Do all staff have to participate?
No. The security and approval rules apply to everyone; taking part in the pilot is voluntary. Mandated use combined with a punitive attitude to mistakes reliably leads to tools being used secretly and results being accepted unchecked.
We already have licenses. Isn’t that enough?
Access is not usage. A license does not answer which data may be included, who approves a result, how quality is measured, and what happens in the event of an error. That is where the work lies.
Where do we start?
With a recurring, checkable internal process — triaging requests, preparing impact data, classifying documents. Explicitly not with donor communication: people react sensitively to machine-generated personal messages, and your credibility carries the risk.
Situation assessment
In the initial consultation we clarify where AI is already used, which data is involved and who decides what happens to freed-up capacity. After that it is clear whether a time-limited pilot fits.
Request a situation assessmentBefore anything runs, a framework is needed: what is permitted and who is responsible.
What to settle before you startPhotos by Centre for Ageing Better and Annie Spratt on Unsplash.