How to Measure Whether a Workplace Workshop Was Actually Useful
Evaluation goes wrong at the first step more often than the last. A form gets sent because sending a form is what happens after training, the results are pleasant, and nobody can say what decision they informed. This is the method we use, the actual questions we ask attendees, and the part usually left out: a straight account of what our own data cannot establish.
Name the decision before you design the measure
Measurement is only worth its cost if a result would change something. So write down the decision first: whether to run this session for the other four teams, whether to change the format, whether the subject was the right one, whether to commission anything at all next year. Each decision needs different evidence, and some need surprisingly little. Deciding whether to repeat a session for a second cohort is well served by asking the first cohort whether the takeaways were usable.
Government evaluation guidance makes the same point in a more formal register: scope the evaluation alongside the intervention, not after it, and design it around the questions someone will actually act on. A dashboard built after the fact, from whatever data happened to exist, tends to answer questions nobody asked.
If no plausible result would change what you do next, the exercise is documentation rather than evaluation.
Evidence: 2
Four levels, and the causal chain they do not guarantee
The familiar framework distinguishes reaction (what attendees thought), learning (what they now know or can do), behaviour (what they do differently at work) and results (what changed for the organisation). It is a useful map of where you could look. Its well-documented weakness is that the levels are routinely read as a chain in which each one causes the next, and reviews of the model point out that this progression is assumed rather than demonstrated, with intervening variables ignored.
The practical consequence is blunt: a room full of fives at level one predicts very little about level three. Enjoyment and learning are not the same variable, and a session can be rated highly because it was entertaining, or rated lower because it corrected a belief the room was attached to. Most commercial workshop evaluation stops at level one and reports it in language borrowed from level four. Knowing which level a number belongs to is most of the skill.
What we actually ask, and why it is shaped that way
After a Brainwave session, attendees get a short form. It opens by identifying which session it was, then asks three questions on a one-to-five scale: the session overall, how usable the takeaways were, and how likely they are to recommend Brainwave. Then it asks two open questions (what stayed with you, and what would you change) plus what we should run next. Name and email are optional and the form says so; nothing after the three ratings is required.
Every one of those choices is deliberate and each has a cost. The ratings come before the writing because three taps are answerable while standing up, and a form whose first screen is an empty textarea is a form most people close. But ordering the questions this way anchors the free text against a rating that has already been given. Name and email are optional because feedback that costs someone their name is not the feedback we needed; the trade is that we cannot follow up on a specific comment or track one person's view over time. The “what would you change” box carries a note asking people to be blunt, because the polite half of a feedback form is the half that teaches you nothing.
- Three 1–5 ratings: overall, usability of the takeaways, likelihood to recommend.
- Two open questions: what stayed with you, and what would you change.
- A “what should we run next?” selection that shapes the following session.
- Name and email optional; permission to quote asked explicitly, on the spot.
What that data cannot tell you
It is collected immediately after the session, from people who chose to respond, with no comparison group and no follow-up. That means it measures reaction and self-rated usefulness: level one, and the thinnest edge of level two. It cannot tell you whether anyone slept better, whether a team's focus improved, or whether anything about the way the work is organised changed a month later. Anyone presenting numbers of this shape as evidence of behaviour change is overreading them, and that includes us if we ever do it.
Three specific biases are worth naming. Response bias: the people most and least engaged answer at different rates, so the mean describes respondents, not the room. Immediacy: a rating taken minutes after a good session captures the session's felt quality, which decays at a different rate from what was learned. And the absence of a counterfactual: without a comparable group who did not attend, a before-and-after difference cannot be separated from everything else happening that month. The systematic-review literature on workplace interventions is full of exactly these design limitations, which is the honest context for any single provider's feedback data.
- Self-report, self-selected, immediate, no control group.
- Measures reaction and perceived usefulness, not behaviour or organisational outcomes.
- The mean describes who answered, not who attended.
- Useful for improving the next session. Not evidence that work changed.
The measurement your organisation has to own
The parts of the picture a provider cannot see belong to you, and they are cheap to collect if you decide in advance. Pick one or two indicators tied to the specific change you hoped for, take a baseline before the session, and measure again six to eight weeks later, long enough for the glow to fade and any real change to appear. If a focus session was paired with a rule about meeting-free mornings, count the meetings. If a sleep session was aimed at a shift pattern, look at the roster alongside people's own reports. Broader instruments such as the NIOSH Worker Well-Being Questionnaire are designed for this kind of understanding, on the explicit condition that they are not clinical judgements.
Pair the workshop with one change to how the work is organised, and evaluate the pair. This is what occupational-health and NICE guidance both point at: organisational-level conditions come first and individual approaches sit on top of them. It also makes the evaluation more honest, because you are now measuring something you actually did rather than something you hoped people absorbed. And keep the trust conditions intact: say what is collected, protect small groups from being identifiable, and report back what changed. A survey with no visible consequence teaches people that answering is pointless.
- Baseline before, repeat six to eight weeks after.
- One or two indicators tied to the change you actually wanted.
- Change one work condition alongside the session, and evaluate the pair.
- Protect anonymity in small groups; report back what you did with the answers.
Why we publish no ROI figure
We are asked for one, and other providers supply them. Producing a defensible return figure for a single workshop would mean isolating its effect from hiring, workload, leadership changes, seasonality and every other wellbeing initiative running that quarter, with a comparison group. That is a research programme, and it is not what a proposal document contains. The numbers in circulation are typically built by multiplying a self-reported improvement by an assumed value, which produces a precise-looking result from two soft inputs.
So the claim we do make is narrower and, we think, more useful: attendees consistently rate the takeaways usable, they tell us what to change and we change it, and a session works best paired with something the organisation itself alters. If you need a return figure to get approval internally, we would rather help you design a real before-and-after comparison you own than hand you a number we cannot stand behind.
A number nobody can reconstruct is decoration, however precise it looks.
Read the research
Sources
- 01Kirkpatrick and Beyond: A Review of Models of Training Evaluation (IES Report 392) Institute for Employment Studies
- 02
- 03Evidence of Workplace Interventions: A Systematic Review of Systematic Reviews International Journal of Environmental Research and Public Health
- 04NIOSH Worker Well-Being Questionnaire (WellBQ) CDC / NIOSH
- 05Mental wellbeing at work (NG212) National Institute for Health and Care Excellence
This article is for general education, not diagnosis or medical advice. If a health concern is affecting you, speak with a qualified professional.
How we work with evidence
