The heat map is on the screen and two teams are in red. Everyone in the room understands that something now has to happen, and nobody in the room knows what. The consultant proposes a facilitated session. It gets scheduled for a Thursday. This is the point at which most psychological safety work ends without anyone saying so, and it is worth being precise about why.
The evidence on what to do about a low psychological safety score is thin and mixed. A systematic review of fourteen interventions found inconsistent results, and only three of the fourteen actually measured psychological safety as an outcome. There is also direct evidence that a poorly run debrief lowers a team's psychological safety rather than raising it.
So the honest position is that the measurement industry has built a good instrument and a weak response, and it sells them as one thing.
What Does the Evidence Say About Fixing a Low Score?
Less than you would hope, and the numbers are worth holding in mind. Róisín O'Donovan and Eilish McAuliffe screened nearly nine thousand studies for their 2020 systematic review in BMC Health Services Research. From that, they found fourteen interventions anywhere in the healthcare literature that had tried to improve psychological safety, speaking up or voice behaviour. Fourteen.
Their results were mixed. Some interventions improved something. This was not consistent across studies. And the finding that ought to be quoted more often: of those fourteen interventions, only three evaluated psychological safety as an outcome at all. The rest targeted it and never checked.
One randomised experiment in the set gives a sense of how stubborn the problem is. Raemer and colleagues trained anaesthesiologists to speak up in the operating theatre, then tested them in simulation. Only a quarter of participants spoke up to a colleague of their own discipline. This was after the training, in an exercise they knew was an exercise.
The review's own conclusion on the educational approach is that teaching people about speaking up does not reliably change speaking up, because the behaviour is rooted in something education does not touch.
Can a Debrief Make Things Worse?
Yes, and there is a study on it that almost nobody in this market cites.
Dufresne simulated a critical incident with anaesthesia teams and examined what the person leading the debrief did. The finding, reported in the same review, is specific and uncomfortable. When leaders made negative evaluative statements, the team's psychological safety was lower. When leaders balanced advocacy and inquiry language in the first ten minutes, psychological safety was also lower. The recommendation that follows is that whoever leads the session should avoid evaluative statements about team or individual performance early on.
Read that again with a Thursday session in mind. The first ten minutes of a debrief about why your team does not feel safe can reduce how safe your team feels. Not through malice. Through an inexperienced facilitator saying something reasonable at the wrong moment.
A psychological safety debrief is the only intervention we know of that can be disproved by its own delivery. Run it badly and you have not failed to help. You have supplied fresh evidence for the thing people were already afraid of.
This is the argument for a clinician in the room rather than a facilitator with a deck. Not because clinicians are wiser. Because controlling depth, protecting an individual from exposing more than they intended in front of their manager, and keeping the unit of discussion at the level of the pattern rather than the person are trained skills, and the study above is what happens when they are absent.
Why Do Team Level Fixes Fail?
Because the team is usually not where the problem is held. O'Leary ran action research meetings with two newly formed teams. In one, a psychologically safe space developed. In the other, it did not. The difference was not the facilitation. It was organisational norms and team stability: shared decision making as a company habit, and a stable core group of people.
Which means a team can be facilitated skilfully and still fail to become safe, because the thing preventing it sits above the team. The review says as much in its recommendations, which include targeting the organisational level as well as the team, securing visible leader support, and running interventions over longer periods rather than as single events.
Notice what that rules out. A one day workshop for the two red teams, delivered by an external, with no change above them and no follow up, is the intervention the evidence least supports. It is also the intervention most commonly sold, because it is the one that fits a budget line and threatens nobody's calendar. We have written about the same pattern in workplace wellbeing generally, and it is the same pattern.
Who Is the Red Score Actually Made Of?
People, and this is the part the heat map is designed to help you forget.
A team scores in the red because several individuals answered honestly that it is not safe here to admit a mistake, ask for help, or say the difficult thing. Those are not abstract attitudes. Someone answered that way because of a specific meeting, a specific review, a specific manager. And having answered, they are now waiting to see what the company does with it.
You have also just run a disclosure exercise. Asking a group about fear and silence surfaces distress, and some of what surfaces will belong to people who are struggling well beyond work. If your process ends at a report, you have collected that and filed it. The Mental Healthcare Act, 2017 governs what a clinician may do with such information, and an employer is not on the list of exceptions. So the route out has to exist, it has to be private, and it has to reach an actual clinician rather than a resources page.
We sell that route. Read the paragraph above knowing it.
What Would a Serious Response Actually Look Like?
Six things, drawn from what the evidence supports rather than from what is easy to invoice.
Decide how leadership will receive a bad result, before anyone answers. A leadership team that punishes a red score has confirmed it. Say this out loud in advance, because that statement is the first intervention and it is free.
Do not let an untrained person lead the session. The Dufresne finding is the whole reason. Whoever holds the room needs to be able to hear something heavier than the format expects and manage it in one sentence.
Build in the right to pass. One of the few structured tools in the review, the CENTRE guidelines, includes confidentiality, equal airtime, non-judgemental listening and an explicit right to say nothing. Note that the review also records that no formal evaluation of that tool had been published, which is the state of this field.
Change something above the team. Workload, meeting structure, how errors are handled, who gets credited. If nothing at that level moves, the team has been asked to feel differently about unchanged conditions. Most of those levers sit in the psychosocial hazard territory that ISO 45003 describes.
Re measure, and report the teams that did not move. Use the same instrument on the same response format, per team, with a floor on group size. We have set out what the instrument can and cannot tell you, and how to run the measurement honestly.
Keep the care line open longer than the engagement. The person who decides to get help six weeks after the session is the point of the whole exercise, and they will not do it during the workshop.
None of this resolves neatly. The evidence base for improving psychological safety is fourteen studies with mixed results, three of which measured the outcome, in a literature with no longitudinal work. Anyone quoting you a guaranteed improvement is quoting you something that does not exist yet.
What is knowable is smaller and more useful. A red score tells you where people have stopped speaking. The response either changes what happens when they do, or it does not. Everything else is a Thursday.
What Else Do People Ask About Low Psychological Safety Scores?
What should you do if a team scores low on psychological safety?
Start by deciding how leadership will respond publicly, because punishing a low score confirms it. Then run a facilitated session led by someone trained to manage disclosure, change at least one condition above the team such as workload or how errors are handled, keep a private route to professional support open, and measure the same team again on the same instrument after a stated period.
Do psychological safety interventions actually work?
The evidence is thin. A 2020 systematic review by O'Donovan and McAuliffe found only fourteen relevant interventions across the healthcare literature, with mixed results, and only three of them measured psychological safety as an outcome. The review concluded that education alone does not reliably change speaking up behaviour and that longer, multifaceted interventions targeting the organisation as well as the team are needed.
Can a psychological safety workshop backfire?
It can. Research by Dufresne, reported in the same review, found that psychological safety was lower when debriefing leaders made negative evaluative statements, and also lower when leaders balanced advocacy and inquiry language in the first ten minutes. The recommendation is that session leaders avoid early evaluative statements about team or individual performance.
Why did our team score improve but nothing changed?
Because a score can move on attention alone. Being asked about something is itself an event, and early readings often reflect the novelty of being consulted rather than a change in conditions. This is why a baseline, a stated period, and a repeat measurement matter, and why a single post intervention reading should not be treated as evidence of improvement.
Who should facilitate a psychological safety debrief?
Someone trained to manage what surfaces. In these sessions a participant will occasionally disclose more than the format was designed for, in front of colleagues and sometimes their manager. Managing that safely means controlling depth, protecting the individual from over exposure, keeping discussion at the level of the pattern rather than the person, and following up privately afterwards.
A note on the cover image
The image at the top of this piece was generated by AI, to a brief written by us. It is not a photograph and does not depict a real person or place.
Sources
O'Donovan, R., & McAuliffe, E. (2020). A systematic review exploring the content and outcomes of interventions to improve psychological safety, speaking up and voice behaviour. BMC Health Services Research, 20, 101. Open access, CC BY 4.0. Fourteen interventions reviewed from 8,947 records screened. Full text
The findings attributed here to Dufresne (2007), Raemer and colleagues (2016), O'Leary (2016) and Cave and colleagues (2016) are reported as summarised in the O'Donovan and McAuliffe review above, which we read in full. We have not read those individual studies directly.
Edmondson, A. C. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350-383. DOI 10.2307/2666999. Paywalled.
The Mental Healthcare Act, 2017 (Act No. 10 of 2017), Section 23. Full text, India Code
Nothing in this article is legal advice. Take your own.







