Right About the Room
Taking criticism badly gets diagnosed as thin skin. I think the flinch is a piece of learning instead: it reads the room correctly, because criticism really does buy its author something, and then it files that answer in the slot reserved for whether the sentence is true. What that costs you is not growth. It is that people quietly stop telling you things, and the silence looks exactly like agreement.
Almost everything written about taking criticism badly treats it as a defect of character. Thin skin. A fragile ego. Not enough growth mindset. Someone with work to do on themselves. The diagnosis varies and the prescription never does, and it always comes out as some version of relax, it isn't personal.
That has it backwards. The flinch is not a failure of nerve, it is a piece of learning. It gets assembled over years of watching what criticism actually does in rooms that have other people in them, and about the thing it was built to detect it is usually right. What goes wrong is what it does next with a correct answer.
It shows up in three shapes and we file them as three personalities. Someone goes quiet and stays quiet for a fortnight. Someone shuts you down inside four words, before the sentence has finished arriving. Someone argues back point by point with an energy wildly out of proportion to the remark that started it. The three look nothing like each other and they are the same reflex, which is why advice written for one of them never works on the other two.
What it is picking up
In 1981 Teresa Amabile took two book reviews out of the same Sunday edition of the New York Times, written by the same critic. One was glowing, one was savage. She trimmed them, presented them as two editors writing about the same invented novel, and asked fifty-five students what they made of each reviewer. The savage one came out ahead on intelligence and competence, and a long way behind on kindness, fairness and likeability.1
There was a hole in that study and Amabile is the one who points at it. A separate group of raters judged the positive review the better piece of writing, better organised and more forceful. So she rebuilt the materials: the reviews rewritten word for word into matched positive and negative versions with each phrase swapped for its opposite, pretested until neither version read as the better-written one. A hundred students this time. The negative reviewer was still rated significantly more intelligent, and still less kind, less fair, less open-minded, less likeable.2
Criticism buys perceived intelligence. It is cheap to produce, available to anyone, and it pays on delivery. The same exchange rate shows up in a nastier currency: asked who should handle a set of cognitive tasks, people handed around seventy per cent of them to the cynic, in a literature where cynicism is negatively correlated with measured ability.3 Sounding unimpressed is one of the most reliable ways to look intelligent ever discovered, and it requires nothing of you.
That bonus, though, is paid by the audience. It does not survive contact with the target. When the criticism is aimed at the person forming the impression, or at someone they have just been talking to, the effect turns over, and the critic comes out rated less likeable and less intelligent than someone who praised.4
The critic and the person being criticised are reading the same event off two different instruments, and both instruments are working. The room is impressed. The person being described is not. Neither of them is imagining it, and the flinch is roughly what an accurate reading feels like from the wrong end.
How much of it is even about you
There is a second thing the reflex has learned, and it learned it from counting.
When 2,350 managers were each rated on three dimensions by seven raters, two bosses, two peers, two subordinates and themselves, the raters' own idiosyncrasies accounted for 62 per cent of the variance in the scores. Everything the model could pin to the person being rated, the part of the score that held steady across raters, accounted for 21. A second sample of 2,142 managers came in at 53 and 25.5 Hand someone a performance rating and most of what they are holding is a description of whoever filled it in.
Then there is the question of whether the feedback does any good. The best-known meta-analysis of the field pooled 607 measured effects and found that feedback interventions improved performance on average by 0.41 standard deviations, which sounds like a settled case until you read the next line: more than 38 per cent of those effects were negative. The feedback group scored below its comparison group.6 The obvious explanation is wrong, too. It is not that the harsh ones backfire and the kind ones work. Once one over-represented author's studies were pulled, the sign of the feedback stopped being a significant moderator at all, and what predicted damage was how far the intervention dragged the recipient's attention off the task and onto themselves.
Multi-source feedback, the whole 360-degree apparatus that organisations spend real money on, moves subsequent ratings by 0.15 from direct reports, 0.15 from supervisors, and 0.05 from peers.7
Put it together and the defensive person is not being irrational. They have a working model, and the model has data behind it. Most of any given verdict is about the person delivering it. Better than a third of the measured effects in that meta-analysis came out negative. And the sentence coming at them is buying its author something, in front of an audience, at their expense.
Two questions, one answer
The error is small, specific, and nothing to do with character.
There are two independent questions on the table. Is this person scoring points? And is the sentence true? The flinch answers the first one, fast and usually correctly, and then files that answer in the slot reserved for the second. That is the whole defect. Not fragility, a filing mistake, and one that a very intelligent person can make at speed for forty years.
The two are independent because they were never connected. A status move can be accurate. Someone can be showing off, enjoying it, aiming for the soft part on purpose, and still be right about the thing, and their motives leave the truth value exactly where it was. The reverse holds as well, which is the half nobody defends against: a generous, careful, kindly-worded note from someone with no stake at all can be completely wrong, and it arrives with none of the alarms going off.
The folk theory has to go, because most of the advice is built on it. Defensiveness is not the tell of a person with low self-esteem. When Bushman and Baumeister had participants insulted over an essay they had written, aggression toward the evaluator went up across the board, more sharply among the narcissistic, and trait self-esteem predicted nothing whatsoever.8 Nor is wanting to be flattered the driver. When the desire for accuracy is pitted against the desire for praise, the meta-analytic split is that people's feelings go to whoever flatters them while their judgments and their actual choices go to whoever confirms what they already believe about themselves.9 The machinery is not fishing for compliments. It is defending a self-description, and it will defend an unflattering one just as hard.
The room adjusts
The cost is usually described as missed growth, which is both true and useless, because nobody has ever changed a habit to collect a benefit that abstract.
Telling you something already costs the person who tells you, before you have reacted at all. Eleven experiments found that people deem innocent bearers of bad news less likeable. In the preregistered lead study the messenger merely announced a random draw they visibly had no control over, and was liked less for it. The penalty shrank when recipients were made aware the messenger meant well, which is to say it never disappeared, it only got cheaper.10
So there is a baseline price for opening your mouth, and every flinch you produce is a surcharge on top of it. People are quick to notice the new rate. Given a choice between material that supports what they already think and material that challenges it, people take the supportive option roughly twice as often, and that preference gets stronger the more committed to the position they are, disappearing altogether among those who are not.11 They will pay money to not find out.12 Your colleagues are running the same calculation about you that you are running about the evidence, and they are running it more often.
Code review, postmortems, incident reviews: my trade built a whole ritual apparatus for writing down what went wrong, with a named rule that nobody gets blamed for it. None of it exists because engineers are unusually wise. It exists because organisations discovered, expensively, that people conceal problems when naming them costs something. I have spent twenty minutes arguing with a review comment and then made the change anyway, and the twenty minutes went somewhere other than deciding whether the comment was right.
What makes this hard to catch is that you never experience the loss, because what you lose is a thing that does not happen. There is no notification for the email nobody sent. The meeting where three people decided separately that it was not worth it looks, from the inside, exactly like a meeting with no objections.
Grading the advice
Most of the standard advice does not survive contact with the evidence.
Growth mindset interventions move academic achievement by 0.08 standard deviations across 43 effects, and by a non-significant 0.02 when the comparison is a passive control group; the largest preregistered national trial lifted lower-achieving students by a tenth of a grade point and did nothing detectable for anyone else.13 Self-affirmation exercises, where you first reflect on something you value about yourself, produce a 0.17 improvement in how far people accept threatening information, on a confidence interval that comes within a hair of zero, and a well-powered replication of the school-based version, in the same district and running the same protocol, found nothing at all, ruling out any benefit above a tenth of a standard deviation.14 They are real effects, and they are too small to build a practice on.
Against that, four things hold.
Wait. Cortisol does not respond to every kind of stress. The big response comes from one specific combination: a task you cannot control, performed where you are being judged. That pairing produced an effect nearly three times that of either ingredient alone, and it was the only condition where cortisol was still elevated 41 to 60 minutes after the thing had finished.15 In the hour after criticism you are not thinking about it, you are metabolising it. Nothing true expires overnight. No experiment tests whether the wait improves what you then do with the criticism, so treat that part as an inference from the physiology rather than a finding.
Take the person off the sentence. Write the claim down with no author attached and ask whether you would accept it from someone you liked. This sounds like a trick and it is a trick, but it separates the two questions by force, which is the only operation that actually needs performing.
Count sources rather than volume. Given that most of a single verdict is the person giving it, one critic is a data point and should be weighted like one, no matter how loudly it was delivered. Three people with no connection to each other saying the same thing is a different object entirely. Most people do this backwards, treating one memorable attack as a verdict and three mild independent mentions as noise.
And the fourth is inconvenient, because it is not addressed to the person receiving anything. The most reliable manipulation anyone has found for making critical feedback land is a change in what the critic does. Attach a note to a piece of harsh feedback saying the standard is high and you believe the person can reach it, and revision rates among the Black students in that study went from 17 per cent to 71.16 That study is 44 seventh-graders and should be carried gently. But it points the same way as the messenger research, where the penalty for bad news falls when motives are visibly benevolent. A large part of what gets sold as a skill for receiving criticism is really a duty of the person giving it, and it gets sold to the receiver because the receiver is the one who feels bad.
When not to absorb it
Not all of it should be absorbed, and an essay that argued otherwise would be describing a room nobody actually works in.
Some of what arrives is a fact about your position rather than your work. Speaking up does not pay uniformly: across a field study and an experiment, suggesting improvements went with higher peer-rated standing and, through it, with emerging as a leader, while flagging problems did not, and the benefit tilted toward men.17 The popular version of this, that women get systematically vaguer feedback, is genuinely contested. The largest test of it, a preregistered randomised trial covering 46,176 feedback responses about 4,328 senior managers, found women receiving more specific feedback about strengths and less about development, with no overall gender difference, and the intervention built to fix the bias failed.18 The headline version does not survive that test. Filing a piece of criticism as information about the room rather than about yourself is sometimes the correct filing, and the fact that it is also the favourite excuse of everyone who has never learned anything is not an argument against doing it when it's true.
The motive and the content are two separate readings, and the first is not evidence about the second. I take the flinch seriously as a report about the room. I no longer let it vote on whether the sentence is correct.
The reason to bother is not self-improvement. It is that the alternative degrades quietly and does not announce itself. The room gets easier, and the meetings start going well.
There is no way to check this from the inside. The people who would have told you are the ones who stopped.
Teresa M. Amabile, "Brilliant but Cruel: Perceptions of Negative Evaluators," Journal of Experimental Social Psychology 19, no. 2 (1983): 146–156. The figures here are from the pre-publication version of the same two studies, presented at the Eastern Psychological Association in April 1981 and deposited as ERIC ED211573; the published article is behind a paywall at DOI 10.1016/0022-1031(83)90034-3. Study 1, N = 55: intelligence F(1,53) = 4.89, p < .05; kindness F(1,53) = 49.03, p < .001; likeability F(1,53) = 35.48, p < .001. On the 40-point scales the intelligence gap is 24.51 against 27.69 and the likeability gap is 25.18 against 12.47, so the effect on perceived intelligence is real and much the smaller of the two. ↩
Amabile, Study 2, N = 100, ERIC ED211573, pp. 10–16. Intelligence F(1,96) = 20.71, p < .001; kindness F(1,96) = 320.99; likeability F(1,96) = 64.44; open-mindedness F(1,96) = 33.86. Intelligence was the trait on which, in Amabile's words, "there were no significant effects other than the main effect of valence" — the parallel advantages on competence and literary expertise held only under some presentation orders. ↩
Olga Stavrova and Daniel Ehlebracht, "The Cynical Genius Illusion: Exploring and Debunking Lay Beliefs About Cynicism and Competence," Personality and Social Psychology Bulletin 45, no. 2 (2019): 254–269. PMC6328999. Participants allocated roughly 70 per cent of cognitive tasks to a cynical rather than a non-cynical target, while across large international datasets cynicism runs negatively against measured cognitive ability and education. ↩
Nurit Tal-Or and Ayelet Gal-Oz, "Evaluating the evaluator: The effects of the valence, sequence and target of evaluations on the perception of the evaluator," International Journal of Psychology 54, no. 1 (2019): 80–87. DOI 10.1002/ijop.12425; PubMed. ↩
Steven E. Scullen, Michael K. Mount and Maynard Goff, "Understanding the Latent Structure of Job Performance Ratings," Journal of Applied Psychology 85, no. 6 (2000): 956–970. DOI 10.1037/0021-9010.85.6.956. Each manager was rated by two bosses, two peers, two subordinates and themselves. Random measurement error was counted separately again, at 11 per cent in the first sample and 18 in the second. ↩
Avraham N. Kluger and Angelo DeNisi, "The Effects of Feedback Interventions on Performance: A Historical Review, a Meta-Analysis, and a Preliminary Feedback Intervention Theory," Psychological Bulletin 119, no. 2 (1996): 254–284, at 258, 273 and 277. DOI 10.1037/0033-2909.119.2.254; full text. The share of negative effects falls to 33 per cent when the 91 effects from a single over-represented author are dropped and to 32 in the fully trimmed set of 470, where the observed variance of 0.45 still dwarfs the 0.09 expected from sampling error. The authors state that the pattern "cannot be explained by sampling error, feedback sign, or existing theories." ↩
James W. Smither, Manuel London and Richard R. Reilly, "Does Performance Improve Following Multisource Feedback? A Theoretical Model, Meta-Analysis, and Review of Empirical Findings," Personnel Psychology 58, no. 1 (2005): 33–66. Full text. Corrected mean effect sizes across 24 longitudinal studies; self-ratings came in at −0.04. ↩
Brad J. Bushman and Roy F. Baumeister, "Threatened Egotism, Narcissism, Self-Esteem, and Direct and Displaced Aggression: Does Self-Love or Self-Hate Lead to Violence?," Journal of Personality and Social Psychology 75, no. 1 (1998): 219–229. Full text. Study 1, N = 260: the insult raised aggression with d = 0.25, and narcissism predicted aggression more strongly after an insult (r = .37) than after praise (r = .18). A later meta-analysis of 437 studies and 123,043 participants puts the narcissism-aggression link at r = .26 overall, more than twice as strong under provocation: DOI 10.1037/bul0000323. ↩
Tracy Kwang and William B. Swann Jr., "Do People Embrace Praise Even When They Feel Unworthy? A Review of Critical Tests of Self-Enhancement Versus Self-Verification," Personality and Social Psychology Review 14, no. 3 (2010): 263–280. Full text. Cognitive processing favoured self-verification (r = .30 against .18), as did the actual choice of evaluator and of feedback; affect favoured self-enhancement (r = .29 against .13). ↩
Leslie K. John, Hayley Blunden and Heidi Liu, "Shooting the Messenger," Journal of Experimental Psychology: General 148, no. 4 (2019): 644–666. Full text; DOI 10.1037/xge0000586. ↩
William Hart, Dolores Albarracín, Alice H. Eagly, Inge Brechan, Matthew J. Lindberg and Lisa Merrill, "Feeling Validated Versus Being Correct: A Meta-Analysis of Selective Exposure to Information," Psychological Bulletin 135, no. 4 (2009): 555–588. Full text. Overall d = 0.36, which the authors convert to an odds ratio of 1.92; d = 0.42 among the highly committed against 0.01 among the uncommitted. ↩
Russell Golman, David Hagmann and George Loewenstein, "Information Avoidance," Journal of Economic Literature 55, no. 1 (2017): 96–135. Full text. People decline information that is free, useful and relevant to a decision they are about to make, and in the medical cases will pay to avoid it. ↩
Victoria F. Sisk, Alexander P. Burgoyne, Jingze Sun, Jennifer L. Butler and Brooke N. Macnamara, "To What Extent and Under Which Circumstances Are Growth Mind-Sets Important to Academic Achievement? Two Meta-Analyses," Psychological Science 29, no. 4 (2018): 549–571 (full text); David S. Yeager et al., "A National Experiment Reveals Where a Growth Mindset Improves Achievement," Nature 573 (2019): 364–369 (PMC6786290), 12,490 ninth-graders in 65 schools. ↩
Tracy Epton, Peter R. Harris, Rachel Kane, Guido M. van Koningsbruggen and Paschal Sheeran, "The Impact of Self-Affirmation on Health-Behavior Change: A Meta-Analysis," Health Psychology 34, no. 3 (2015): 187–196, 34 tests covering 3,433 people, d = 0.17 with a confidence interval of 0.03 to 0.31 (PubMed). The failed replication is Paul Hanselman, Christopher S. Rozek, Jeffrey Grigg and Geoffrey D. Borman, "New Evidence on Self-Affirmation Effects and Theorized Sources of Heterogeneity From Large-Scale Replications," Journal of Educational Psychology 109, no. 3 (2017): 405–424 (PMC5403146). What it fails to reproduce is an earlier district-wide field experiment of the same kind, not the laboratory paradigm of Geoffrey L. Cohen, Joshua Aronson and Claude M. Steele, "When Beliefs Yield to Evidence: Reducing Biased Evaluation by Affirming the Self," Personality and Social Psychology Bulletin 26, no. 9 (2000): 1151–1164, where affirmed participants were more persuaded by evidence against their own position on capital punishment. ↩
Sally S. Dickerson and Margaret E. Kemeny, "Acute Stressors and Cortisol Responses: A Theoretical Integration and Synthesis of Laboratory Research," Psychological Bulletin 130, no. 3 (2004): 355–391. Full text; PubMed. Across 208 laboratory studies, tasks combining uncontrollability with social-evaluative threat produced d = 0.92, against no significant cortisol change at all for passive tasks and for performance tasks with neither ingredient. ↩
David Scott Yeager et al., "Breaking the Cycle of Mistrust: Wise Interventions to Provide Critical Feedback Across the Racial Divide," Journal of Experimental Psychology: General 143, no. 2 (2014): 804–824, at 809. Full text. 44 seventh-graders, 22 Black and 22 White, all mid-range students; the 71 and 17 per cent figures are covariate-adjusted revision rates among the Black students only, the White students' rates being 87 and 62 per cent and not significantly different, raw rates 64 and 27 per cent, p = .045. Teachers were blind to which note each student received. ↩
Elizabeth J. McClean, Sean R. Martin, Kyle J. Emich and Todd Woodruff, "The Social Consequences of Voice: An Examination of Voice Type and Gender on Status and Subsequent Leader Emergence," Academy of Management Journal 61, no. 5 (2018): 1869–1891. DOI 10.5465/amj.2016.0148. ↩
Behavioural Insights Team, Improving Performance Management Feedback: A Randomised Controlled Trial (2021). Report. The pre-registered trial covered 46,176 feedback responses written about 4,328 senior managers in a large UK public-sector body. On the earlier and more widely quoted finding, see Shelley J. Correll and Caroline Simard, "Research: Vague Feedback Is Holding Women Back," Harvard Business Review, 29 April 2016, an industry research brief rather than a peer-reviewed study; the peer-reviewed work from the same group is Correll, Weisshaar, Wynn and Wehner, "Inside the Black Box of Organizational Life: The Gendered Language of Performance Assessment," American Sociological Review 85, no. 6 (2020): 1022–1050. ↩
Enjoyed this? Get the next one.
New essays, in your inbox. Double opt-in, unsubscribe anytime — no tracking pixels.