The HIDDEN Problem in Science
There’s a really major problem in science that I’m pretty sure you’ve never heard about. In fact, almost nobody knows about it. It didn’t even have a name until we gave it one. And I’m not talking about the replication crisis, which maybe you have heard of. So today, I’m going to break down for you what this major problem in science is, how it functions, and give some suggestions for how we might improve the situation. In empirical sciences such as
Psychology, which is my focus, there are three main strategies that people are aware of in order to get published in top journals. The first strategy is what everyone wants to do and also what we want everyone to do. It’s to make a novel or important scientific contribution. The journals publish you because you’ve made a genuine contribution to science. The second strategy is known as P hacking and this is largely the cause of the replication crisis. Packing is when you make choices
In the way you analyze or report your results such that you get a statistically significant result. But if someone were to redo the study from scratch on new data, they would not find the same result. Essentially, you’re leveraging noise in order to get a false positive that’s not a real effect. Now, Packaging has been a very major problem in some sciences. For instance, in psychology, if you were to look back 20 years, you’d find that maybe 40% of all
Papers were Packed. And if you were to redo the study, you would not find the original finding. The third strategy for publication is fraud. This is where a researcher makes up a result or even makes up the study itself. Thankfully, this is very rare. It’s so unethical that most researchers are not willing to do it and it would be career ending if it were caught. On the other hand, it’s much more common than we’d want it to
Be. You probably heard about some famous faked experiments. In a survey we ran of academic psychologists, we had them estimate the rate of fraud in studies in their field and they put it at about 6% on average. So, that covers the three main ways of publishing top journals that most people are aware of. make a valuable contribution, P hack, or commit fraud. But here’s where things get weird. We run a project called transparent replications. And the way it
Works is that when new papers come out in top psychology journals, we pick some of them at random and we redo a study from them to see if it really holds up. As we went to replicate paper after paper, we kept finding serious problems, but they don’t match any of the three descriptions I gave you. They’re not valuable contributions, they’re not paging, and they’re not fraud. So, what the heck is going on? To understand this, let’s look at a real example we
Encountered. The study investigated whether people’s views on where wealth comes from is linked to their views on social welfare policy. So the idea is if you have the rewarding view on where wealth comes from, you might support the incentivizing policies. That is the policies that incentivize hard work. If you have the rigged view on where wealth comes from, you might, on the other hand, support the redistributing policies because you think the wealth wasn’t gotten fairly so we should
Redistribute it. On the other hand, if you have the random view on where wealth comes from, you might support risk pooling policies that help reduce the risk for everybody. As the paper puts it, rewarding, rigged, and random beliefs uniquely predict rated importance of incentivizing, redistributing, and risk pooling goals for social welfare policy, respectively. So reading the study, you come away with the view that there are three different views on where wealth comes from and that each of these is linked to a
Specific goal for social welfare policy that people support. So once we picked this paper at random and we’d selected this particular study from that paper, we redid the study from scratch. We rebuilt the materials. We recruited new study participants and we reanalyzed it using the original statistical methods. Guess what we found? Well, we got the same result as the original paper. This was not a case of peaking and it was not a case of fraud. And yet, there was
Something deeply flawed about this paper. The thing is, the statistics were quite complicated. It’s hard to intuitively understand what they were doing. So, we asked ourselves, what’s the simplest way to analyze this result? Since there are three views on where wealth comes from and three policies, we can simply make a table of all of them. We’d expect that each of the views on where wealth comes from has a positive correlation with its corresponding policy, but not a meaningful correlation
With the other policies. Unfortunately, this is not what we find. The rewarding view is supposed to be linked to the incentivizing policy, but it actually has no correlation. However, it has a negative correlation to the other two policies. This explains why they did find that it was more associated to its corresponding policy than the others, but it doesn’t have the meaning that you think. It’s not actually associated with its own policy. Moreover, the random view on where wealth comes from is not
More associated with its corresponding policy, which is the risk pooling policy. In fact, it’s just as associated with the redistributing policy, which is not supposed to be associated with. The only view on where wealth comes from that has the correct pattern is the rigged view. It may seem like the original research team was lying, but they weren’t. What they did is a complex statistical test that was hard to interpret. It really did have the result that it had. it just didn’t mean what it
Seemed to mean. So, while this study was misleading and probably was misinterpreted by many that read it, it was published in a top journal anyway. Now, that’s not to say the entire paper was worthless. The way our method works is spot-checking. We pick one study from the paper to focus on and this was the study we analyzed. Let’s take a step back here. This paper was not pact and it was not fraud and yet it was heavily
Flawed. So, what was that flaw exactly? Well, it didn’t mean what it seemed to mean. In other words, it probably shouldn’t have been published as is, or at least should have been heavily caveed. This leads us to our fourth missing publication strategy. It literally didn’t have a name, so we gave it one. We call it importance hacking, which is an analog to P hacking, but they’re very different from each other. With Packing, you use fishy statistical
Methods, and you get a statistical significant result that isn’t really there. If you redo the study, you don’t get the same effect. With importance hacking, if you redo the study, you do get the same effect. the effect just doesn’t mean what it seems to mean. Importance hacking is when you obscure the meaning or value of a result to make it seem more valuable than it is in order to get it published. So, importance hacking, in order to work at
All, has to work on peer reviewers and editors. It misleads them into thinking that there’s more value in work than there really is. There are many different flavors of importance hacking. It could be overgeneralizing a result, making a result seem more beautiful than it is, making a result seem more useful than it is, and so on. Our bizarre conclusion running replication after replication of top psychology studies is that importance hacking was occurring more often than P hacking. Even though
Importance hacking didn’t even have a name and P hacking is widely known about. While I do want us to run many more studies to make sure our observation holds up, we do actually have another line of evidence suggesting that at least in psychology, importance hacking is a huge problem. We recently ran a study on academic psychologists asking them to break down all the studies in their field that were recently published in top journals into four mutually exclusive categories. One
Is valuable research, two is packing and other forms of false positives. Three is fraud and fourth is importance hacking which we provided them a definition of. They assigned 27% of the studies in their field to the importance hacking group which is slightly higher even than P hacking. We also asked them how severe a problem they think P hacking is and importance hacking is in their field right now. And they actually rated importance hacking as a more severe
Problem than they did P hacking. What other forms can importance hacking take? Well, sometimes it looks like a measurement not reflecting the real world. For instance, in one study they used an extremely simple apple picking game and assumed that exploration in the game had implications for exploration in real life. A second example is that things that are obvious can be made to seem non-obvious through clever use of language. Suppose that a new work conscientiousness scale has really good
Predictive performance at predicting how good people are at their jobs. Well, that might seem interesting, but if it turns out the work conscientiousness scale contains questions about how good people are at their job, then it depletes the value. A third way this can happen is that when an effect is so small that it has no practical significance, but because it’s technically statistically significant, they’re able to still get the paper published. This can be done by focusing
On the p value and not on the actual size of the effect. Since my focus is the field of psychology, that’s what I have the most evidence on. But I believe importance hacking is actually substantial problem in many different disciplines in science. Now, I want to be really clear here. Importance hacking is usually not lying. We can think about a spectrum. At one end, people are talking about the result in completely objective terms with no hype, no spin,
No overgeneralization. At the other end of the spectrum is complete lying about the result. Importance hacking is often somewhere in the middle. They’re talking about a real result, but they’re doing it in a way that ends up confusing people, making them think that it has more value than it really does. However, it’s actually quite common that researchers themselves are misled about their own result. So, they may have no knowledge that they’re even doing this. Now, we would all hope
That peer reviewers and editors would catch this. But importance hacking is specifically referring to the stuff of this type that peer reviewers and editors are not catching, which unfortunately seems to be quite a bit. It seems that actually quite a lot of this is not getting detected during the peer review process. Suppose that we’re right and important hacking is a really big problem in science. What do we do about this? Well, we’re just beginning to understand this problem, but I have
Four recommendations. First, peer reviewers need better training in detecting importance hacking. they can better understand what to look for to sus it out. They don’t actually have an incentive to let bad research through. So, if they detect it and they know it’s bad, they’re likely to actually stop it. Second, it’s important that the materials for a paper are readily available and the peer reviewers are encouraged to look at them. The paper itself is to some extent a marketing
Document about the research. It’s important that reviewers actually look at the research itself because sometimes that’s the only way to tell that importance hacking has happened. Third, papers should be required to provide the simplest valid analysis. This is the name that we give to the simplest possible valid way of analyzing the data. It’s fine if researchers also want to report a really complex statistical analysis. The problem is if they only report a complex analysis, it can
Actually hide importance act results and make it hard for peer reviewers to spot them. My fourth recommendation is to use study diagrams. This is an idea that we came up with to make it easier to understand what exactly a study did at a simple glance. Here’s a sample study diagram from one of the papers we replicated. And they’re there to explain very clearly what exactly a study did at a simple glance. It may be surprising, but studies can often obscure exactly
What the study design was, making it harder for reviewers to tell precisely what the study involved. We all want science to produce valuable results. P hacking and importance hacking are barriers to doing that. We think that P hacking has actually improved quite a bit in some disciplines. That’s a great sign. We think importance hacking is going to be the next frontier, something we’re going to have to tackle to produce the next level of value in science. It’s
Crazy that it didn’t even have a term until we gave it one. We need to work on ways to improve it to make science better. To go much deeper on this and to see which papers actually replicated and which didn’t, check out replications.thinking.org or see the link in the description below. If you enjoy learning about psychology, we’d really appreciate it if you’d subscribe. We deep dive on psychology topics and do a lot of our own original research.