Showing posts with label causation. Show all posts
Showing posts with label causation. Show all posts

Tuesday, September 19, 2023

Percentages: A desire to claim authority

 

     Stereotypes often have a grain of truth within. How large of a grain? It varies a lot. Any time that you find yourself saying -- Xs do this or Ys do that -- you know that there are those who don't fit that criterion. Many males do not like to ask for directions (in my household, my wife is the one who doesn't like to ask). Women, in general, have faster reflexes and higher pain thresholds. Much of the time, such statements are stated as absolutes. "Men don't ask for directions." This is true even when we are quite aware of exceptions -- with either men asking for directions or other genders not wanting to ask.

     But yet another generalization which doesn't fit everyone -- people (especially men) want to be perceived as being an authority -- knowing what they are talking about. "70% of all men have dandruff." "93% of all drivers don't come to a complete stop at stop signs." People could easily say "most drivers don't come to a complete stop" but which sounds more as if the person has done their research and knows what they are talking about -- "most drivers ..." or "93% of all drivers ..."?

     How do such percentages come about? Occasionally, the person may have read a paper stating some percentage. That percentage may, or may not, have been accurate and the person may, or may not, have remembered the percentage correctly. But there WAS some percentage associated with the event so recreate it.

    "Most becomes 70%." "Almost all becomes 94%." "Almost no one become 1%." You may have some favorite numbers that you use.

     This doesn't happen just in everyday life. It also happens in studies. Each study has a statistical range for all numbers and there is a general recognition that there may be constraints that have been omitted from the study or aspects of the population pool might make a difference with other followup studies. Yet, the results don't get summarized as "35% chance with a +/- error range of 5% under specific conditions."

     Finally, there is sloppy usage of well-defined words. The word "cause" is an oft-misused one. Somehow, in the summaries, "correlates to" or "seems to usually be present" becomes "causes". If A causes B then EVERY time A exists then B happens. There aren't a lot of absolutes in life. Even doses of arsenic will not necessarily cause death (though it certainly can raise the likelihood -- botox or plutonium have even smaller likelihoods of not being fatal). Each time I read of a study that indicates causation, I say to myself "uh-huh. How about under this condition? How about this 'exception' that I read about? ..."

     Does this apply to you? Will at least 82.5% of you read this and think about how it may apply within your life?

Saturday, February 13, 2021

Causation and Correlation: confusion or simplification

 

     We often encounter statements in the media of the nature "X causes Y". "G is the cause of Q". "You need to stop H or Q will happen."

     Of course, the rejoinder often is "She did H and Q never happened." or "I have Xed all of my life and Y has never entered into the picture."

     This isn't to say that X, G, or H are not things that are detrimental to health, or the economy, or the environment, or whatever other category it may be applied to. It is just a matter that, as is true of many other situations applied to health, politics, religion, and so forth the use of the word "cause" is not accurate.

     Simplification is not, in itself, a bad thing. Simplification can help in describing something for a wide audience with widely different backgrounds and experiences. Simplification should be recognized as such and not be considered "the whole truth". As mentioned in a prior blog, the Body Mass Index (BMI) number is a good analysis of healthy weight for about 80% of the population. And, as long as the 20% is not penalized via use of the BMI, it is not a bad thing. Unfortunately, easily obtained numbers often get misused -- on purpose or out of lack of effort.

     As mentioned in the previous blog, if you take two large, truly random, groups of people then you can statistically compare the groups when one group has something done and the "control group" has not. This group had H. The control group did not have H. Five percent more of the group with H developed Q compared to the group. Is five percent a significant difference? It depends on group sizes and confidence in the comparability of the two groups. Beyond that, the actual number where it becomes "statistically significant" should be done by a statistician (which I am not).

     At this point, perhaps it is determined that it IS statistically significant. The media, in their desire to inform but strong tendency to oversimplify, will shout "H causes Q!" But what about all those other people in either group that did not develop Q?

     If I hit an unshielded thumb, placed on a concrete slab with no protections (note the careful listing of conditions) and hit it hard with a hammer, it will cause pain and it will cause an injury (of varying seriousness depending on additional factors). In this situation, since it will happen to anybody who is in this scenario, we can say "hitting an unprotected thumb on a concrete slab with a hammer causes pain and injury". No exceptions, if you do A it causes B.

     Note, there can always be something that is not specified that might make this not true. What if the thumb was part of a strong prosthetic hand? What if the concrete slab had not set yet and the hand and thumb went into the concrete when the hammer hit? But in the everyday world, you can say that hitting the thumb with a hammer causes pain and injury. You can say that the earth's rotation causes the appearance of the sun coming up over the horizon to the east. In each case, causing is (for all practical purposes) 100% true for the specified situation.

     In these situations it is also very important to be specific. For example, there is the statement that "burning fossil fuels causes an increase in carbon dioxide and other waste products in the atmosphere". Seems obviously true, doesn't it? But, what if the factory had a carbon dioxide "scrubber" on the stack and a reclamation chamber prior to release of particulates to the atmosphere? So, a much more accurate (and something that can be addressed in more ways) statement is "burning fossil fuels without adequate filtering and reclamation causes an increase in carbon dioxide and other waste products in the atmosphere". It isn't the burning of the fossil fuels that is the foundational problem -- it is the non-cleansed release of the  exhaust products. Possible solutions increase. We can change fuels or we can install filtration and reclamation remedies.

     In the situation where the group with H had Q develop 5% more often than the group without H, it is NOT true that "H causes Q" no matter how much the media, or others, want to simplify the situation. What we have here is a correlation. A correlation can be insignificant (below the threshold of what the statisticians consider to be significant), significant (at threshold), moderately significant, or strongly significant between H and Q developing. But it cannot accurately be said that "H causes Q".

     And, the inaccurate use of cause in saying "H causes Q" reduces credibility of the message.

Saturday, September 27, 2014

Why aren't studies steady? The problems with studies on humans.


     If you don't like the results of a study ... wait for the next study. It appears that the results from studies involving people vary drastically from study to study ... and they do. At one time butter is bad and margarine is good and then, later, margarine is bad and butter is better (not necessarily good). Fats are bad. No, the right fats are good. Olive oils are the right fats. No, more polyunsaturated fats are even better. Why do studies that involve humans vary in their results so much?

     There are quite a few reasons why it is difficult to have consistent results from studies of humans. Some are inherent problems. Some are political problems. And a large number of problems arise from the way studies are reported in the media rather the actual study. In other words, a study may be done very well and present results that are interesting but not conclusive -- but some specific parts are taken up by the media as "startling results". What are the problems with studies?

  • People are not mice. Many studies on health effects are done with mice, or guinea pigs, or monkeys, or chimpanzees, or some other more easily studied animal. In addition, many studies may make use of one gender but generalize to both genders. It is rather obvious that, for the best results, use of humans of the appropriate category must be used. Why isn't this the case?

    Cost. People want to be paid (in money or value) to participate in studies. Animals can be purchased -- and, until some group starts recognizing what is being done to them, can be treated with the least care needed.

    Morality. Except in certain situations (such as Nazi Germany where psychopaths had full permission to experiment) it is not acceptable to put people's lives at risk. This is associated with control groups where they are NOT treated the same as the group which is being tested as well as with the groups being treated with undetermined results.

  • Patience. It takes TIME to determine what long-term results are. Of course, it doesn't take much time for an immediately lethal poison to be known but most substances aren't as immediate in effect. Taking longer periods of time means results are delayed. It also increases the costs of the study.

    Two types of studies of this nature are longitudinal and cross-generational. One examines individuals for substantial portions of their lives and the other examines the effects from parent to child to grandchild. Humans have pretty long lives (unless stopped by disease, accident, or violence) and this also leads to use of shorter lived animals as subjects.

  • Control groups. A control group is simple in concept but much harder to create and use. The idea is that one group has the variation (tries a medicine, eats a food, does an exercise, endures environmental conditions, ...) and the other does not. Comparing the results of the two groups is hoped to be able to isolate the effects of the specific effect being studied.

    A huge difficulty is that it is impossible to prevent the confusion of combinations of variables. You are testing item A. It turns out that A does one thing in the presence of items B and C. It does another, different, thing in the presence of items C and D -- and it may even do something else in the presence of B and D. If B, C, and D are all known and defined then useful results may still be obtained -- but often they are not. These combinatorial variables are unknown -- but they can drastically affect the results. Two of these variables involve environment and genetics.

    Environmental variables. What is the effect of electricity in a house or city? In modern society, it is impossible to eliminate -- and, if taken to a part of the world where electricity is NOT present, other variables will exist. What is the effect of plastics? What are the effects of pesticides, hormones, or antibiotics in the food? While these can be minimized, they cannot be eliminated. What about specific pollutants in the air? And so forth.

    Genetics. Leading up to the next bullet item on correlation versus causation, I once read a statistical study on smoking and cancer rates in various countries of the world. It turned out that countries with the greatest amount of smoking had among the lowest amounts of cancer. The effect of smoking depended upon the population. (It didn't get a lot of media attention since the outcome was not politically popular.) A homogenous (identical in nature) population is needed for studies and sets of identical octuplets are hard to find.

  • Correlation versus causation. This is understood by scientists doing studies but easily distorted by politicians, business owners, groups, or media people who have a bias towards a particular result. Causation means the variable causes the result. If I hit my toe with a hammer it will be bruised (or broken). This is true if any toe hit by any hammer causes these results. The effects can be changed -- a hammer hitting a toe that is shielded by a steel-toed boot is NOT hurt (but it also means the toe of the foot is not actually hit).

    Some people are allergic to monosodium glutamate (MSG). Some are not. So, MSG AND allergic people cause a particular effect but MSG AND non-allergic affect does NOT cause the effect. This is verifiable contributory causation.

    Everyone who drinks water will eventually die. Does drinking water cause people to die? No (unless, of course, the water is contaminated), but this is the type of (often statistical) result that biased groups enjoy mis-interpreting and reporting.

  • Binary Results. People like simple results. They like "yes" or "no". They do not like long combinations of possibilities. So, a report that says "butter is bad" is much easier to distribute than a report that says "butter, in combination with lack of exercise and excessive refined carbohydrates and a genetic tendency towards high cholesterol, can contribute to high blood pressure".

    Properly done studies rarely have simple results.


So, in summary, it is difficult to create a useful, consistent study on the effects of anything with people. Even if done properly, it is difficult to give results without also giving all of the controlled variables along with the result.

Smoke Gets in Your Lungs (updated)

     This is an article that I published in here on February 22, 2013. I try to make my articles “timeless” as I try to work with “foundatio...