Pages

 
Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Stop abusing statistical significance

0 comments
I just made my first edit on Wikipedia, on the article on 'statistical power'. Here's the old text, with the deleted parts in bold:

There are times when the recommendations of power analysis regarding sample size will be inadequate. Power analysis is appropriate when the concern is with the correct acceptance or rejection of a null hypothesis. In many contexts, the issue is less about determining if there is or is not a difference but rather with getting a more refined estimate of the population effect size. For example, if we were expecting a population correlation between intelligence and job performance of around .50, a sample size of 20 will give us approximately 80% power (alpha = .05, two-tail). However, in doing this study we are probably more interested in knowing whether the correlation is .30 or .60 or .50. In this context we would need a much larger sample size in order to reduce the confidence interval of our estimate to a range that is acceptable for our purposes. These and other considerations often result in the true but somewhat simplistic recommendation that when it comes to sample size, "More is better!"

However, huge sample sizes can lead to statistical tests becoming so powerful that the null hypothesis is always rejected for real data. This is a problem in studies of differential item functioning.


Leaving the cost of collecting data aside, larger (appropriately collected) samples are ALWAYS BETTER. At the end of the day, if your sample is *too* large (for example if your statistical software restricts the amount of information you can load on it and you don't need the extra information anyways) you can always obtain a smaller random sample from your larger random sample. So, the 'more is better' recommendation is simple, but not simplistic.

The last paragraph reveals a fundamental misconception about statistical significance that refuses to go away. If the effect of an independent variable on the dependent variable is zero, using a very large sample will result to an estimated effect that is 0 to many decimal places; as the sample size increases further, the effect will approach *exactly* zero even more. NEVER USE STATISTICAL SIGNIFICANCE AS A PROXY FOR PRACTICAL SIGNIFICANCE. I have no clue whether large sample sizes have been seen as a problem in the past in studies of differential item functioning, but if that is the case then the researchers are idiots.

Here is another post on problematic applications of statistical significance.
Read more from this post in Review topics and articles of economics »

Stata is blogging

0 comments
Stata now has an official blog, Not Elsewhere Classified. And here's a list of a few unofficial ones.
Read more from this post in Review topics and articles of economics »

Race and brains, once more unto the breach

0 comments
If the previous post did not persuade you to stop wasting your time thinking about it, McMegan links to Jim Manzi who addresses the question. He concludes his well researched piece with this:

Do genetic differences accounts for any material portion of the difference in IQ scores by self-identified racial groups in the US? The only honest answer is that we don’t know. This, not political correctness is why the American Psychological Association’s formal consensus point of view on this question is stated without qualification: “At present, this question has no scientific answer.”

All right and proper, but that's not the right question to ask. What you really want to know is this: are 'black genes' leading to materially less intelligence than 'white genes'? And the answer is simple: IQ tests can't tell you that.

My understanding is that IQ scores say nothing about 'absolute' intelligence, they only provide a ranking. Or to use terminology more familiar to some of my readers, IQ scores only have an ordinal, not a cardinal meaning. It is very well likely that someone scoring 110 is only trivially more intelligent than someone scoring 90; what the difference between 110 and 90 actually means in terms of 'amount of intelligence' is anyone's guess.

OK, I hear you say, but don't we use quasi-cardinal interpretations for IQ scores? (e.g. isn't 'normal intelligence' supposed to lie between 90 and 109?) Quoting Manzi again:

There are statistically significant differences in IQ test performance between self-identified racial and ethnic groups in the US, and these differences have been sustained over long periods of time. The specific difference that is most widely discussed is the fact that in the US Non-Hispanic whites score, on average, about 15 points (~1 STDEV) higher than African-Americans. (Leaving aside the complication that it matters exactly how we define “long periods of time”, since, for example, there is circumstantial evidence that the black-white IQ gap may have been reduced substantially over the past several decades.)

So, 15 points is the maximum possible difference between the races. We also know with certainty that environment plays a role in determining intelligence, so the difference that can be attributed to genetics is a maximum of 10 points or so, with the actual difference (if it exists) likely to be much smaller. That's nothing: let me remind you that the average (and I think also median) person scores 100 by design, that 'average' or 'normal' intelligence is a 20 point band, and that in any case IQ is a flawed measure of 'intelligence' as used in everyday language (and a 'better than random' - but not by much - predictor of 'success' in life).

To bring the human element into this, my own results from several IQ tests are uniformly distributed across a 35 points range, and while I'm pretty good at arriving to answers to almost every IQ test type question thrown at me - questions that the average person won't answer at all - I take more time than the average person to do so (does that make me more or less intelligent?). And since the human element sells, I got more for you: here is a long list of highly successful and intelligent people who were very likely autistic, and here's the corresponding long list of people with dyslexia. What's the point? Intelligence is not a uni-dimensional variable.

And just to make sure, I should also mention the Flynn effect, quoting from this (excellent) paper (free access):

Since 1932 and probably prior to that, test scores have been increasing at a rate of 3 to 6 IQ points per decade, depending on the IQ test used. The preponderance of evidence indicates that scores are continuing to rise at a constant rate (Flynn, 2006b).

And since you apparently have to be a genius to read Bluematter., I won't even draw out the implications of the following observation on the likely (non)persistence of any currently observed differences amongst races (ala Manzi's circumstantial evidence):

There is at least one exception, however. The Scandinavian countries currently are showing little or no rise in their test scores (Flynn, 2006a). As large IQ increases were seen in Norway prior to 1968, Flynn suggests that Scandinavia might have experienced early increases that have since abated. This raises the possibility that IQ increases in other industrialized nations will also end.

So in IQ we have a metric with no cardinal interpretation, with a weak correlation to 'general intelligence' and an even weaker one to 'success in life', with the observed differences between races being pretty small and most likely diminishing even before controlling for environmental characteristics.

What's the issue again?
Read more from this post in Review topics and articles of economics »

I hate seasonally adjusted series

0 comments
To all data providers, wherever they may be:

Please, please stop seasonally adjusting the series you give me. I can run a regression with seasonal dummies myself, thank you very much. So can everyone else. The difficult thing is to get from the seasonally adjusted series to the original one, and I don't see why you should deprive me of the privilege.

Yes, I know I shouldn't want to in most cases, but let me be the judge of that, OK? What's your problem anyway, what is it to you? Please, just do me this favour.
Read more from this post in Review topics and articles of economics »

Assume blacks have lower IQ than whites and DNA is to blame

0 comments
Now name one thing you would do differently - as a politician, as a citizen or as a human being. I'll be damned if you can come up with a single example.

I really, really can't understand what this debate is all about (other than in an immature 'I did not evolve from the apes' or 'I am really frustrated the earth is not at the centre of the solar system' kind of a way). Hell, even this debate is more relevant.

Now can we please, as a culture, move on?
Read more from this post in Review topics and articles of economics »

Disaggregating annual variables

0 comments
A reader emails me:

When doing econometrics on quarterly time series data, if there are some key variables that are available only annually, is there merit in interpolating the annual data to create a quarterly series or should the variables be discarded? What is the general advice on interpolation?


As a general rule, you should not discard the annual data. As with any econometric problem, the question is: do the additional data contain potentially useful information? If the answer is yes, then the next step is to find the best way to disaggregate the annual observations into quarterly ones.

There are many ways to do this, and the most appropriate one will depend on the problem at hand. Also, keep in mind that determining the appropriate standard errors for your included variables, especially the disaggregated ones, can be a bit tricky in this setting.

Here are some relevant papers (only the first free access)

Also keep in mind that, depending on the problem you are facing, it may even make sense to aggregate variables – e.g. making annual variables out of quarterly ones. Yes, you shed information in that case, but an even more critical question to ask is whether your assumptions are satisfied (usually E(u|X)=0). In many cases, you would have reason to expect the error to be correlated with your dependent variables in a ‘quarterly’ model but not in an ‘annual’ model, in which case it would be most probably preferable to use the latter.
Read more from this post in Review topics and articles of economics »

What's in a name?

0 comments
Andrew Gelman quotes this paper (free access) by Leif Nelson and Joseph Simmons:

In five studies, we found that people like their names enough to unconsciously pursue consciously avoided outcomes that resemble their names. Baseball players avoid strikeouts, but players whose names begin with the strikeout-signifying letter K strike out more than others (Study 1). All students want As, but students whose names begin with letters associated with poorer performance (C and D) achieve lower grade point averages (GPAs) than do students whose names begin with A and B (Study 2), especially if they like their initials (Study 3). Because lower GPAs lead to lesser graduate schools, students whose names begin with the letters C and D attend lower-ranked law schools than students whose names begin with A and B (Study 4). Finally, in an experimental study, we manipulated congruence between participants’ initials and the labels of prizes and found that participants solve fewer anagrams when a consolation prize shares their first initial than when it does not (Study 5). These findings provide striking evidence that unconsciously desiring negative name-resembling performance outcomes can insidiously undermine the more conscious pursuit of positive outcomes.

The explanation? (Keep in mind this is a paper published in Psychological Science)

People like their names and initials (Nuttin, 1987). In fact, this name-letter effect (NLE) is influential enough to encourage the pursuit of name-resembling life outcomes and partners. [...]

Do people consciously or unconsciously pursue name-resembling outcomes? Do a few people named Jack deliberately move to Jacksonville for its Jack-resembling appeal, or are they driven by an unconscious desire? Researchers have certainly argued that the latter is true. The NLE is described as an indicator of implicit egotism (e.g., Koole, Dijksterhuis, & van Knippenberg, 2001; Jones et al., 2004; Pelham, Carvallo, & Jones, 2005; Pelham et al., 2002; Sherman & Kim, 2005), as own-name liking is thought to indicate unconscious self-liking.

I'm not convinced. To refer back to one of the quoted studies, how about students whose surnames start with 'F'? Shouldn't they be performing much worse than the C's and D's?

Two potential explanations here:

1. Omitted variable bias. For example, say that names that start with C or D are way less frequent in the population of Asian students compared to Anglo-Saxon surnames. Further, assume that Asians are discriminated against when it comes to college admission, perhaps due to uncertainty about the quality of the schools they attend. That way, the average Asian in college will be a better student than the average Anglo-Saxon, and he will also be less likely to have a name that starts with C or D.

2. There are an infinite number of hypotheses, and a finite but very large number of original datasets. In other words, datamining - or if we want to be somewhat less harsh on the researcher, pure luck.

And talking of names, here's Levitt and Dubner approaching the issue from a completely different angle.
Read more from this post in Review topics and articles of economics »

What are the odds?

0 comments
You are on holiday in some strange land, and you bump into Sue, an old friend. What are the odds?

Well, the probability is 1. 100%. It's certain. It bloody happened.

OK, you say, fair point - but that's not what you meant. What you meant is what is the probability you would bump into Sue, assuming you hadn't just bumped into him. I reply that that's a silly assumption to make as you just did bump into him, but you insist.

Well then, the probability is whatever you want it to be - pick a number, and I'll explain to you why it's plausible. You look perplexed and ask what I mean.

I explain: first of all, you need to specify the probability of what you are interested in.
-Do you want to know the probability you'd bump into Sue at the time and place you did
-Do you want the probability you would bump into an old friend while on holiday in general
-Or do you want the probability that something 'remarkable' enough would happen to you at some point in your life that would make you start asking silly questions about 'what is the probability of that happening'?

You say the first, obviously, and that I should stop being clever. I wasn't done, I say, and proceed to ask what is the information set I should base my probability estimate at - quantum mechanics aside, randomness is in the eye of the beholder after all, and if I was all-knowing God the probability of whatever it is that happened would be 1 even before it happened.

You throw your pina colada on my head and vow never to speak to me again.

A few days later, you are kind of missing me but don't feel like talking to me yet, so you visit bluematter. as a first step in rebuilding the relationship. And the first thing you see is this delightful little story, via Andrew Gelman:

In the city of Syracuse, the strangest thing happened in Tuesday's Democratic presidential primary.

Sen. Hillary Clinton and Sen. Barack Obama received the exact same number of votes, according to unofficial Board of Election results.

Clinton: 6,001.

Obama: 6,001.

The odds of Clinton and Obama tying were less than one in 1 million, said Syracuse University mathematics Professor Hyune-Ju Kim.

Elaborating on Thursday, she [Professor Hyune-Ju Kim] noted: "The "almost impossible" odd is obtained when we assume the Syracuse voter distribution follows the New York state distribution. Since it is almost impossible to observe what we have observed, statistically we can conclude that Syracuse voter distribution is significantly different from the New York state distribution."

There would be less than one in 1 million chance of a tie occurring between Clinton and Obama in voting by a randomly selected group of 12,346 New York Democratic voters, she said.


To which Andrew replies:

Not to pick on some harried mathematics professor who'd probably rather be out proving theorems, but . . . of course Syracuse voters are not a randomly selected group of New Yorkers. You don't need a statistical test to see that. Regarding the probability of an exact tie: I don't think that's so low: a quick calculation might say that either Clinton or Obama could've received between, say, 5000 and 7000 votes, giving something like a 1/2000 chance of an exact tie. That's gotta be the right order of magnitude.

If there was one thing you were ever certain about, it is that you don't want to read what I have to say on this. A baseball bat happens to lie next to you (what are the odds!?). You grab it with both your shaky hands and smash the computer monitor to pieces.
Read more from this post in Review topics and articles of economics »

Roses are red, violets are blue, and correlation is not causation

0 comments
Merv emails me this article:

Football clubs with red team strips are more successful than those with other colours, according to a study released Wednesday.

The fact that English clubs Manchester United, Liverpool and Arsenal regularly top league tables is not a coincidence, say the experts from Durham University and the University of Plymouth.

Red shirts give the team an advantage due to deep-rooted biological response to the colour. "In nature, red is often associated with male aggression and display," they said, giving the example of the red-breasted robin.

"It is a testosterone-driven signal of male quality, and its striking effect has even been harnessed by soldiers in the past," added the researchers, after analyzing data on English football league results since World War II.


The red-breasted robin? That's the most fearsome red beast they can think of?

I can't access the paper, but just reading the abstract is enough to convince me it's bonkers. The authors also have a 2005 paper - in Nature no less - entitled 'Red enhances human performance in contests'. If any reader has a access to Nature, would you be kind enough to email me the article so I can -ahem- review it?
Read more from this post in Review topics and articles of economics »

Dodging the Vietnam draft

0 comments
What percentage of draft eligible men did not, or would not, join the US military despite being drafted? How many men's enlistment in the military ultimately depended on the outcome of the lottery? I will let you ponder these questions for a moment, and put the answer under the fold...


The numbers come from Angrist's 1990 AER paper (published paper, working paper - both free access), with the table being adapted from the Imbens/Wooldridge lecture 5.




Whether someone was drafted was randomly determined (Angrist has more details on the Vietnam lottery). Now, to make sense of the table above:

A 'never-taker' is someone who does not serve in the military, regardless of the draft lottery outcome. A startling (at least to me) 69% falls in this category - in other words, 7 out of 10 of men in the draft pool did not/ would not have served regardless of whether they were drafted or not.

An 'always-taker' is the exact opposite: he would serve if drafted, and volunteer to serve if not drafted. (note here is that 'always-takers' may not have volunteered in the absence of the draft - many young men at high risk of being drafted may have volunteered to take advantage of the better terms offered to volunteers compared to draftees). The proportion of always-takers was 19%.

A 'complier' is someone who would serve if drafted, but not volunteer to serve otherwise. A mere 12% falls in this category, meaning that only 1 in 8 men in the draft pool served (or not) depending on the lottery outcome.

[For completeness, a defier is defined as someone who chooses not to serve if selected, while he actually volunteers to serve if not selected. Defiers are simply assumed away, and reasonably enough too]

Wikipedia has some more background information:

The large cohort of Baby Boomers who became eligible for military service during the Vietnam War also meant a steep increase in the number of exemptions and deferments, especially for college and graduate students. This was the source of considerable resentment among poor and working class young men, who could not afford a college education.

As U.S. troop strength in Vietnam increased, more young men were drafted for service there, and many of those still at home sought means of avoiding the draft. For those seeking a relatively safer alternative to the Army, Marine Corps, Navy, or the Air Force, the Coast Guard was an option (provided one could meet the more stringent enlistment standards). Since only a handful of National Guard and Reserve units were sent to Vietnam, enlistment in the Guard or the Reserves became a favored means of draft avoidance. Vocations to the ministry and the rabbinate soared, because divinity students were exempt from the draft. Doctors and draft board members found themselves being pressured by relatives or family friends to exempt potential draftees.

According to the Veteran's Administration, 9.2 million men served in the military between 1964 and 1975. Nearly 3.5 million men served in the Vietnam theater of operations. From a pool of approximately 27 million, the draft raised 2,215,000 men for military service during the Vietnam era. It has also been credited with "encouraging" many of the 8.7 million "volunteers" to join rather than risk being drafted.

Of the nearly 16 million men not engaged in active military service, 96% were exempted (typically because of jobs including other military service), deferred (usually for educational reasons), or disqualified (usually for physical and mental deficiencies but also for criminal records to include draft violations). Draft offenders in the last category numbered nearly 500,000 but less than 10,000 were convicted or imprisoned for draft violations. Finally, as many as 100,000 draft eligible males fled the country.

And here's a couple of interesting paragraphs from Angrist:

Selection of individuals for induction for the draft-eligible, non-deferred 'high priority pool' was based on a number of criteria, the most important of which were the pre-induction physical examination and the examination of mental aptitude. In 1970, for example, half of all registrants failed pre-induction examinations and 20 percent of those remaining were eliminated by physical inspections conducted at the time of induction.

[Also], it is sometimes argued that during the Vietnam era students went to college to avoid the draft and that educational standards were reduced so as to avoid having to flunk students out of school. Baskir and Strauss (1978) claim that Vietnam era college enrollment was 6-7 percent higher because of the draft.
On a vaguely related personal note, this blogger was drafted and served nine months in the military, although that was in Greece during peace-time. I wrote something about the experience here and here, though both posts are mainly about other topics. During that time, I wrote a book consisting of 148 μαντινάδες, exclusively during guard duty (I did innumerable 4-hour stints of guard duty- in the middle of the night, in freezing temperatures, alone, with a loaded machine gun). Any reader who would like to read it, just drop me a line and I'll email you the professionally put-together pdf. Reader beware: it's in Greek.
Read more from this post in Review topics and articles of economics »

Are double-blind trials underestimating drug effectiveness?

0 comments
In a double-blind trial, patients exhibit the placebo (and nocebo) effects because of the expectation they might be on a real drug.

If expectations are so important, could it be that patients being administered the real drug don't react to it fully due to the expectation they might be on the placebo?

Addendum: Steve Waldman (of Interfluidity and Naked Capitalism) posts in the comments:
@llimllib on twitter - Bill Mill - posted a cite (in response to a tweet on this post) that seems like a nice confirmation of your conjecture. Almost perfect.
Wow.
Read more from this post in Review topics and articles of economics »

Google's broken hiring process

0 comments
From Google's director of research:
One of the interesting things we've found, when trying to predict how well somebody we've hired is going to perform when we evaluate them a year or two later, is one of the best indicators of success within the company was getting the worst possible score on one of your interviews. We rank people from one to four, and if you got a one on one of your interviews, that was a really good indicator of success.

Ryan Tate uses this as evidence that Google's interview process is broken.

Craig Newmark sets the record straight.

On related news, I was surprised by suggestions that Larry Page still reviews CVs personally, and half-surprised to find out that lots of Google employees struggle with bureaucracy within the company.
Read more from this post in Review topics and articles of economics »

Does being beautiful mean being average?

0 comments
Yes, according to this research (via Andrew Gelman):

The debate over the definition of beauty has been waged by both scientists and philosophers for centuries. We tested the idea that a facial configuration close to the population mean is fundamental to attractiveness.

First, we digitized images of faces of male and female college students (i.e., transformed the facial images into little dots of lightness and darkness called "pixels"). Each face is represented by a matrix of pixel values that can be mathematically averaged with the matrices of other faces. Once digitized and averaged together, we can turn the averaged pixel values back into images and have the composite faces rated for attractiveness.

College students rated the male and female composite faces as significantly higher in attractiveness than the individual faces used to create them, if the composites had at least 16 different faces in them. Thus, averaged faces are attractive. Note that when we use the word, "average," we mean the arithmetical mean, and not an average-looking person. If, for example, you take a female composite (averaged) face made of 32 different faces and overlay it on the face of an extremely attractive female model, the two images line up almost perfectly indicating that the model's facial configuration is very similar to the composites' facial configuration.

[...] we view averageness as fundamental and necessary to facial attractiveness. Averageness is not the only component of attractiveness, but without it, no face will be attractive.


Here are some selected publications on the matter.

My two cents: what makes you ugly are extreme characteristics (e.g. big nose or ears); averaging simply takes care of these 'large errors'. The same principle is behind the frequently superior performance of composite forecasts (e.g. of economic variables), where the arithmetic mean of a number of forecasts is often more accurate than any of the individual components.

Postscript: Note that averaging doesn't quite work with hair.
Read more from this post in Review topics and articles of economics »

The normal distribution

0 comments
This beautiful image comes courtesy of W. J. Youden.

I've been reading Edward Tufte's superb The Visual Display of Quantitative Information - perhaps the most perfect book I've ever come across. Almost every page is a revelation, so expect me to be posting more of the wonderful graphs Tufte has collected in the future.
Read more from this post in Review topics and articles of economics »

Man, good effort but then you mess it up

0 comments
Hai hai (as my Urdu-speaking partner would say), Derek Lowe starts well but messes up his conclusion (via Megan McArdle):

The news of a possible diagnostic test for Alzheimer’s disease is very interesting [...]

But let’s run some numbers. The test was 91% accurate when run on stored blood samples of people who were later checked for development of Alzheimer’s, which compared to the existing techniques is pretty good. Is it good enough for a diagnostic test, though? We’ll concentrate on the younger elderly, who would be most in the market for this test.The NIH estimates that about 5% of people from 65 to 74 have AD. According to the Census Bureau (pdf), we had 17.3 million people between those ages in 2000, and that’s expected to grow to almost 38 million in 2030. Let’s call it 20 million as a nice round number.

What if all 20 million had been tested with this new method? We’ll break that down into the two groups – the 1 million who are really going to get the disease and the 19 million who aren’t. When that latter group gets their results back, 17,290,000 people are going to be told, correctly, that they don’t seem to be on track to get Alzheimer’s. Unfortunately, because of that 91% accuracy rate, 1,710,000 people are going to be told, incorrectly, that they are. You can guess what this will do for their peace of mind. Note, also, that almost twice as many people have just been wrongly told that they’re getting Alzheimer’s than the total number of people who really will.

Meanwhile, the million people who really are in trouble are opening their envelopes, and 910,000 of them are getting the bad news. But 90,000 of them are being told, incorrectly, that they’re in good shape, and are in for a cruel time of it in the coming years.

The people who got the hard news are likely to want to know if that’s real or not, and many of them will take the test again just to be sure. But that’s not going to help; in fact, it’ll confuse things even more. If that whole cohort of 1.7 million people who were wrongly diagnosed as being at risk get re-tested, about 1.556 million of them will get a clean test this time. Now they have a dilemma – they’ve got one up and one down, and which one do you believe? Meanwhile, nearly 154,000 of them will get a second wrong diagnosis, and will be more sure than ever that they’re on the list for Alzheimer’s.

Meanwhile, if that list of 910,000 people who were correctly diagnosed as being at risk get re-tested, 828 thousand of them will hear the bad news again and will (correctly) assume that they’re in trouble. But we’ve just added to the mixed-diagnosis crowd, because almost 82,000 people will be incorrectly given a clean result and won’t know what to believe.

I’ll assume that the people who got the clean test the first time will not be motivated to check again. So after two rounds of testing, we have 17.3 million people who’ve been correctly given a clean ticket, and 828,000 who’ve been correctly been given the red flag. But we also have 154,000 people who aren’t going to get the disease but have been told twice that they will, 90,000 people who are going to get it but have been told that they aren’t, and over 1.6 million people who have been through a blender and don’t know anything more than when they started.

Sad but true: 91% is just not good enough for a diagnostic test.

Yes, doctors need to be able to calculate the probability a patient has a given disease taking into account not only the accuracy of the test but also other available information (e.g., for random testing, prevalence of the disease amongst an age-group); and they need to communicate this information clearly to the patient. This misunderstanding is a real problem, and something that doctors and everyone else need to be educated about.

But to go from that to '91% is just not good enough' is a huge leap.

As long as there isn't a 100% accurate test, we can never be certain whether the disease is present or not; but the test does give a lot of relevant information and we can lower the probability of a false alarm as much as we like by administering the test again and again.

If a disease affects 1 in 20 people and the test is 90% accurate, a 'positive' result means you have a mere 32% probability you are actually ill. If you administer the test a second time and you get a second positive, this probability jumps to 81%, and this keeps rising with the number of positive results. For a negative test result, the news are even better: the first negative result translates to a 99.5% you are healthy, the second negative to a .999% that you are.

(18% of the people will get one positive and one negative, which simply means there is a 95% probability they are healthy - i.e. the same as before taking any tests. Instead of 'not knowing what to believe', as Lowe speculates, their doctors should just explain to them that they need more testing if they want to increase the accuracy of the standard, pre-test prediction (healthy) above 95%)

Pay attention now, here comes the correct conclusion: If you don't have any symptoms, a positive test result for most diseases doesn't mean much - in most cases, you are still more likely to be healthy than not.

Next time you take a test, ask your doctor to calculate the probability you are actually ill or healthy; and if you want more certainty, take the test again, and again, until you are content with the degree of certainty on offer. And thank all those nice researchers for them 90% accurate tests - at least if they are not painful.
Read more from this post in Review topics and articles of economics »

Planning an informed jump off a bridge

0 comments
Zubin Jelveh has the inside scoop on where to go to avoid the crowds:

Read more from this post in Review topics and articles of economics »

The importance of being clear

0 comments
A formatting fubar involving an Excel spreadsheet has left Barclays Capital with contracts involving collapsed investment bank Lehman Brothers than it never meant to acquire.

Working to a tight deadline, a junior law associate at Cleary Gottlieb Steen & Hamilton LLP converted an Excel file into a PDF format document. The doc was to be posted on a bankruptcy court's website before a midnight purchase offer deadline on 18 September, just four hours after Barclays sent the spreadsheet to the lawyers. The Excel file contained 1,000 rows of data and 24,000 cells.

Some of these details on various trading contracts were marked as hidden because they were not intended to form part of Barclays' proposed deal. However, this "hidden" distinction was ignored during the reformatting process so that Barclays ended up offering to take on an additional 179 contracts as part of its bankruptcy buyout deal, Finextra reports.


The Register has the full story. As Merv has always warned, 'horrible things happen when you hide cells in excel'.

I see this as a manifestation of a wider lack of education on the importance of communicating information efficiently. The Spartans, Tufte, Strunk and White, the Economist, Picasso and numerous econometricians have done a lot to improve things, but management-speak, TV advertising and other such phenomena show we still have a long way to go.
Read more from this post in Review topics and articles of economics »

Statistical significance

0 comments
...is not a measure of confidence in the point estimate; it says nothing about accuracy.

Statistical significance simply means that the true value of the statistic being estimated has a higher than 5% or 1% probability (the levels conventionally chosen) to be away from a range around zero (the range being determined by sample size and degrees of freedom of the estimator in question). Any statistic which is large enough will be found to be statistically significant even in small samples; this doesn't mean, however, that the accuracy of the point estimate can't be very poor.
Read more from this post in Review topics and articles of economics »

Funny graphs

0 comments
Dave sent those via email - I've seen some of them before, but they are still hilarious (note: there's more graphs under the fold).














 
Btw, this is the only kind of pie chart that doesn't make me angry.
Read more from this post in Review topics and articles of economics »

Econometric causality

0 comments
James Heckman has an excellent paper on the subject (free access).
Read more from this post in Review topics and articles of economics »