Pages

 
Showing posts with label econometrics. Show all posts
Showing posts with label econometrics. Show all posts

Stop abusing statistical significance

0 comments
I just made my first edit on Wikipedia, on the article on 'statistical power'. Here's the old text, with the deleted parts in bold:

There are times when the recommendations of power analysis regarding sample size will be inadequate. Power analysis is appropriate when the concern is with the correct acceptance or rejection of a null hypothesis. In many contexts, the issue is less about determining if there is or is not a difference but rather with getting a more refined estimate of the population effect size. For example, if we were expecting a population correlation between intelligence and job performance of around .50, a sample size of 20 will give us approximately 80% power (alpha = .05, two-tail). However, in doing this study we are probably more interested in knowing whether the correlation is .30 or .60 or .50. In this context we would need a much larger sample size in order to reduce the confidence interval of our estimate to a range that is acceptable for our purposes. These and other considerations often result in the true but somewhat simplistic recommendation that when it comes to sample size, "More is better!"

However, huge sample sizes can lead to statistical tests becoming so powerful that the null hypothesis is always rejected for real data. This is a problem in studies of differential item functioning.


Leaving the cost of collecting data aside, larger (appropriately collected) samples are ALWAYS BETTER. At the end of the day, if your sample is *too* large (for example if your statistical software restricts the amount of information you can load on it and you don't need the extra information anyways) you can always obtain a smaller random sample from your larger random sample. So, the 'more is better' recommendation is simple, but not simplistic.

The last paragraph reveals a fundamental misconception about statistical significance that refuses to go away. If the effect of an independent variable on the dependent variable is zero, using a very large sample will result to an estimated effect that is 0 to many decimal places; as the sample size increases further, the effect will approach *exactly* zero even more. NEVER USE STATISTICAL SIGNIFICANCE AS A PROXY FOR PRACTICAL SIGNIFICANCE. I have no clue whether large sample sizes have been seen as a problem in the past in studies of differential item functioning, but if that is the case then the researchers are idiots.

Here is another post on problematic applications of statistical significance.
Read more from this post in Review topics and articles of economics »

Stata is blogging

0 comments
Stata now has an official blog, Not Elsewhere Classified. And here's a list of a few unofficial ones.
Read more from this post in Review topics and articles of economics »

I hate seasonally adjusted series

0 comments
To all data providers, wherever they may be:

Please, please stop seasonally adjusting the series you give me. I can run a regression with seasonal dummies myself, thank you very much. So can everyone else. The difficult thing is to get from the seasonally adjusted series to the original one, and I don't see why you should deprive me of the privilege.

Yes, I know I shouldn't want to in most cases, but let me be the judge of that, OK? What's your problem anyway, what is it to you? Please, just do me this favour.
Read more from this post in Review topics and articles of economics »

Disaggregating annual variables

0 comments
A reader emails me:

When doing econometrics on quarterly time series data, if there are some key variables that are available only annually, is there merit in interpolating the annual data to create a quarterly series or should the variables be discarded? What is the general advice on interpolation?


As a general rule, you should not discard the annual data. As with any econometric problem, the question is: do the additional data contain potentially useful information? If the answer is yes, then the next step is to find the best way to disaggregate the annual observations into quarterly ones.

There are many ways to do this, and the most appropriate one will depend on the problem at hand. Also, keep in mind that determining the appropriate standard errors for your included variables, especially the disaggregated ones, can be a bit tricky in this setting.

Here are some relevant papers (only the first free access)

Also keep in mind that, depending on the problem you are facing, it may even make sense to aggregate variables – e.g. making annual variables out of quarterly ones. Yes, you shed information in that case, but an even more critical question to ask is whether your assumptions are satisfied (usually E(u|X)=0). In many cases, you would have reason to expect the error to be correlated with your dependent variables in a ‘quarterly’ model but not in an ‘annual’ model, in which case it would be most probably preferable to use the latter.
Read more from this post in Review topics and articles of economics »

What's in a name?

0 comments
Andrew Gelman quotes this paper (free access) by Leif Nelson and Joseph Simmons:

In five studies, we found that people like their names enough to unconsciously pursue consciously avoided outcomes that resemble their names. Baseball players avoid strikeouts, but players whose names begin with the strikeout-signifying letter K strike out more than others (Study 1). All students want As, but students whose names begin with letters associated with poorer performance (C and D) achieve lower grade point averages (GPAs) than do students whose names begin with A and B (Study 2), especially if they like their initials (Study 3). Because lower GPAs lead to lesser graduate schools, students whose names begin with the letters C and D attend lower-ranked law schools than students whose names begin with A and B (Study 4). Finally, in an experimental study, we manipulated congruence between participants’ initials and the labels of prizes and found that participants solve fewer anagrams when a consolation prize shares their first initial than when it does not (Study 5). These findings provide striking evidence that unconsciously desiring negative name-resembling performance outcomes can insidiously undermine the more conscious pursuit of positive outcomes.

The explanation? (Keep in mind this is a paper published in Psychological Science)

People like their names and initials (Nuttin, 1987). In fact, this name-letter effect (NLE) is influential enough to encourage the pursuit of name-resembling life outcomes and partners. [...]

Do people consciously or unconsciously pursue name-resembling outcomes? Do a few people named Jack deliberately move to Jacksonville for its Jack-resembling appeal, or are they driven by an unconscious desire? Researchers have certainly argued that the latter is true. The NLE is described as an indicator of implicit egotism (e.g., Koole, Dijksterhuis, & van Knippenberg, 2001; Jones et al., 2004; Pelham, Carvallo, & Jones, 2005; Pelham et al., 2002; Sherman & Kim, 2005), as own-name liking is thought to indicate unconscious self-liking.

I'm not convinced. To refer back to one of the quoted studies, how about students whose surnames start with 'F'? Shouldn't they be performing much worse than the C's and D's?

Two potential explanations here:

1. Omitted variable bias. For example, say that names that start with C or D are way less frequent in the population of Asian students compared to Anglo-Saxon surnames. Further, assume that Asians are discriminated against when it comes to college admission, perhaps due to uncertainty about the quality of the schools they attend. That way, the average Asian in college will be a better student than the average Anglo-Saxon, and he will also be less likely to have a name that starts with C or D.

2. There are an infinite number of hypotheses, and a finite but very large number of original datasets. In other words, datamining - or if we want to be somewhat less harsh on the researcher, pure luck.

And talking of names, here's Levitt and Dubner approaching the issue from a completely different angle.
Read more from this post in Review topics and articles of economics »

Dodging the Vietnam draft

0 comments
What percentage of draft eligible men did not, or would not, join the US military despite being drafted? How many men's enlistment in the military ultimately depended on the outcome of the lottery? I will let you ponder these questions for a moment, and put the answer under the fold...


The numbers come from Angrist's 1990 AER paper (published paper, working paper - both free access), with the table being adapted from the Imbens/Wooldridge lecture 5.




Whether someone was drafted was randomly determined (Angrist has more details on the Vietnam lottery). Now, to make sense of the table above:

A 'never-taker' is someone who does not serve in the military, regardless of the draft lottery outcome. A startling (at least to me) 69% falls in this category - in other words, 7 out of 10 of men in the draft pool did not/ would not have served regardless of whether they were drafted or not.

An 'always-taker' is the exact opposite: he would serve if drafted, and volunteer to serve if not drafted. (note here is that 'always-takers' may not have volunteered in the absence of the draft - many young men at high risk of being drafted may have volunteered to take advantage of the better terms offered to volunteers compared to draftees). The proportion of always-takers was 19%.

A 'complier' is someone who would serve if drafted, but not volunteer to serve otherwise. A mere 12% falls in this category, meaning that only 1 in 8 men in the draft pool served (or not) depending on the lottery outcome.

[For completeness, a defier is defined as someone who chooses not to serve if selected, while he actually volunteers to serve if not selected. Defiers are simply assumed away, and reasonably enough too]

Wikipedia has some more background information:

The large cohort of Baby Boomers who became eligible for military service during the Vietnam War also meant a steep increase in the number of exemptions and deferments, especially for college and graduate students. This was the source of considerable resentment among poor and working class young men, who could not afford a college education.

As U.S. troop strength in Vietnam increased, more young men were drafted for service there, and many of those still at home sought means of avoiding the draft. For those seeking a relatively safer alternative to the Army, Marine Corps, Navy, or the Air Force, the Coast Guard was an option (provided one could meet the more stringent enlistment standards). Since only a handful of National Guard and Reserve units were sent to Vietnam, enlistment in the Guard or the Reserves became a favored means of draft avoidance. Vocations to the ministry and the rabbinate soared, because divinity students were exempt from the draft. Doctors and draft board members found themselves being pressured by relatives or family friends to exempt potential draftees.

According to the Veteran's Administration, 9.2 million men served in the military between 1964 and 1975. Nearly 3.5 million men served in the Vietnam theater of operations. From a pool of approximately 27 million, the draft raised 2,215,000 men for military service during the Vietnam era. It has also been credited with "encouraging" many of the 8.7 million "volunteers" to join rather than risk being drafted.

Of the nearly 16 million men not engaged in active military service, 96% were exempted (typically because of jobs including other military service), deferred (usually for educational reasons), or disqualified (usually for physical and mental deficiencies but also for criminal records to include draft violations). Draft offenders in the last category numbered nearly 500,000 but less than 10,000 were convicted or imprisoned for draft violations. Finally, as many as 100,000 draft eligible males fled the country.

And here's a couple of interesting paragraphs from Angrist:

Selection of individuals for induction for the draft-eligible, non-deferred 'high priority pool' was based on a number of criteria, the most important of which were the pre-induction physical examination and the examination of mental aptitude. In 1970, for example, half of all registrants failed pre-induction examinations and 20 percent of those remaining were eliminated by physical inspections conducted at the time of induction.

[Also], it is sometimes argued that during the Vietnam era students went to college to avoid the draft and that educational standards were reduced so as to avoid having to flunk students out of school. Baskir and Strauss (1978) claim that Vietnam era college enrollment was 6-7 percent higher because of the draft.
On a vaguely related personal note, this blogger was drafted and served nine months in the military, although that was in Greece during peace-time. I wrote something about the experience here and here, though both posts are mainly about other topics. During that time, I wrote a book consisting of 148 μαντινάδες, exclusively during guard duty (I did innumerable 4-hour stints of guard duty- in the middle of the night, in freezing temperatures, alone, with a loaded machine gun). Any reader who would like to read it, just drop me a line and I'll email you the professionally put-together pdf. Reader beware: it's in Greek.
Read more from this post in Review topics and articles of economics »

Are double-blind trials underestimating drug effectiveness?

0 comments
In a double-blind trial, patients exhibit the placebo (and nocebo) effects because of the expectation they might be on a real drug.

If expectations are so important, could it be that patients being administered the real drug don't react to it fully due to the expectation they might be on the placebo?

Addendum: Steve Waldman (of Interfluidity and Naked Capitalism) posts in the comments:
@llimllib on twitter - Bill Mill - posted a cite (in response to a tweet on this post) that seems like a nice confirmation of your conjecture. Almost perfect.
Wow.
Read more from this post in Review topics and articles of economics »

New developments in econometrics, by Wooldridge and Imbens

0 comments
Jeffrey Wooldridge (a hero of mine) and Guido Imbens delivered this excellent 3-day cemmap masterclass a few months ago, with yours truly in attendance.

The lecture notes and presentation slides are now available online. If you were looking for a book offering an overview of developments in econometrics in the past decade or two, you've just found it - and it's absolutely free, so get downloading.

You can also find more cemmap goodies here (make sure you go through the different years, and navigate a bit around the site). Their seminars and masterclasses are consistently excellent, so if you are interested in econometrics this is a little treasure chest waiting to be discovered.
Read more from this post in Review topics and articles of economics »

Does being beautiful mean being average?

0 comments
Yes, according to this research (via Andrew Gelman):

The debate over the definition of beauty has been waged by both scientists and philosophers for centuries. We tested the idea that a facial configuration close to the population mean is fundamental to attractiveness.

First, we digitized images of faces of male and female college students (i.e., transformed the facial images into little dots of lightness and darkness called "pixels"). Each face is represented by a matrix of pixel values that can be mathematically averaged with the matrices of other faces. Once digitized and averaged together, we can turn the averaged pixel values back into images and have the composite faces rated for attractiveness.

College students rated the male and female composite faces as significantly higher in attractiveness than the individual faces used to create them, if the composites had at least 16 different faces in them. Thus, averaged faces are attractive. Note that when we use the word, "average," we mean the arithmetical mean, and not an average-looking person. If, for example, you take a female composite (averaged) face made of 32 different faces and overlay it on the face of an extremely attractive female model, the two images line up almost perfectly indicating that the model's facial configuration is very similar to the composites' facial configuration.

[...] we view averageness as fundamental and necessary to facial attractiveness. Averageness is not the only component of attractiveness, but without it, no face will be attractive.


Here are some selected publications on the matter.

My two cents: what makes you ugly are extreme characteristics (e.g. big nose or ears); averaging simply takes care of these 'large errors'. The same principle is behind the frequently superior performance of composite forecasts (e.g. of economic variables), where the arithmetic mean of a number of forecasts is often more accurate than any of the individual components.

Postscript: Note that averaging doesn't quite work with hair.
Read more from this post in Review topics and articles of economics »

The normal distribution

0 comments
This beautiful image comes courtesy of W. J. Youden.

I've been reading Edward Tufte's superb The Visual Display of Quantitative Information - perhaps the most perfect book I've ever come across. Almost every page is a revelation, so expect me to be posting more of the wonderful graphs Tufte has collected in the future.
Read more from this post in Review topics and articles of economics »

The importance of being clear

0 comments
A formatting fubar involving an Excel spreadsheet has left Barclays Capital with contracts involving collapsed investment bank Lehman Brothers than it never meant to acquire.

Working to a tight deadline, a junior law associate at Cleary Gottlieb Steen & Hamilton LLP converted an Excel file into a PDF format document. The doc was to be posted on a bankruptcy court's website before a midnight purchase offer deadline on 18 September, just four hours after Barclays sent the spreadsheet to the lawyers. The Excel file contained 1,000 rows of data and 24,000 cells.

Some of these details on various trading contracts were marked as hidden because they were not intended to form part of Barclays' proposed deal. However, this "hidden" distinction was ignored during the reformatting process so that Barclays ended up offering to take on an additional 179 contracts as part of its bankruptcy buyout deal, Finextra reports.


The Register has the full story. As Merv has always warned, 'horrible things happen when you hide cells in excel'.

I see this as a manifestation of a wider lack of education on the importance of communicating information efficiently. The Spartans, Tufte, Strunk and White, the Economist, Picasso and numerous econometricians have done a lot to improve things, but management-speak, TV advertising and other such phenomena show we still have a long way to go.
Read more from this post in Review topics and articles of economics »

Statistical significance

0 comments
...is not a measure of confidence in the point estimate; it says nothing about accuracy.

Statistical significance simply means that the true value of the statistic being estimated has a higher than 5% or 1% probability (the levels conventionally chosen) to be away from a range around zero (the range being determined by sample size and degrees of freedom of the estimator in question). Any statistic which is large enough will be found to be statistically significant even in small samples; this doesn't mean, however, that the accuracy of the point estimate can't be very poor.
Read more from this post in Review topics and articles of economics »

Econometric causality

0 comments
James Heckman has an excellent paper on the subject (free access).
Read more from this post in Review topics and articles of economics »

Multicollinearity

0 comments
Advance warning: This is a tedious post, and it is extremely unlikely you will find it either interesting or informative.

Santosh Anagol is an economics PhD student at Yale and he blogs at Brown Man's Burden. Going through his stuff, I came across a short paper he wrote back in 2004 about the implications of multicollinearity (I won't link to Wikipedia on this, as the article on multicollinearity is lacking and potentially misleading. For more information, read a standard econometrics textbook.)

What he does is simple enough:


with the error normally distributed and uncorrelated with the x's, etc. He then proceeds to run the regression three times, with σ12 (the covariance of x1 with x2) going from zero to .99.
At correlations below .999 our statistical model nails the point estimates and has large t-values. So we don’t need to worry about correlated regressors unless the correlation is EXTREMELY high.
Talking about variables with a correlation of .99 is not very relevant for practical purposes (For many popular datasets, I doubt the correlation between the recorded values and their true values is even as high as .95). In any case, the sample size chosen (1000) is large, and it is not surprising that the OLS estimators yield estimates close to the true value even in the presence of .95 correlation (it is not surprising to an experienced econometrician; see the conclusion to the post). What is more interesting to observe is how the confidence interval around these point estimates changes as σ12 is chosen to be higher. With x's barely correlated, x1 is roughly .13 points wide, with σ12=.5 it goes to .15 units and at σ12=.95 it reaches almost .4 units.

Continuing with the results:
With regressors that have correlations around .99, we get some bad results. In this case the point estimates are off, and one of them is significant. This would obviously be the wrong conclusion about the DGP.



'Statistical significance' is often misunderstood to be a measure of confidence in the point estimate, but it is nothing of the sort. Finding an estimate to be 'statistically significant' simply means that the (95% in this case) confidence interval does not include zero - in other words, there's a low chance that the true value of the statistic in the population is zero, and thus the variable of interest is likely to have an effect on y.

So, the conclusions we would draw about the DGP from the above results are actually the right ones: x1 is not likely to be equal to zero (and it isn't; it equals 2 by construction), and there's a 95% chance that x2 lies between -3.12 and 2.9897 (which it does; by construction, x2=1). The only reason β1 is found to be statistically significant and β2 not is the fact that x1 was picked to equal 1 and x2 was picked to equal 2, so we need to feed our estimators with more information in order to establish that x1 has an effect on y than is the case with establishing the same thing for x2.

The point estimates are indeed off, but this is purely due to the particular random sample - and the large confidence intervals alert us as to the possibility this is the case. Run the same model with a larger sample size (or pick a large number of other random samples and draw the probability distribution of your estimators), and the OLS estimates will be spot on.

And a final observation:

If two variables are highly correlated, will it screw up coefficients on other, exogenous variables? I ran the model with another regressor x3 that was uncorrelated with x1 and x2 , and with a coefficient of 3 in the data generating process. The degree of correlation between x1 and x2 DOES NOT CHANGE point estimates and t-stats of our coefficient on x3.


...which is to be expected from theory. Any explanatory variable that is not correlated with the x's of interest does not need to enter the model at all - it can safely reside in the error term without any bias being introduced as a result. (the Gauss Markov assumptions only call for the error to equal zero given x). The coefficient on x3 would be the same even if x1 and x2 were not included in the regression, and the coefficients on x1 and x2 are not affected by the inclusion of x3 in the specification.

Before leaving this post, I should make clear that I am not critical of Santosh's note; in fact, I think it's great and his effort is to be applauded. From the introduction to the paper:

I’ve been confused for a while about the effects of having x variables that are correlated. This is pretty embarrassing, given this is undergrad metrics stuff. But I’ve also seen enough grad students and professors throw around ”multicollinearity” without really understanding its implications that its worth straightening out.

This is not 'undergrad metrics stuff' at all. It is true that economics undergrads learn about the qualitative effect of 'multicollinearity', but developing an understanding of its significance in practice only comes after substantial exposure to the literature and hands-on experience (as with so many things econometrics). Santosh's attitude is the right one, and playing around with simulated data is a great, low cost way to digest the theory and really understand econometrics - one that tutors should be encouraging far more than is currently the case.
Read more from this post in Review topics and articles of economics »

Econometrics versus calibration

0 comments
Some macroeconomists use calibration, some use econometrics, and some use both. There’s no real methodological debate left in the field on this issue.

What is true is that most people outside of macro do not like calibration. I don’t know why. I spent seven years of my life thinking about whether econometrics was better than calibration … and pretty much decided that the answer is: “it depends”.

That's N. Kocherlakota, from his essay on the state of modern macro. He's absolutely right, and keep in mind that I am an econometrician.

Part of the reason is that macroeconometrics is, well, dodgy - and that's 'at best'. A lot of macroeconometrics is outright bonkers.

Another issue is that econometrics is generally misunderstood; most of the 'data' on any issue usually resides on the researcher's head, meaning that in most cases the informational contribution of hard-coded data is much less important that that of theory ('soft-coded data'?). There's no such thing as 'letting the data speak for itself', but the notion sounds appealing so people keep pretending there is.

And if you take issue with the previous sentence, good luck discovering the next Phillips curve, deciding whether GDP has a unit root, and estimating any elasticity with a non-laughable degree of precision.
Read more from this post in Review topics and articles of economics »

Stata lessons & other resources

0 comments
A friend asked for a quick list, so here goes:

UCLA's excellent resources to help you learn and use Stata

Another great collection of Stata Resources by Park Hun Myoung, as well as a stata command cheat-card

London School of Economics Stata Resources

Syracuse University's Stata tutorial

Program in Statistics and Methodology by the pol sci's at Ohio State University

Duke Stata tutorials
Read more from this post in Review topics and articles of economics »