Saturday, October 18, 2008

Mysteries of financial risk, plus: House on fire

The idea floating around (one version proposed by McCain) that the government should buy up "troubled" mortgages is as misguided notion as they come. "Troubled" itself is hard to define. Is it determined by the borrowers' difficulties in making loan payments? Or by loan default? Or is it just a mortgage "under water," valued at more than the house it's attached to? "Troubled" should be limited to, at most, the first two cases.

The housing crisis -- or rather, the house financing crisis -- will have, not one, but many endings, covering a range of possibilities:
  • Borrowers paying reliably on mortgages "under water"
  • Borrowers with payment difficulties who renegotiate their loans (lower interest rate)
  • Borrowers in default who might renegotiate or just move out, to rentals
  • Borrowers in foreclosure who must move out, or stay and rent with option to buy
At this point, the last category is small, a little over a percent of mortgages. The range of options is enough to make simply throwing people out on the street an unnecessarily harsh choice. Banks and lenders will not want to sit on unoccupied, non-income-generating property in any case. Both lenders and borrowers will unavoidably take some losses along the way.

Except for directly intervening with borrowers with Fannie and Freddie loans, it's hard to see what role government should take here, except to act as a catalyst. Government should certainly not be engaged in perpetuating the housing bubble; for example, in trying to prop up house prices or encouraging any more subprime lending. If it does anything for the housing market, it should be terminating the ingredients that went into the bubble in the first place.



The general financial crisis, centered in the credit markets and impacting others (like the stock market), was certainly triggered by the weakness in the subprime mortgage market and exacerbated by falling house prices across the board. But the financial system, as evolved over the last thirty years, has developed intrinsic weaknesses of its own that falling house prices merely exposed. Those dangers are embodied in excessive debt and rationalized in turn by faulty theories about controlling risk.

Many of the supposed culprits -- mortgage bonds and "derivative" securities (essentially, complex, composite repackagings of existing securities); the non-existent "deregulation" of Wall Street; and the alleged merging of investment and commercial banking -- are bogus. These supposed factors are either not real or not capable of producing an unforeseeable credit crisis of this magnitude.

Over the last generation or so, the financial world, American and non-American, the regulated and the regulators, has developed an unhealthy and misplaced confidence in its ability to quantify and manage risk. The crisis we see unfolding now has nothing in the slightest to do with "fraud" or malfeasance on any individual's part.* Traditional regulation is designed to deter and punish such misbehavior, which is multiply times over illegal anyway. A crisis of this type is a result of collective misjudgment and collectively-held false ideas about risk, mixed with a certain level of hubris.

Viewed this way, our present financial troubles start to look less like a crime caper and more like the failure of a complex technological system, like the explosion of the space shuttle Challenger or the sinking of the Titanic. Megan McArdle had an interesting post on this point a while back.

To follow Megan, it's especially enlightening to compare the failure of financial risk management with the Challenger explosion, on which topic she recounts the story in Richard Feynman's famous What Do You Care What Other People Think? and captured in detail in Feynman's appendix to the Rogers Commission report. The key comparison: the different ways that different people interpreted "small" risks. Based on decades of prior experience with rockets, the engineers knew in their bones that the "small" risk of a fatal shuttle accident was about one in a 100. (And we know now, with over 25 years of shuttle experience, that they were right.) But they couldn't articulate and defend their point of view in the face of managerial and political figures, whose notion of "small" was more like one in a 100,000 or one in a 1,000,000. Each near-fatal incident, instead of being interpreted correctly as a warning, was instead rosily misinterpreted as "great, we survived another close one" and falsely built up NASA's confidence.

That difference -- "small" as one in a 100 or one in a 1000, versus "small" as one in a million or ten million -- is precisely the difference between the "wild" and the "mild" in risk, "Extremistan" versus "Medocristan." Readers of previous posts on finance and statistics will know of Nassim Nicholas Taleb's The Black Swan and all about such misperception of risk. A one-in-a-hundred incident is something likely to happen more than once in a person's lifetime. A one-in-a-million or ten-million incident is unlikely to happen in anyone's.

It makes the crucial difference to social systems created and run by humans. All of us, especially the college-trained, are prone to the Tyranny of the Cookbook, falsely believing that some answer is better than no answer, even if that answer is wrong. Much of the financial world still wrongly assumes the mild risk of Medocristan and rationalizes the powerful evidence to the contrary by handwaving.
---
* It has even less to do with "corruption," something outside of Wall Street's power, since that requires the granting of political favors. You have to look to K Street (in Washington) for that.

Labels: , , , , , ,

Friday, June 06, 2008

Entropy, information, and the ice cube

Looked at the right way, a humble ice cube can teach you a lot about thermodynamics. As Nicholas Taleb points out in The Black Swan, the simplest facts about how it melts contain the kernel of the Second Law of thermodynamics by implication. Elaborating upon the simplest case illuminates many other situations far more complex.

The Second Law is not deterministic, but probabilistic. It doesn't say, the system has to evolve in such and such a way. It just indicates, given where a system is now, what the most probable direction it will evolve in. Thermodynamics comes into play when we can't know the exact state of every water molecule. Instead, all we know are certain fixed totals about the system: the total energy, the total volume, the total number of molecules. In the very simplest case, total isolation, the ice cube does nothing.

In the next simplest case, the cube is isolated except for contact with a heat bath, which is defined by a temperature, which we'll assume is at or above the ice melting temperature. The Second Law tells us the overwhelmingly probable evolution of the cube: it will change in the direction that increases its entropy or disorganization. That includes heating up (acquiring heat from the heat bath). In this case, that includes melting.

When the entropy increases, information is lost. Imagine that the ice cube has some features on its surface, or that it was carved into some shape. After it's fully melted, all that's left is puddle of water. The puddle of water is the final result, regardless of whatever funky features the solid ice had in its shape.

Since its discovery in the 19th century, the Second Law has had a sad countenance, apparently nothing but a tale of decay and decline. Certainly the thought that an elaborately carved ice sculpture and a plain ice cube of identical mass might end the same way - as a large, featureless puddle of water - made thermodynamics seem like the truly dismal science.



But these are only the two simplest possibilities for the ice cube. It could be in contact with a chemical bath of some substance that reacts and binds with water (hydrates). In that case, two processes, melting and hydration, proceed simultaneously. The larger and/or faster will predominate, since it increases the entropy faster. Melting might not happen at all, because the Second Law would then have another and better avenue to satisfy itself. The Second Law doesn't tell you that melting has to happen; just that, whatever avenues of change are available, the one that gets you to higher entropy faster wins. And it only specifies a probabilistic tendency, specifying nothing in general about rate, except that it's positive.

By historical convention, systems in passive contact with external "baths" are not considered "open." They are in thermodynamic equilibrium with themselves and any "baths" in contact. Truly "open" systems are ones with "flow-through," where matter, radiation, and/or heat flow in and flow out. "Open" systems are not in equilibrium, in general, either with themselves or with the outside. In that case, the Second Law still applies, but it only applies to the whole system and its environment. Any part of the whole can see its entropy decline, so long as the entropy of the whole rises. If an open system exhibits a strong spontaneous tendency under certain conditions to lower its entropy and acquire structure, it has to expel the excess entropy outside of itself. The system is said to be self-organizing.

Hence, biology, evolution, and weather.

The competing rates of different processes become a more complex but even more critical tangle in the self-organizing case. For self-organization to succeed, the spontaneous structure has to form faster than any competing process (dissipation) importing entropy back into the system from the outside. Thermodynamic equilibrium is not valid for "open" systems as a whole, but might be valid for parts of the system. Typically, equilibrium is an excellent approximation for suitably "small" part of the system, where local temperature and pressure can be defined and local thermodynamic equilibrium (LTE) holds. But it remains true that on intermediate to the largest scales of the system, LTE is badly violated. These are exactly the scales over which flows of matter, radiation, and heat are most obvious.

One of the things that makes weather and climate prediction so hard is that on intermediate to large scales, the evolution of the atmosphere is not, on the one hand, a simple application of determinism: the system is chaotic; but on the other, not a simple application of thermodynamic arguments either. On short scales (a few to a few tens of meters), the atmosphere respects LTE pretty well, and thermodynamics can be used to predict its evolution (if boundary conditions are known). But on the scales of storms, cyclones, and fronts, the atmosphere, while "thermodynamic" in some sense, is nowhere close to thermodynamic equilibrium, and the simple probabilistic arguments of thermodynamic equilibrium don't apply. It's chaotic enough to make long-term, detailed prediction impossible; but not so chaotic that simple statistical arguments can be used instead. It's somewhere in between: highly sensitive to poorly known initial conditions and past history.

This realm - in between simple linear predictability and simple statistical equilibrium - is not only a result of chaos, but constitutes a distinct area of dynamics and physics, usually given the name complexity. It's a dynamical regime rich with unpredictable structures that repeatedly form and dissipate - like weather, or living things. (June 2)

POSTSCRIPT: The conventional global temperature index continues its recent precipitous drop. In case you've been hiding in a basement the last few months, this was the coldest spring, and the coldest May, in many years.

It should be stressed again that this conventional global temperature index (one of a handful of composite statistical indexes used by the IPCC and others) is not the temperature of anything. The Earth has no single temperature: it's not in thermal equilibrium, with either itself or a "heat bath." It's a complex weighted statistical composite of many individual temperature measurements of the air, ocean surface, and radiation. (The first two are local; the third is nonlocal, by its nature.) The attempt to pass this composite off as the unique "temperature of the Earth" is one of the many major fallacies of the climate change hysteria.

The direction and timing of the trend are not in doubt. Something has been happening in the last decade and has accelerated. That something is cooling. But the exact nature and magnitude of the trend can only be understood by disaggregating the composite index back into its originating individual temperature measurements and looking at their trends in time and space.

Labels: , , , , , ,

Wednesday, April 09, 2008

Chaos and markets II

For those who teach finance, a number seems better than no number — even if it’s wrong.

- Mandelbrot and Taleb

It's the Tyranny of the Cookbook, to which we can reply: No number is better than some number - especially when it's wrong.

Much financial advice is common-sensical, but in recent decades has incorporated misguided notions of implicitly Gaussian or bell-curve statistics in analysis of price movements, as well as false concepts of "efficient markets." You often hear the jargon of means, variances, and betas. (A beta is just price volatility defined as a second moment, or a variance. The square root of the variance is the standard deviation.) We've already seen a truckload of examples of where and why such concepts break down and why methods based on such assumptions are wrong. The fact that this approach to finance has a Nobel prize is irrelevant.* The methods and concepts have spread from academic finance and economics departments to the desktops and minds of investment specialists in the last 30 years and done significant damage: the Long Term Capital Management crisis in 1998 and the mortgage crisis of 2007-08 were both made possible, in part, by such "professional consensus" malpractice. Here we have legendary cases of Platonified false expertise and the "empty suit" syndrome. The price change distributions are fractal-driven power laws, not bell curves, a fact first presented to the economics world almost 50 years ago by Mandelbrot - and then rejected because it didn't fit convenient, if unempirical, Mediocristan assumptions. The missing practical key is the widely unrecognized enhanced risk of large fluctuations, especially downward moves. Individuals and institutions adopting wrong rules expose themselves unwittingly to much larger risks than they realize.

If we drop the assumptions of bell-curve price fluctuations and efficient markets, where do we stand?

The first is basic math and science: get your units straight. People who practice finance usually get this right, but it's amazing to see ignorance even in the business pages about this. Economics, like mechanics, has three basic types of units: money (a universal store of value and medium of exchange), things or activities (count them distinctly and don't commit the Fallacy of Aggregation, lumping bananas and pork bellies, say), and time. The essential point is that wealth is an accumulation of flows. The flows are prices (measured in money) times things or activities (quantified somehow) divided by increments of time. Interest rates are prices divided by prices divided by time, or just 1/time. Wages are money per unit of labor (an activity) per time. And so on.

The principle of diversification remains, but its rationale changes. It's not "everything will even out" (it doesn't always), but "we don't know very well how individual investments and investment classes will perform - sample all of them." Diversification, not only within investment classes, but especially across classes, is even more important in Extremistan than in Mediocristan.

More basic to the uncorrelated, Gaussian price movement picture is the efficient market hypothesis, which has failed in a number of crucial respects. Market timing matters, especially if you're making large moves (investing or liquidating). The market analysis based on this wisdom is called "technical analysis" or "charting," and its advocates are called "chartists." They stare at price chart patterns. In the "uncorrelated random walk" picture, these patterns mean nothing. But in fact they do mean something. Market moves are indeed correlated across time. Only after three to five years do they start to lose their memory, and it's not clear that they ever entirely do.

Furthermore, there are investment classes that consistently under- and overperform the whole market average. The best-known underperformer is the class of "growth stocks," because they're hyped by the media and analysts to the point where buyers demand them strongly - they're consistently overpriced relative to their long-term performance. OTOH, there are underpriced investments: so-called "value" stocks, for example. Warren Buffet and others have made a fortune hunting for undervalued but worthy investments. It's all boils down to not paying more for an investment than it's worth.

Finally, the "fat tail" phenomenon should make everyone suspicious of probability distribution moments (means and variances). If misanalyzed using Gaussian assumptions, fat-tailed distributions appear to be non-stationary: if you keep sampling such distributions to estimate moments, your results will not, in general, converge as you add more data points. The estimated moments will just keep growing. After an infinite amount of sampling, they diverge to infinity. While means and variances are measures of performance, they're not good measures.

The devil's staircase. A better approach than looking at daily movements is to look at cumulants (integrals) and at absolute linear ranges (price highs - price lows). The cumulant is more stable than the daily changes in value, and sudden jumps in the total value of an asset or flow of goods and services show up clearly. (The fact that such sudden jumps often dominate the total or cumulative history of an asset or market also stands out clearly.) The absolute linear range grows with time, but gives you some sense of the best and worst the market can do. These are the rules of the road in Extremistan. "Mild" variables change by a large number of small increments. "Wild" variables change by a small number of large increments, and "really wild" variables change mainly by a handful of very large increments.

Markets with an incomplete cookbook. The investment community at large still has not fully absorbed Mandelbrot's message about fractals and the uselessness of Gaussian, bell-curve statistics in understanding and prospering in markets. The normal and the Levy-type distributions look similar when you compare them for small deviations from the mean.** It's the large deviations that constitute the acid test, and it is here where investment professionals often start waving their hands.† In a Gaussian world, such large changes shouldn't occur almost ever, and the history of Gaussian markets would be dominated by many, many small changes. But real markets are strongly shaped by a limited set of rare, large, and consequential events. A new investment science to replace the rigorous, Platonified irrelevancies of contemporary financial theory is badly needed.

POSTSCRIPT: Here's a short note on market risk by Mandelbrot and Taleb from a few years ago.

References

= B. Malkiel, A Random Walk Down Wall Street, rev. ed. Classic presentation of efficient-market, Gaussian random walk theory to the masses. Much of the technical side is wrong as a picture of markets, but the basic investment advice (the trade-off between active and passive investment, diversification) is sound.

This posting is a sketch of what's needed to replace the bell-curve price movement framework. Just noted today: the embarrassing underperformance of stock index funds since the 2000 market peak, compared with even lowly bonds, not to speak of value stocks.

= R. Haugen, The Inefficient Stock Market. Nice short, if technical, study of systematic inefficiencies (over- and underpricings) in markets.
---
* Black and Scholes won it in 1997, and Taleb and others have railed against this as a perfect example of rewarding Platonified bullshit with its origins in academic circles, with highly restrictive assumptions, applied to real life where those assumptions don't hold. The LTCM crises occurred less than a year after the award - again suggesting a just G-d, or perhaps one with a refined sense of humor.

A larger objection can be made against the economics Nobel prize altogether, and Taleb and others argue that as well. It's actually a Nobel foundation prize paid for by the Royal Bank of Sweden, not specified in Nobel's will. Although some great and deserving economists have won it (Hayek and Friedman among them), in general, it's difficult to argue with the reality that economics has often been subject to both fads and conveniently cookbook pseudoknowledge. The standards for the Nobel prizes in the natural sciences are much stricter, and I hope they remain thus, so that at least those Nobel prizes mean something.

** Actually, the log-normal. The Gaussian bell-curve is applied, not to prices, but to the logarithms of prices. Small changes in prices are then translated into small percentage changes. (For price P, the differential dP is replaced by dP/P.) For small ΔP's, the log-normal and Lévy-type distributions look almost identical - it is here that the theorists of the Gaussian random walk go astray.

The "random walk" idea can be taken beyond the Gaussian or normal type and recast into a more general form of Lévy flights, dropping the requirement of finite distribution moments. To handle correlations over time between events, it can also be generalized in another way, to have memory: fractal random walks. Such erratic "random" or "drunkard's walks" are an important tool for applying statistical methods to dynamics under conditions of limited knowledge. The random walk is also central to analyzing diffusion (both standard Gaussian and "anomalous" fractal types). In chemistry and biology, the random walk is sometimes called Brownian motion.

† In the last generation, improvisations have grown up around the failure of Gaussian methods, but this series of ad hoc patches and fixes doesn't get to the root of the problem. Some analysts still just take out large deviations ("outliers") by hand, a kind of data denial. Others appeal to the notion of "exogenous" (outside-the-system) shocks, which destroys the method's predictive (if not its retrospective) powers.

The most sophisticated patch is to make the Gaussian parameters depend on time, the common version being GARCH. This is the best you can do within the misguided Gaussian framework; in that wrong framework, the actual (and probably stationary) distribution of price movements looks non-stationary. The time-dependent parameters are supposed to mimic this, but at the cost of largely destroying the method's predictive power.

Labels: , , , , ,

Tuesday, April 08, 2008

Chaos and markets I

Chaos as an idea and metaphor applies not only to the natural sciences, but the study of human society as well. It's always important to clearly distinguish between metaphors, and models and analogies. The former are loose and poetic, and treacherous if you try to extract precise conclusions from them. The latter are misleadingly precise and suffer from the illusion that everything about chaos can be reduced to cookbook. Fuzzy is comprehensive but imprecise; precise only seems under control, because the untamable part of chaos gets excluded before you even start. When facing chaos, it's better to be approximately right than precisely wrong. So caveat emptor.

Financial and other economic markets are prime examples of chaotic behavior in human life. They feature individual agents acting rationally, but with limited information and often conflicting goals. The torrent of financial information available gives many people the illusion that, somewhere, someone knows what's going on.* Actually, the people in charge of large institutions and the power to set rules of the game are often some of the more poorly-informed actors, precisely because the scope of their responsibility is so broad and the impact of their decisions so difficult to fathom ahead of time.

There are experiments in behavioral economics that do yield important and controlled information about human economic reasoning and decision-making. But the whole subject, while fascinating and full of insights about the limitations of "economic rationality," is in its infancy. Hopefully, in coming years, the results of behavioral economics will come to displace the "likely-story" Platonified and often false mathematical models that have ruled in economics and finance since the 1960s.**

Economies and cycles. Economic evolution does show some characteristics of irregular waves and more regular cycles. The best known is the six-to-ten-year business cycle, which is an investment-driven cycle in which consumption of what is produced is the final step closing the loop. Recessions occur at the end of these cycles, when investment and consumption across the whole economy tend to get weak all at once. There are shorter-term cycles, of roughly two to four years in length, which are inventory or "reservoir" cycles associated with economic demand rising and falling in various sectors (like housing, in the current bust). They're waves of building up and depletion of inventories. These waves of bubble and bust can be amplified by bad government policy (again, as we're seeing now).

Longer, irregular waves of economic activity are harder to pin down, but well-attested in the historical record. The best-known, if still controversial, is the Kondratieff wave, of approximately 50 to 70 years in length. It's the "two-generation" economic wave.

Cycles! None of these phenomena is simply periodic, but irregular - multiperiodic, shot through with some chaos.

Economics and statistics: Normal, log-normal, and power laws. If we forget about specific events, specific times, and specific histories, we fall back on a statistical description of economic change. Individual events get binned by type, character, and frequency. When we look at economic change as fluctuations of prices and flows of goods, we see the effect of both "normal" and "fat tail" processes everywhere. These form an object lesson in the power of "black swans" and the larger crowd of "grey swans" to shape economic history.

The mainstay of quantitative finance is the log-normal distribution (figure at the top of the posting), where individual instances are assumed to multiply, not add (hence the logarithm; the sum of the logarithms of individual factors is the log of their product). But it's easier to compare normals with their Lévy generalizations. The Lévy distributions have the Gaussian as a limiting case (family of distributions labeled by exponent α, with α=2 as the Gaussian case).

Here's a graphed set of Lévy distributions. The normal curve is the black curve:



The Lévy distributions relevant to finance are those with α close to, but less than, two. Notice that for small deviations x from the mean (zero), the "α-close-to-2" distributions don't look that different from one another. It's the "outliers" that make the difference clear. For deviations x far from the mean, the non-Gaussian curves fall off slowly; in fact, as power laws ~ |x|-(1+α). These distributions are sometimes called scalable, because they have no fixed, intrinsic scale of deviation that sharply limits how big the deviations can be and that forces the distribution to remain close to the mean.†



This log-log plot shows how much more sharply the Gaussian (black curve) falls off with x than do the the Levy generalizations. Large deviations remain less likely than small; but they are far more likely in the non-Gaussian, power-law, case than in the Gaussian.

The mild versus the wild. The financial world straddles two paradigms.††
  • Mediocristan is a world of conserved or almost-conserved total quantities. They tend to get subdivided in roughly equal ways among all possibilities. The flows of goods, labor, and services (as opposed to their prices) tend to have more of a "mild" behavior, at least over limited periods of time. When they change, the usually change slowly. Rapid changes are rare (but not unknown); large changes are more frequent, but usually happen over months and years.

  • Extremistan is a world of non-conserved total quantities. For example, the total flow of economic value associated with the flow of some thing or some activity is its price (say, dollars/donut) times its physical flow (say, donuts/day). Everything associated with prices (including interest rates and wages, which are prices for capital and labor, respectively) is inherently subject to full-blown "wildness." Cumulative change is often a result of a fairly small number of big events, with the large crowd of small events making not much difference to the total.
Coping with the chaos of prices and related information requires learning two apparently contradictory lessons:
  • How to ignore daily fluctuations, not take numbers in isolation, and not worry about undefined hypotheticals. Few days are important in the grand scheme of things.

  • How to keep an open mind to the occasional "grey swan" and the rarer but consequential "black swan" - which, when it happens, can happen in a few days, or even hours.
We'll look at these next.

References

= B. Mandelbrot, The Fractal Geometry of Nature and (with et al.) Fractals and Scaling in Finance. The first is a modern classic and should be read by anyone with the slightest interest in mathematics. The second is an empirical study of price movements-cum-critique of Gaussian quantitative finance.

= N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. Reviewed here.

= M. Lax, Random Processes in Physics and Finance. Much more technical, part of the burgeoning field of econophysics.
---
* If they are assumed to also be in control of everything, we have a conspiracy theory.

** There is a branch of statistical physics, called frustration or quenched disorder theory (spin glasses) that treats systems evolving under multiple, conflicting, and random constraints.

† The Lévy distributions are the class of probability distributions that enjoy the property that a sum of Lévy-distributed variables is also distributed according to a (slightly different) Lévy distribution. This is like the central limit theorem of the Gaussian bell curve, but more general. It doesn't require the distribution moments, or weightings, to be defined. In the general Lévy case, they're infinite anyway, because of the "fat tails" for large deviations from the mean.

†† This distinction, earlier than Taleb's, is due to Mandelbrot.

Labels: , , , , ,

Sunday, March 23, 2008

Fat tails and outliers: A closer look

No, it's not about the Fat Tonys of the world, Taleb's proverbial cabdrivers who know at least as much about events as so-called experts, not because they're so smart, but because the so-called experts know far less than they think. But Fat Tony might appreciate the world of "fat-tailed" probability distributions, since they provide the mathematical way of capturing, in part, the phenomenon of the black swan: why large deviations from the mean ("outliers") are less common than small ones, but still much more common than expected on the basis of the normal or Gaussian bell-curve distribution.

The Gaussian distribution is used so much because of an important mathematical result, the Central Limit Theorem (CLT). It states that, if we consider a large number of instances of a random process, the collective "distribution of distributions" is Gaussian, if certain conditions hold. These conditions are that:
  • The individual instances making up the distribution must be independent of one another.
  • The moments, or weighted averages, of the original probability distribution must be finite.
What happens in the "large numbers" limit, if these conditions hold, is that, of all the moments of the original distribution, only three matter after the dust settles - the total population size, the mean, and the variance (the zeroth, first, and second moments - see below). All the other moments either vanish or are controlled by the first three. These three are exactly the ones needed to define a Gaussian bell curve.

A simple example. Let's consider a population of particular instances of some property or attribute, quantified by a random variable x, allowed to range from -∞ to +∞. Its probability density is f(x); within an infinitesimal range dx, the total number of instances between x and x+dx is f(x) dx. The cumulative number of all instances of x < X is the integral of f(x) from -∞ to X. Define the nth moment (or weighted area under the curve) as M(n) = ∫ xn f(x) dx. The non-negative integer n = 0, 1, 2, ... ∞.

The Gaussian with zero mean and variance of one is f(x) = exp(-x2/2)/√(2π). (The normalization is chosen such that M(0) = 1.) It is strongly peaked at x = 0 (the mean) and falls off rapidly for deviations from the mean.

The "fat tail" case occurs when, whatever f(x) is doing for small x, it decreases for large x as |x|-a, a > 0, apart from overall multiplicative constants. f(x) falls off for large x, but far more slowly than the Gaussian does. Then M(n) ~ ∫ |x|n-a dx. Replace the upper (lower) limit of the integral with +X (-X), X → +∞. Then M(n) ~ Xn-a+1. There are three possibilities:
  • n - a + 1 < 0. The moment M(n) is defined (convergent or finite).
  • n - a + 1 = 0. The moment M(n) is infinite, diverging logarithmically.
  • n - a + 1 > 0. The moment M(n) is infinite, diverging as a positive power.
For a "fat-tailed" distribution behaving this way, while some moments (for lower n) might be defined, the remaining moments n > a - 1 are undefined. Therefore the CLT does not hold, and it is not correct to use Gaussian-based statistical methods for such populations.*

Long before Fat Tony.... Such distributions are called, in the mathematical literature, Lévy flights, after the French mathematician Paul Lévy, who first worked with them in the decade prior to the Second World War. Mandelbrot, the geometer of fractals, was a student of Lévy. Both Lévy and Mandelbrot went into hiding after the French defeat in 1940, avoiding the Nazi and Vichy dragnet of French Jews.

After the war, they were also intellectual refugees from a certain style of mathematics that swept over the French academic world and had a strong influence elsewhere. Collectively named the Bourbaki school, it drove applied and "heuristic" mathematics to the margins of the field and favored a lean, abstract approach of theorem-proof, with no pictures, diagrams, or applications. (It was the same period that the artistic avant-garde moved strongly in the same direction: away from sense perception, toward "pure" abstraction.) The situation relaxed in the 1970s and 1980s, followed by a strong revival of interest in applied mathematics both among mathematicians and scientists and engineers who use mathematics. While rigor and precision are essential to mathematics, it can't survive or even make sense without contact with applied problems and the world of the senses, and the Bourbaki revolution petered out.

Using Lévy's results, Russian mathematicians Gnedenko and Kolmogorov proved a generalization of the Central Limit Theorem that allows for systematic statistical methods to be applied even in such Extremistan cases. But the resulting "distribution of distributions" is not Gaussian. If we want to study the statistics of events in a chaotic system, like the climate or financial markets, say, we must use these generalized methods pioneered by Lévy, not the 19th-century methods of binomials, Poisson, and Gauss. Like 20th-century artistic palettes and musical styles, it's a 20th-century statistics cookbook of expanded possibilities and greater generality. In the next posting, we'll meet a recent climate case where appropriate statistical methods were applied, with striking results, to a situation where wrong methods were long used.
---
* Usually, a > 1 in practice. If 0 < a < 1, then even the zeroth moment M(0), the total number in the population, is infinite. (The mean and variance are undefined as well.) Mathematicians can still cope with cases where some or all the moments diverge, by using something called the generating function of the probability distribution.

Labels: , , , , ,

Thursday, March 20, 2008

Meet the thinkers: The curious aviary of Dr. Taleb

Cygnus atratusWe also know there are known unknowns; that is to say we know there are some things we do not know. But there are also unknown unknowns - the ones we don't know we don't know.
- Donald Rumsfeld



Some of us been waiting for something like this book for a long time, and its Levantine author has come a long way - all the way from the hills of northern Lebanon and the Syro-Greek Orthodox town of Amyoun. The book is The Black Swan: The Impact of the Highly Improbable, and the author, Nassim Nicholas Taleb, former financial trader and now extraordinary professor of the inexact sciences at the University of Massachusetts, Amherst, etc., etc. - essentially, the Dean's pet, and they don't know where to put him. The Black Swan is one of the most important science books for a non-science audience in many years. Like the best chaos and complexity books of a decade or two ago, The Black Swan deals with scientific questions arising from the stuff of everyday life, not far-off galaxies and times long ago.

The core of Taleb's point is the impact of what we don't know, the improbable, and how "randomness" is really another name for our ignorance. But Taleb has a larger target: a whole book was needed to attack and dismantle the legitimacy of bell curve statistics, "Mediocristan" methods wrongly applied to "Extremistan," and explain why so much of the world doesn't follow the "middle of the road" behavior prescribed by the Gaussian-normal distribution and its cousins, such as the binomial or Poisson distributions.

The book is rich with fallacies exploded:
  • The Ludic Fallacy. This is the fallacy we pick up when we learn probability based on tightly constrained assumptions, "rule of the game," that make understanding statistical methods based on them as easy as an elementary cookbook. (Ludus is Latin for "game" or "fun.") Real life often presents us with situations of limited knowledge, where probabilistic thinking is appropriate, but where we don't know the "rules of the game," at least not all of them. Many trained in probability and statistics apply the cookbook methods anyway, for the lack of anything better. They capture risk - the known unknowns - but not true uncertainty - the unknown unknowns.

  • The Narrative Fallacy. This is a biggie, practiced on an industrial scale by the news media, every day. We draw connections between dots where the real connections are different, or don't exist, or there are no dots to be found. The news media does it to keep our attention with frequently made-up stories, or "narratives," to use the post-modern jargon, that seem better than no story, or a different one.

  • The Narrative Fallacy supports a related fallacy, one of Misattributed or Reified Intentionality, the fallacy that human society is collectively a result of human intentions or consciousness. In fact, most of it is not, and attempts to force it to be so have led to one disaster after another. Our minds are too limited and possess too narrow a scope of awareness to make this possible. Human society is mostly made behind our backs, so to speak. Taleb's developed views on this question end up very close to the views of the famous Austrian school of economics and sociology.
Taleb has an outrageously funny time explaining what went wrong with statistics and the social sciences in the 19th century, when they were invaded by the concept of the Average Man, and everything was reduced to bell curves, means, and small variations.* All this would hold if our world were Mediocristan. But much of our world is not.

Who's Stan, and what's the difference? Mediocristan is tightly constrained by "fixed totals," or what physicists call "conservation laws." We've already met these and seen what they do. They force the collective behavior into highly restricted patterns, with "equipartitions" of energy, or number, or volume. This certainly is an aspect of our world, and not just in thermodynamics. Heights and weight, for example, both of them strongly limited by gravity and metabolic limits, are distributed in a way close to the bell curve. But then again, consider the distribution of weights in aquatic animals, and you can already see: without gravity, the maximum size is much bigger (think of whales and octopi).

The key to Mediocristan is the Central Limit Theorem. If a population's distribution (of whatever attribute) is made up of independent instances and has well-defined moments (weightings), then the distribution approaches the bell curve in the limit of "large numbers." The presence of "fixed totals" guarantees well-defined distribution weights (moments).

But in many, perhaps the majority of, cases, it fails. The instances are not independent of one another, not distributed with well-defined weights, or neither. The distribution then has much less reason to clump near the mean. In fact, in such cases, many of our usual statistical clichés (means, variances, medians, etc.) fail to capture what's going on.

This is the world Taleb calls Extremistan.** If there's no "fixed total" of something being distributed (like economic wealth, or the total number of books sold by a single author, say), there's no reason to think that the total will be broken up in a roughly even way among instances. Here is the key to understanding much of our world - economic markets, wealth, and income in particular. Many days on markets are boring. Some are interesting. A few are extraordinary - and it these days, the black swans of the financial world, that end up dominating the cumulative history of the market. Just look at the last few months' newspapers.

We encounter similar truths in biological evolution, in contrast to the anodyne but wrong gradualism still dominantly taught. Most of the cumulative change in biological evolution is due to a small number of extraordinary turns of events that have outsized impacts echoing through the millennia. Ditto for human history.

And of course, on Taleb's home ground of financial markets, the reality of black swans, and fractal or fat-tailed distributions, is of intense interest. The disastrous application of bell curve-based statistical methods to quantitative finance in the last generation has not made markets better-behaved or investment strategies sounder. On the contrary: the 1998 Long Term Capital Management and 2008 mortgage crises make clear just how wrong these methods are. They're "state-of-the-art" in some sociological sense, but it's a mistake to call them an art, much less a science.

We've met these strange birds already: Taleb's black swans are the stream of unique events of chaos. His grey swans are those occasional, semi-tamable events at the low frequency end of the spectrum.

Plato in Nerdistan. As the book develops in its middle, Taleb wanders through the thickets of epistemology, how we know what we know. This part is somewhat weaker than the book's earlier and last parts, because the argument goes too far afield and loses a bit of focus. Taleb over-blurs the distinction between event (his specialty) and entity. Before European explorers reached Australia, they believed that all swans are white. The whiteness was not an essential part of the definition of "swan," nor was the belief obviously false. It was a contingent statement about two different properties of things: "swanness" and "whiteness." This supposed connection met its end when the explorers encountered the black swans of Australia. A deeper lesson took a longer to sink in, and some still resist it: disproving something is much easier than proving it. Proving something requires understanding its nature more deeply and thoroughly than our knowledge often runs.

Even this middle part is rich with deserving targets. Taleb calls them "Platonified abstractions," the stuff of academic knowledge. They're thrown around confidently by people who don't know what they don't know. This might almost be a definition of nerdity: what you know fits into cut-and-dried abstractions, and you confuse these with the actual world only known to us very imperfectly. Nerds stand in counterpoise to Taleb's foil, the Fat Tonys, the proverbial cabdrivers of the world who know better and who understand that when it comes to Platonicity, you can take it or leave it.

What do you know, and how do you know it? Exact human knowledge is coined under laboratory control or by precise logic. Most of the knowledge we use in everyday life is approximate knowledge in well-defined, if not controlled, conditions. At the edges of what we know is amorphous knowledge, often mixed in with a lot of prejudice and guessing. And if we want more and better knowledge, we face the reality of trade-offs. I can be sure something will happen today, but I don't know its significance. I can also be sure something significant will happen in the next year, but I don't know when.

Modern science is not based on induction, contrary to common belief. It's based on a mixture of hypothesis, deduction, controlled experiment, and controlled mathematics. It's not because scientists are dogmatists that they live by deduction. It's because deduction allows one's reasoning to be kept under precise control, with all the assumptions on the table and the steps clear. Induction (like statistical correlation) can certainly be strongly suggestive of hypotheses, and it's essential for developing logical definitions. But you can't prove anything with it. One counterexample - the black swan - destroys it. Silent evidence is always lurking to upset the induction cart.

The essence of probability. Coping with limited knowledge means falling back on probabilistic arguments, and this is in fact the origin of statistics. Its modern founders (Pascal, Bayes, Laplace, Gauss) all identified probability with a greater or lesser sense of certainty about something, not its frequency. This distinction fueled a great 19th-century debate between Bayesians and frequentists. Until the 1920s, the frequentists had the upper hand. But modern mathematics has abandoned frequentism, except as an approximation in carefully circumscribed situations where the Ludus isn't a Fallacy (like sports or gambling, for example). With frequentism came many long-unexamined false assumptions; for example, that "noise" and "randomness" are "theory-free" concepts. In fact, few things are more loaded down with theoretical assumptions than "randomness," if taken as a metaphysical category. Taking it as a statement about the limits of human knowledge, OTOH, makes it almost a truism. In most cases, the Ludic Fallacy will come back to bite us: we often don't know all of the rules of the game.

Unfortunately, the frequentist approach to statistics is still taught because it's cookbook. Even in situations where a canned approach is not appropriate, a recipe feels comforting, relieving people of having to think. I might even call this the Cookbook Fallacy: having a wrong recipe is better than no recipe. Actually, no recipe is better than a bad one - at least it's honest and doesn't force us into wrong assumptions.

Taleb in his garden. Along with his skeptical empiricism, Taleb exhibits other exquisitely refined scientific tastes, paralleling his capacious gourmand tastes in literature and food. This might seem an affectation, but it points to an important truth.

Richard Feyman, another man of powerful scientific intuition, said: to do good science, you gotta have taste! Science, like the arts, has its forms of kitsch: rules mechanically applied without the imagination and drive for the fully worked-out development, but without indulging in useless repetition. Science is still and will always remain partly an art. To have taste is to avoid weak arguments and rationalizing, and to avoid applying methods and concepts where and when they are not valid. It is to think that, if there's no deep fundamental principle that prevents something, then why not? What would the world look like if it were so? Maybe you lack the imagination to see that that is our world. Taste is seeing that not taking obvious things for granted is a true royal road to discovery. It is paying attention to the silent evidence, to the dog that didn't bark, and to the pious who prayed and drowned anyway: survivor bias.

Taste in science also requires revisiting fundamental issues, ones never completely resolved. The progress of science has solved many problems defined more narrowly. But deep issues remain, even if transformed. Science has its classics and its literature, history, and philosophy; progress doesn't erase their importance. Read them and avoid being a cultural philistine.†

Taleb reminds us that what we don't know can hurt us, and that what we don't know is often more important than what we do. The Black Swan is a fine book. Buy, read, and enjoy it, patiently and slowly. And if nothing else, be charmed by the bittersweet tale of Yevgenia and her unknown masterpiece.
---
* Hayek attacked much the same in his Counterrevolution of Science, laying out the 19th-century origins of Platonified pseudo-knowledge in the social sciences and the pretensions of social planning that often went with it. Plenty of perceptive people, like our old friend Poincaré, resisted this development, this misapplication of inappropriate mathematical methods to society. But the Tyranny of the Cookbook is an unrelenting one.

** Not to be confused with Wackistan. That's where Ahmadinejad lives.

† Taleb uses the German term, Bildungsphilister, just to show, I suppose, that he isn't one.

Be alert to a real affectation, indulging in philosophical problems isolated from anything real. As Taleb points out, most philosophical issues worth bothering with are suggested by something outside philosophy.

Labels: , , , , , , ,