Saturday, July 22, 2017

Yep. Tron.

It was winter when I started writing this, but writing posts about physics is hard, at least if you're not a physicist.  This one was particularly hard because I had to re-learn what I thought I knew about the topic, and then realize that I'd never really understood it as well as I'd thought, then try to learn it correctly, then realize that I also needed to re-learn some of the prerequisites, which led to a whole other post ... but just for the sake of illustration, let's pretend it's still winter.

If you live near a modest-sized pond or lake, you might (depending on the weather) see it freeze over at night and thaw during the day.  Thermodynamically this can be described in terms of energy (specifically heat) and entropy.  At night, the water is giving off heat into the surrounding environment and losing entropy (while its temperature stays right at freezing).  The surrounding environment is taking on heat and gaining entropy.  The surroundings gain at least as much entropy as the pond loses, and ultimately the Earth will radiate just that bit more heat into space.  When you do all the accounting, the entropy of the universe increases by just a tiny bit, relatively speaking.

During the day, the process reverses.  The water takes on heat and gains entropy (while its temperature still stays right at freezing).  The surroundings give off heat, which ultimately came from the sun, and lose entropy.  The water gains at least as much entropy as the surroundings lose*, and again the entropy of the universe goes up by just that little, tiny bit, relatively speaking.

So what is this entropy of which we speak?  Originally entropy was defined in terms of heat and temperature.  One of the major achievements of modern physics was to reformulate entropy in a more powerful and elegant form, revealing deep and interesting connections, thereby leading to both enlightenment and confusion.  The connections were deep enough that Claude Shannon, in his founding work on information theory, defined a similar concept with the same name, leading to even more enlightenment and confusion.

The original thermodynamic definition relies on the distinction between heat and temperature.  Temperature, at least in the situations we'll be discussing here, is a measure of how energetic individual particles -- typically atoms or molecules -- are on average.  Heat is a form of energy, independent of how many particles are involved.

The air in an oven heated to 500K (that is, 500 Kelvin, about 227 degrees Celsius or 440 degrees Fahrenheit) and a pot full of oil at 500K are, of course, at the same temperature, but you can safely put your hand in the oven for a bit.  The oil, not so much.  Why?  Mainly because there's a lot more heat in the oil than in the air.  By definition the molecules in the oven air are just as energetic, on average, as a the molecules the oil, but there are a lot more molecules of oil, and therefore a lot more energy, which is to say heat.

At least, that's the quick explanation for purposes of illustration.  Going into the real details doesn't change the basic point: heat is different from temperature and changing the temperature of something requires transferring energy (heat) to or from it.  As in the case of the pond freezing and melting, there are also cases where you can transfer heat to or from something without changing its temperature.  This will be important in what follows.

Entropy was originally defined as part of understanding the Carnot cycle, which describes the ideal heat-driven engine (the efficiency of a real engine is usually given as a percentage of what the Carnot cycle would produce, not as a percentage of the energy it uses).  Among the principal results in classical thermodynamics is that the Carnot cycle was as good as you can get even in principle, but not even it can ever be perfectly efficient, even in principle.

At this point it might be helpful to read that earlier post on energy, if you haven't already.  Particularly relevant parts here are that the state of the working fluid in a heat engine, such as the steam in a steam engine, can be described with two parameters, or, equivalently, as a point in a two-dimensional diagram, and that the cycle an engine goes through can be described by a path in that two-dimensional space.

Also keep in mind the ideal gas law: In an ideal gas, the temperature of a given amount of gas is proportional to pressure times volume.  Here and in the rest of this post, "gas" means "a substance without a fixed shape or volume" and not what people call "gasoline" or "petrol".

If you've ever noticed a bicycle pump heat up as you pump up a tire, that's (more or less) why.  You're compressing air, that is, decreasing its volume, so (unless the pump is able to spill heat with perfect efficiency, which it isn't) the temperature has to go up.  For the same reason the air coming out of a can of compressed air is dangerously cold.  The air is expanding rapidly so the temperature drops sharply.

In the Carnot cycle you first supply heat a to gas (the "working fluid", for example steam in a steam engine) while maintaining a perfectly constant temperature by expanding the container it's in.  You're heating that gas, in the sense of supplying heat, but not in the sense of raising its temperature.  Again, heat and temperature are two different things.

To continue the Carnot cycle, let the container keep expanding, but now in such a way that it neither gains nor loses heat (in technical terms, adiabatically).  In these first two steps, you're getting work out of the engine (for example, by connecting a rod to the moving part of a piston and attaching the other part of that rod to a wheel).  The gas is losing energy since it's doing work on the piston, and it's also expanding, so the temperature and pressure are both dropping, but no heat is leaving the container in the adiabatic step.

Work is force times distance, and force in this case is pressure times the area of the surface that's moving.    Since the pressure, and therefore the force, is dropping during the second step you'll need to use calculus to figure out the exact amount of work, but people know how to do that.

The last two steps of the cycle reverse the first two.  In step three you compress the gas, for example by changing the direction the piston is moving, while keeping the temperature the same.  This means the gas is cooling in the sense of giving off heat, but not in the sense of dropping in temperature.  Finally, in step four, compress the gas further, without letting it give off heat.  This raises the temperature.  The piston is doing work on the gas and the volume is decreasing.  In a perfect Carnot cycle the gas ends up in the same state -- same pressure, temperature and volume -- as it began and you can start it all over.

As mentioned in the previous post, you end up putting more heat in at the start then you end up getting back in the third step, and you end up getting more work out in the first two steps than you put in in the last two (because the pressure is higher in the first two steps).  Heat gets converted to work (or if you run the whole thing backwards, you end up with a refrigerator).

If you plot the Carnot cycle on a diagram of pressure versus volume, or the other two combinations of pressure, volume and temperature, you get a a shape with at least two curved sides, and it's hard to tell whether you could do better.  Carnot proved that this cycle is the best you can do, in terms of how much work you can get out of a given amount of heat, by choosing two parameters that make the cycle into a rectangle.  One is temperature -- steps one and three maintain a constant temperature.

The other needs to make the other two steps straight lines.  To make this work out, the second quantity has to remain constant while the temperature is changing, and change when temperature is constant.  The solution is to define a quantity -- call it entropy -- that changes, when temperature is constant, by the amount of heat transferred, divided by that temperature (ΔS = ΔQ/T -- the deltas (Δ) say that we're relating changes in heat and entropy, not absolute quantities; Q stands for heat and S stands for entropy, because reasons).  When there's no heat transferred, entropy doesn't change.  In step one, temperature is constant and entropy increases.  In step two, temperature decreases while entropy remains constant, and so forth.

To be clear, entropy and temperature can, in general, both change at the same time.  For example, if you heat a gas at constant volume, then pressure, temperature and entropy all go up.  The Carnot cycle is a special case where only one changes at a time.

Knowing the definition of entropy, you can convert, say, a pressure/volume diagram to a temperature/entropy diagram and back.  In real systems, the temperature/entropy version won't show absolutely straight vertical and horizontal lines -- that is, there will be at least some places where both change at the same time.  The Carnot cycle is exactly the case where the lines are perfectly horizontal and vertical.

This definition of entropy in terms of heat and temperature says nothing at all about what's going on in the gas, but it's enough, along with some math I won't go into here (but which depends on the cycle being a rectangle), to prove Carnot's result: The portion of heat wasted in a Carnot cycle is the ratio of the cold temperature to the hot temperature (on an absolute temperature scale).  You can only have zero loss -- 100% efficiency -- if the cold temperature is absolute zero.  Which it won't be.

Any cycle that deviates from a perfect rectangle will be less efficient yet.  In real life this is inevitable.  You can come pretty close on all the steps, but not perfectly close.  In real life you don't have an ideal gas, you can't magically switch from being able to put heat into the gas to perfectly insulating it, you won't be able to transfer all the heat from your heat source to the gas, you won't be able to capture all the heat from the third step of the cycle to reuse in the first step of the next cycle, some of the energy of the moving piston will be lost to friction (that is, dissipated into the surroundings as heat) and so on.

The problem-solving that goes into minimizing inefficiencies in real engines is why engineering came to be called engineering and why the hallmark of engineering is getting usefulness out of imperfection.



There are other cases where heat is transferred at a constant temperature, and we can define entropy in the same way as for a gas.  For example, temperature doesn't change during a phase change such as melting or freezing.  As our pond melts and freezes, the temperature stays right at freezing until the pond completely freezes, at which point it can get cooler, or melts entirely, at which point it can get warmer.

If all you know is that some water is at the freezing point, you can't say how much heat it will take to raise the temperature above freezing without knowing how much of it is frozen and how much is liquid.  The concept of entropy is perfectly valid here -- it relates directly to how much of the pond is liquid -- and we can define "entropy of fusion" to account for phase transitions.

There are plenty of other cases that don't look quite so much like the ideal gas case but still involve changes of entropy.  Mixing two substances increases overall entropy.  Entropy is a determining factor in whether a chemical reaction will go forward or backward and in ice melting when you throw salt on it.


Before I go any further about thermodynamic entropy, let me throw in that Claude Shannon's definition of entropy in information theory is, informally, a measure of the number of distinct messages that could have been transmitted in a particular situation.  On the other blog, for example, I've ranted about bits of entropy for passwords.  This is exactly a measure of how many possible passwords there are in a given scheme for picking passwords.

What in the world does this have to do with transferring heat at a constant temperature?  Good question.

Just as the concept of energy underwent several shifts in understanding on the way to its current formulation, so did entropy.  The first major shift came with the development of statistical mechanics.  Here "mechanics" refers to the behavior of physical objects, and "statistical" means you've got enough of them that you're only concerned about their overall behavior.

Statistical mechanics models an ideal gas as a collection of particles bouncing around in a container.  You can think of this as a bunch of tiny balls bouncing around in a box, but there's a key difference from what you might expect from that image.  In an ideal gas, all the collisions are perfectly elastic, meaning that the energy of motion (called kinetic energy) remains the same before and after.  In a real box full of balls, the kinetic energy of the balls gets converted to heat as the balls bump into each other and push each other's molecules around, and sooner or later the balls stop bouncing.

But the whole point of the statistical view of thermodynamics is that heat is just the kinetic energy of the particles the system is made up of.  When actual bouncing balls lose energy to heat, that means that the kinetic energy of the large-scale motion of the balls themselves is getting converted into kinetic energy of the small-scale motion of the molecules the balls are made of, and of the air in the box, and of the walls of the box, and eventually the surroundings.  That is, the large scale motion we can see is getting converted into a lot of small-scale motion that we can't, which we call heat.

When two particles, say two oxygen molecules, bounce off each other, the kinetic energy of the moving particles just gets converted into kinetic energy of differently-moving particles, and that's it.  In the original formulation of statistical mechanics, there's simply no other place for that energy to go, no smaller-scale moving parts to transfer energy to (assuming there's no chemical reaction between the two -- if you prefer, put pure helium in the box).

When a particle bounces off the wall of the container, it imparts a small impulse -- an instantaneous force -- to the walls.  When a whole lot of particles continually bounce off the walls of a container, those instantaneous forces add up to (for all practical purposes) a continuous force, that is, pressure.

Temperature is the average kinetic energy of the particles and volume is, well, volume.  That gives us our basic parameters of temperature, pressure and volume.

But what is entropy, in this view?  In statistical mechanics, we're concerned about the large-scale (macroscopic) state of the system, but there are many different small-scale (microscopic) states that could give the same macroscopic picture.

Once you crank through all the math, it turns out that entropy is a measure of how many different microscopic states, which we can't measure, are consistent with the macroscopic state, which we can measure.  In fuller detail, entropy is actually proportional to the logarithm of that number -- the number of digits, more or less -- both because the raw numbers are ridiculously big, and because that way the entropy of two separate systems is the sum of the entropy of the individual systems.

The actual formula is S = k ln(W), where k is Boltzmann's constant and W is the total number of possible microstates, assuming they're all equally probable.  There's a slightly bigger formula if they're not.  Note that, unlike the original thermodynamic definition, this formula deals in absolute quantities, not changes.

When ice melts, entropy increases.  Water molecules in ice are confined to fixed positions in a crystal.  We may not know the exact energy of each individual molecule, but we at least know more or less where it is, and we know that if the energy of such a molecule is too high, it will leave the crystal (if this happens on a large scale, the crystal melts).  Once it does, we know much less about its location or energy.

Even without a phase change, the same sort of reasoning applies.  As temperature -- the average energy of each particle -- increases, the range of energies each particle can have increases.  How to translate this continuous range of energies into a number we can count is a bit of a puzzle, but we can handwave around that for now.

Entropy is often called a measure of disorder, but more accurately it's a measure of uncertainty (as theoretical physicist Sabine Hossenfelder puts it: "a measure for unresolved microscopic details"), that is, how much we don't know.  That's why Shannon used the same term in information theory.  The entropy of a message measures how much we don't know about it just from knowing its size (and a couple of other macroscopic parameters).  Shannon entropy is also logarithmic, for the same reasons that thermodynamic entropy is.

The formula for Shannon entropy in the case that all possible messages are equally probable is H = k ln(M), where M is the number of messages.  I put k there to account for the logarithm usually being base 2 and because it emphasizes the similarity to the other definition.  Again, there's a slightly bigger formula if the various messages aren't all equally probable, and it too looks an awful lot like the corresponding formula for thermodynamic entropy.

The original formulation of statistical mechanics assumed that physics at the microscopic scale followed Newton's laws of motion.  One indication that statistical mechanics was on to something is that when quantum mechanics completely reformulated what physics looks like at the microscopic scale, the statistical formulation not only held up, but became more accurate with the new information available.

In our current understanding, when two oxygen molecules bounce off each other, their electron shells interact (there's more going on, but let's start there), and eventually their energy gets redistributed into a new configuration.  This can mean the molecules traveling off in new paths, but it could also mean that some of the kinetic energy gets transferred to the electrons themselves, or some of the electrons' energy gets converted into kinetic energy.

Macroscopically this all looks the same as the old model, if you have huge numbers of molecules, but in the quantum formulation we have a more precise picture of entropy.  This makes a difference in extreme situations such as extremely cold crystals.  Since energy is quantized, there is a finite (though mind-bendingly huge) number of possible quantum states a typical system can have, and we can stop handwaving about how to handle ranges of possible energy.  This all works whether you have a gas, a liquid, an ordinary solid or some weird Bose-Einstein condensate.  Entropy measures that number of possible quantum states.

Thermodynamic entropy and information theoretic entropy are measuring basically the same thing, namely the number of specific possibilities consistent with what we know in general.  In fact, the modern definition of thermodynamic entropy specifically starts with a raw number of possible states and includes a constant factor to convert from the raw number to the units (energy over temperature) of classical thermodynamics.

This makes the two notions of entropy look even more alike -- they're both based on a count of possibilities, but with different scaling factors.  Below I'll even talk, loosely, of "bits worth of thermodynamic entropy" meaning the number of bits in the binary number for the number of possible quantum states.

Nonetheless, they're not at all the same thing in practice.

Consider a molecule of DNA.  There are dozens of atoms, and hundreds of subatomic particles, in a base pair.  I really don't know how many possible states a phosphorous atom (say) could be in under typical conditions, but I'm going to guess that there are thousands of bits worth of entropy in a base pair at room temperature.  Even if each individual particle can only be on one of two possible states, you've still got hundreds of bits.

From an information-theoretic point of view, there are four possible states for a base pair, which is two bits, and because the genetic code actually includes a fair bit of redundancy in the form of different ways of coding the same amino acid and so forth, it's actually more like 10/6 of a bit, even without taking into account other sources of redundancy.

But there is a lot of redundancy in your genome, as far as we can tell, in the form of duplicated genes and stretches of DNA that might or might not do anything.  All in all, there is about a gigabyte worth of base pairs in a human genome, but the actual gene-coding information can compress down to a few megabytes.  The thermodynamic entropy of the molecule that encodes those megabytes is much, much, larger.  If each base pair represents about a thousand bits worth of thermodynamic entropy under typical conditions, then the whole strand is into the hundreds of gigabytes.

I keep saying "under typical conditions" because thermodynamic entropy, being thermodynamic, depends on temperature.  If you have a fever, your body, including your DNA molecules in particular, has higher entropy than if you're sitting in an ice bath.  The information theoretic entropy, on the other hand, doesn't change.

But all this is dwarfed by another factor.  You have billions of cells in your body (and trillions of bacterial cells that don't have your DNA, but never mind that).  From a thermodynamic standpoint, each of those cells -- its DNA, its RNA, its proteins, lipids, water and so forth -- contributes to the overall entropy of your body.  A billion identical strands of DNA at a given temperature have the same information content as a single strand but a billion times the thermodynamic entropy.

If you want to compare bits to bits, the Shannon entropy of your DNA is inconsequential compared to the thermodynamic entropy of your body.  Even the change in the thermodynamic entropy of your body as you breathe is enormously bigger than the Shannon entropy of your DNA.

I mention all this because from time to time you'll see statements about genetics and the second law of thermodynamics.  The second law, which is very well established, states that the entropy of a closed system cannot decrease over time.  One implication of it is that heat doesn't flow from cold to hot, which is a key assumption in Carnot's proof.

Sometimes the second law is taken to mean that genomes can't get "more complex" over time, since that would violate the second law.  The usual response to this is that living cells aren't closed systems and therefore the second law doesn't apply.  That's perfectly valid.  However, I think a better answer is that this confuses two forms of entropy -- thermodynamic entropy and Shannon entropy -- which are just plain different.  In other words, thermodynamic entropy and the second law don't work that way.

From an information point of view, the entropy of a genome is just how many bits it encodes once you compress out any redundancy.  Longer genomes typically have more entropy.  From a thermodynamic point of view, at a given temperature, more of the same substance has higher entropy than less as well, but we're measuring different quantities.

A live elephant has much, much higher entropy than a live mouse, and likewise for a live human versus a live mouse.  As it happens, a mouse genome is roughly the same size as a human genome, even though there's a huge difference in thermodynamic entropy between a live human and a live mouse.  The mouse genome is slightly smaller than ours, but not a lot.  There's no reason it couldn't be larger, and certainly no thermodynamic reason.  Neither the mouse nor human genome is particularly large.  Several organisms have genomes dozens of times larger, at least in terms of raw base pairs.

From a thermodynamic point of view, it hardly matters what exact content a DNA molecule has.  There are some minor differences in thermodynamic behavior among the particular base pairs, and in some contexts it makes a slight difference what order they're arranged in, but overall the gene-copying machinery works the same whether the DNA is encoding a human digestive protein or nothing at all.  Differences in gene content are dwarfed by the thermodynamic entropy change of turning one strand of DNA and a supply of loose nucleotides into two strands, that in turn is dwarfed by everything else going on in the cell, and that in turn is dwarfed by the jump from one cell to billions.

For what it's worth, content makes even less thermodynamic difference in other forms of storage.  A RAM chip full of random numbers has essentially the same thermodynamic entropy, at a given temperature, as one containing all zeroes or all ones, even though those have drastically different Shannon entropies.  The thermodynamic entropy changes involved in writing a single bit to memory are going to equate to a lot more than one bit.

Again, this is all assuming it's valid to compare the two forms of entropy at all, based on their both being measures of uncertainty about what exact state a system is in, and again, the two are not actually comparable, even though they're similar in form.  Comparing the two is like trying to compare a football score to a basketball score on the basis that they're both counting the number of times the teams involved have scored goals.


There's a lot more to talk about here, for example the relation between symmetry and disorder (more disorder means more symmetry, which was not what I thought until I sat down to think about it), and the relationship between entropy and time (for example, as experimental physicist Richard Muller points out, local entropy decreases all the time without time appearing to flow backward), but for now I think I've hit the main points:
  • The second law of thermodynamics is just that -- a law of thermodynamics
  • Thermodynamic entropy as currently defined and information-theoretic (Shannon) entropy are two distinct concepts, even though they're very similar in form and derivation.
  • The two are defined in different contexts and behave entirely differently, despite what we might think from them having the same name.
  • Back at the first point, the second law of thermodynamics says almost nothing about Shannon entropy, even though you can, if you like, use the same terminology in counting quantum states.
  • All this has even less to do with genetics.

* Strictly speaking, you need to take the Sun into account.  The Sun is gaining entropy over time, at a much, much higher rate than our little pond and its surroundings, and it's only an insignificantly tiny part of the universe.  But even if you had a closed system, of a pond and surroundings that were sometimes warm and sometimes cold, for whatever reason, the result would be the same: The entropy of a closed system increases over time.

Wednesday, July 19, 2017

The human perspective and its limits

A couple more points occurred to me after I hit "publish" on the previous post.  Both of them revolve around subjectivity versus objectivity, and to what extent we might be limited by our human perspective.


In trying to define whether a kind of behavior is simple or complex, I gave two different notions which I claimed were equivalent: how hard it is to describe and how hard it is to build something to copy it.

The first is, in a sense, subjective, because it involves our ability to describe and understand things.  Since we describe things using language, it's tied to what fits well with language.  The second is much more objective.  If I build a chess-playing robot, something with no knowledge of human language or of chess could figure out what it was doing, at least in principle.

One of the most fundamental results in computer science is that there are a number of very simple computing models (stack machines, lambda calculus, combinators, Turing machines, cellular automata, C++ templates ... OK, maybe not always so simple) which are "functionally complete".  That means that any of them can compute any "total recursive function". This covers a wide range of problems, from adding numbers to playing chess to finding cute cat videos and beyond.

It doesn't matter which model you choose.  Any of them can be used to simulate any of the others.  Even a quantum computer is still computing the same kinds of functions [um ... not 100% sure about that ... should run that down some day --D.H.].  The fuss there is about the possibility that a quantum computer could compute certain difficult functions exponentially faster than a non-quantum computer.

Defining a totally recursive function for a problem basically means translating it into mathematical terms, in other words, describing it objectively.  Computability theory says that if you can do that, you can write a program to compute it, essentially building something to perform the task (generally you tell a general-purpose computer to execute the code you wrote, but if you really want to you can build a physical circuit to do the what the computer would do).

So the two notions, of describing a task clearly and producing something to perform it are, provably, equivalent.  There are some technical issues with the notion of complexity here that I'm going to gloss over.  The whole P = NP thing revolves around whether being able to state a problem simply means being able to solve it simply, but when it comes to deciding whether recognizing faces is harder than walking, I'm going to claim we can leave that aside.

The catch here is that my notion of objectivity -- defining a computable function -- is ultimately based on mathematics, which in turn is based on our notion of what it means to prove something (the links between computing and theorem proving are interesting and deep, but we're already in deep enough as it is).  Proof, in turn, is -- at least historically -- based on how our minds work, and in particular how language works.  Which is what I called "subjective" at the top.

So, is our notion of how hard something is to do mechanically -- my ostensibly "objective" definition -- limited by our modes of reasoning, particularly verbal reasoning, or is verbal/mathematical reasoning a fundamentally powerful way of describing things that we happened to discover because we developed minds capable of apprehending it?  I'd tend to think the latter, but then maybe that's just a human bias.



Second, as to our tendency to think that particularly human things like language and house-building are special, that might not just be arrogance, even if we're not really as special as we'd like to think.  We have a theory of mind, and not just of human minds.  We attribute very human-like motivations to other animals, and I'd argue that in many, maybe most, cases we're right.  Moreover, we also attribute different levels of consciousness to different things (things includes machines, which we also anthropomorphize).

There's a big asymmetry there: we actually experience our own consciousness, and we assume other people share the same level of consciousness, at least under normal circumstances, and we have that confirmed as we communicate our consciousnesses to each other.  It's entirely natural, then, to see our own intelligence and consciousness, which we see from the inside in the case of ourselves and close up in the case of other people, as particularly richer and more vivid.  This is difficult to let go of when trying to study other kinds of mind, but it seems to me it's essential at least to try.

Monday, July 17, 2017

Is recognizing faces all that special?

I've seen some headlines recently saying that fish can be taught to recognize human faces.  It's not clear why these would be circulating now, since the original paper appeared in 2016, but it's supposed to be newsworthy because fish weren't thought to have the neural structures needed to recognize faces.  In particular, they lack a neocortex (particularly the fusiform gyrus), or anything clearly analogous to it, which is what humans and primates use in recognizing faces.  Neither do the fish in question normally interact with humans, unlike, say, dogs, which might be expected to have developed an innate ability to recognize people.

The main thesis of the paper appears to be that there's nothing particularly special about recognizing faces.  As a compugeek, I'd say that the human brain is optimized for recognizing faces, but that doesn't mean that a more general approach can't work.  It makes sense that we'd have special machinery for faces.  Recognizing human faces is important to humans, though it's worth pointing out that there are plenty of people who don't seem to have this optimization (the technical term is prosopagnosia).

The authors of the paper also point out that recognizing faces is tricky:
[F]aces share the same basic components and individuals must be discriminated based on subtle differences in features or spatial relationships.
To be sure that the fish are performing the same recognition task we do, though presumably through different means, the experimental setup uses the same skin tone in all the images and crops them to a uniform oval.  Frankly, I found it hard to pick out the differences in what was left, but my facial recognition seems to be weaker than average in real life as well.

This is interesting work and the methodology seems solid, but should we really be surprised?  Yes, recognizing faces is tricky, but so is picking out a potential predator or prey, particularly if it's trying not to be found.

The archerfish used in the experiments normally provide for themselves by spitting jets of water at flies and small animals, then collecting them when they fall.  This means seeing the prey through the distortion of the air/water boundary, contracting various muscles at just the right rate and time, and finding the fallen prey.  For bonus points, don't waste energy shooting down dead leaves and such.

Doing all that requires the type of neural computation that seems easy until you actually try to duplicate it.  Did I mention that archerfish have a range on the order of meters, a dozen or so times their body length? It's not clear why recognizing faces should be particularly hard by comparison.

Computer neural networks can recognize faces using far fewer neurons than a fish has (Wikipedia says an adult zebrafish has around 10 million).  Granted, the fish has other things it needs to do with those neurons, and you can't necessarily compare virtual neurons directly to real ones, but virtual neurons are pretty simple -- they basically add a bunch of numbers, each multiplied by a "weight", and fiddle the result slightly.  Real neurons do much the same thing with electrical signals, hence the name "neural network".

It doesn't seem like recognizing shapes as complex as human faces should require a huge number of neurons.  The question, rather, is what kinds of brains are flexible enough to repurpose their shape recognition to an arbitrary task like figuring out which image of a face to spit at in order to get a tasty treat.

Again, is it surprising that a variety of different brains should have that kind of flexibility?  Being able to recognize new types of shape in the wild has pretty clear adaptive value, as does having flexible brain wiring in general.  Arguably the surprise would be finding an animal that relies strongly on its visual system that couldn't learn to recognize subtle differences in arbitrary shapes.

And yet, this kind of result does seem counterintuitive to many, and I'd include myself if I hadn't already seen similar results.  Intuitively we believe that some things take a more powerful kind of intelligence than others.  Playing chess or computing the derivative of a function is hard.  Walking is easy.

We also have a natural understanding of what kinds of intelligence are particularly human.  We naturally want to draw a clear line between our sort of intelligence and everyone else's.  Clearly those uniquely human abilities must require some higher form of intelligence.  Language with features like pronouns, tenses and subordinate clauses seems unique to us (though there's a lot we don't know about communication in other species), so it must be very high level.  Likewise for whatever we want to call the kind of planning and coordination needed to, say, build a house.

Recognizing each other's faces is a very human thing to do -- notwithstanding that several other kinds of animal seem perfectly capable of it -- so it must require some higher level of intelligence as well.

Now, to be clear, I'm quite sure that there is a constellation of features that, taken together, is unique and mostly universal to humanity, even if we share a number of particular features in that constellation with other species.  No one else we're aware of produces the kind of artifacts we do ... jelly donuts, jet skis, jackhammers, jugs, jujubes, jazz ...  or forms quite the same kind of social structures, or any of a number of other things.

However, that doesn't mean that these things are particularly complex or special.  We're also much less hairy than other primates, but near-hairlessness isn't a complex trait.  Our feet (and much of the rest of our bodies) are specialized for standing up, but that doesn't seem particularly different from specializing to swing through trees, or gallop, or hop like a kangaroo, or whatever else.

Our intuitions about what kind of intelligence is complex, or "of a higher order" are just not very reliable.  Playing chess is not particularly complicated.  It just requires bashing out lots and lots of different possible moves.  Calculating derivatives from a general formula is easy.  Walking, on the other hand, is fiendishly hard.  Language is ... interesting ... but many of the features of language, particularly stringing together combinations of distinct elements in sequence, are quite simple.

What do I mean by "simple" here?  I mean one of two more or less equivalent things: How hard is it to describe accurately, and how hard is it to build something to perform the task.  In other words, how hard is it to objectively model something, in the sense that you'll get the same result no matter who or what is following the instructions.

This is not necessarily the same question as how complex a brain do you need in order to perform the task, but this is partly because brains have developed in response to particular features of their environment.  Playing chess or taking the derivative of a polynomial shouldn't take a lot of neurons in principle, but it's hard for us because we don't have any neurons hardwired for those tasks.  Instead we have to use the less-hardwired parts of our brain pull together pieces that originally arose for different purposes.

Recognizing faces seems like something that requires a modest amount of machinery of the type that most visually-oriented animals should have available, and probably available in a form that can be adapted to the task, even if recognizing human faces isn't something the animal would normally have to do.  Cataloging what sorts of animals do it well seems interesting and ultimately useful in helping us understand our own brains, but we shouldn't be surprised if that catalog turns out to be fairly large.


Sunday, July 16, 2017

Discovering energy

If you get an electricity bill, you're aware that energy is something that can be quantified, consumed, bought and sold.   It's something real, even if you can't touch it or see it.  You probably have a reasonable idea of what energy is, even if it's not a precisely scientific definition, and an idea of some of the things you can do with energy: move things around, heat them or cool them, produce light, transmit information and so forth.

When something's as everyday-familiar as energy it's easy to forget that it wasn't always this way, but in fact the concept of energy as a measurable quantity is only a couple of centuries old, and the closely related concept of work is even newer.

Energy is now commonly defined as the ability to do work, and work as a given force acting over a given distance.  For example, lifting a (metric) ton of something one meter in the Earth's surface gravity requires exerting approximately 9800 Newtons of force over that distance, or approximately 9800 Newton-meters of work altogether.  A Joule of energy is the ability to do one Newton-meter of work, so lifting one ton one meter requires approximately 9800 Joules of energy, plus whatever is lost to inefficiency.

As always, there's quite a bit more going on if you start looking closely.  For one thing, the modern physical concept of energy is more subtle than the common definition, and for another energy "lost" to inefficiency is only "lost" in the sense that it's now in a form (heat) that can't directly do useful work.  I'll get into some, but by no means all, of that detail later in this post and probably in others as well.

I'm not going to try to give an exact history of thermodynamics or calorimetry here, but I do want to call out a few key developments in those fields.  My main aim is to trace the evolution of energy as a concept from a concrete, pragmatic working definition born out of the study of steam engines to the highly abstract concept that underpins the current theoretical understanding of the physical world.



The concept of energy as we know it dates to somewhere around the turn of the 19th century, that is, the late 1700s and early 1800s.   At that point practical steam engines had been around for several decades, though they only really took off when Watt's engine came along in 1781.  Around the same time a number of key experiments were done, heat was recognized as a form of energy and a full theory of heat, work and the relationship between the two was formulated.

What makes things hot?  This is one of those "why is the sky blue?" questions that quickly leads into deep questions that take decades to answer properly.  The short answer, of course, is "heat", but what exactly is that?  A perfectly natural answer, and one of the first to be formalized into something like what we would call a theory, is that heat is some sort of substance, albeit not one that we can see, or weigh, or any of a number of other things one might expect to do with a substance.

This straightforward answer makes sense at first blush.  If you set a cup of hot tea on a table, the tea will get cooler and the spot where it's sitting on the table will get warmer.  The air around the cup also gets warmer, though maybe not so obviously.  It's completely reasonable to say that heat is flowing from the hot teacup to its surroundings, and to this day "heat flow" is still an academic subject.

With a little more thought it seems reasonable to say that heat is somehow trapped in, say, a stick of wood, and that burning the wood releases that heat, or that the Sun is a vast reservoir of heat, some of which is flowing toward us, or any of a number of quite reasonable statements about heat considered as a substance.  This notional substance came to be known as caloric, from the Latin for heat.

As so often happens, though, this perfectly natural idea gets harder and harder to defend as you look more closely.  For example, if you carefully weigh a substance before and after burning it, as Lavoisier did in 1772, you'll find that it's actually heavier after burning.  If burning something releases the caloric in it, then does that mean that caloric has negative weight?  Or perhaps it's actually absorbing cold, and that's the real substance?

On the other hand, you can apparently create as much caloric as you want without changing the weight of anything.  In 1797 Benjamin Thompson, Count Rumford, immersed an unbored cannon in water, bored it with a specially dulled borer and observed that the water was boiling hot after about two and a half hours.  The metal bored from the cannon was not observably different from the remaining metal of the cannon, the total weight of the two together was the same as the original weight of the unbored cannon, and you could keep generating heat as long as you liked.  None of this could be easily explained in terms of heat as a substance.

Quite a while later, in the 1840s, James Joule did precise measurements how much heat was generated by a falling weight powering a stirrer in a vat of water.  Joule determined that heating a pound of water one degree Fahrenheit requires 778.24 foot-pounds of work (e.g., letting a 778.24 pound weight fall one foot, or a 77.824 weight fall ten feet, etc.). Ludwig Colding did similar research, and both Joule and Julius Robert von Mayer published the idea that heat and work can each be converted to the other.  This is not just significant theoretically.  Getting five digits of precision out of stirring water with a falling weight is pretty impressive in its own right.

At this point we're well into the development of thermodynamics, which Lord Kelvin eventually defined in 1854 as "the subject of the relation of heat to forces acting between contiguous parts of bodies, and the relation of heat to electrical agency."  This is a fairly broad definition, and the specific mention of electricity is interesting, but a significant portion of thermodynamics and its development as a discipline centers around the behavior of gasses, particularly steam.


In 1662, Robert Boyle published his finding that the volume of a gas, say, air in a piston, is inversely proportional to the pressure exerted on it.  It's not news, and wasn't at the time, that a gas takes up less space if you put it under pressure.  Not having a fixed volume is a defining property of a gas.  However, "inversely proportional" says more.   It says if you double the pressure on a gas, its volume shrinks by half, and so forth.  Another way to put this is that pressure multiplied by volume remains constant.

In the 1780s, Jacque Charles formulated (but didn't publish) the idea that the volume of a gas was proportional to its temperature.  In 1801 and 1802, John Dalton and Joseph Louis Guy-Lussac published experimental results showing the same effect.  There was one catch: you had to measure temperature on the right scale.  A gas at 100 degrees Fahrenheit doesn't have twice the volume of a gas at 50 degrees, nor does it if you measure in Celsius.

However, if you plot volume vs. temperature on either scale you get a straight line, and if you put the zero point of your temperature scale where that line would show zero volume -- absolute zero -- then the proportionality holds.  Absolute zero is quite cold, as one might expect.  It's around 273 degrees below zero Celsius (about 460 degrees below zero Fahrenheit).  It's also unobtainable, though recent experiments in condensed matter physics have come quite close.

Put those together and you get the combined gas law: Pressure times volume is proportional to temperature.

In 1811 Amedeo Avogadro hypothesized that two samples of the same gas at the same temperature, pressure and volume contained the same number of molecules.  This came to be known as Avogadro's Law.  The number of molecules in a typical system is quite large.  It is usually expressed in terms of Avogadro's Number, approximately six followed by twenty-three zeroes, one of the larger numbers that one sees in regular use in science.

Put that together with the combined gas law and you have the ideal gas law:
PV = nRT
P is the pressure, V is the volume, n is the number of molecules, T is the temperature and R is the gas constant that makes the numbers and units come out right.

The really important abstraction here is state.  If you know the parameters in the ideal gas law -- the pressure, volume, temperature and how many gas particles there are, then you know its state.  This is all you need to know, and all you can know, about that gas as far as thermodynamics is concerned.  Since the number of gas particles doesn't change in a closed system like a steam engine (or at least an idealized one), you only need to know pressure, volume and temperature to know the state.

Since the ideal gas law relates those, you really only need to keep track of two of the three.  If you measure pressure, volume and temperature once to start with, and you observe that the volume remains constant while the pressure increases by 10%, you know that the temperature must be 10% higher than it was.  If the volume had increased by 20% but the pressure had dropped by 10%, the temperature must now be 8% higher (1.2 * 0.9 = 1.08).  And so forth.

You don't have to track pressure and volume particularly, or even two of {pressure, volume, temperature}.  There are other measures that will do just as well (we'll get to one of the important ones in a later post), but no matter how you define your measures you'll need two of them to account for the thermodynamic state of a gas and (as long as they aren't essentially the same measure in disguise), two will be enough.  Technically, there are two degrees of freedom.

This means that you can trace the thermodynamic changes in a gas on a two-dimensional diagram called a phase diagram.  Let's pause for a second to take that in.  If you're studying a steam engine (or in general, a heat engine) that converts heat into work (or vice versa) you can reduce all the movements of all the machinery, all the heating and cooling, down to a path on a piece of paper.  That's a really significant simplification.


In theory, the steam in a steam engine (or generally the working fluid in a heat engine), will follow a cycle over and over, always returning to the same point in the phase diagram (that is, the same state).    In practice, the cycle won't repeat exactly, but it will still follow a path through the phase diagram that repeats the same basic pattern over and over, with minor variations.

The heat source heats the steam and the steam expands.  Expanding means exerting force against the walls of whatever container it's in, say the surface of a piston.  That is, it means doing work.  The steam is then cooled, removing heat from it, and the steam returns to its original pressure, volume and temperature.  At that point, from a thermodynamic point of view, that's all we know about the steam.  We can't know, just from taking measurements on the steam, how many times it's been heated and cooled, or anything else about its history or future.  All you know is its current thermodynamic state.

As the steam contracts back to its original volume, its surroundings are doing work on it.  The trick is to manipulate the pressure, temperature and volume in such a way that the pressure, and thus the force, is lower on the return stroke than the power stroke, and the steam does more work expanding than is done on it contracting.  Doing so, it turns out, will involve putting more heat into the heating than comes out in the cooling.  Heat goes in, work comes out.


This leads us to one of the most important principles in science.  If you carefully measure what happens in real heat engines, and the ways you can trace through a path in a phase diagram, you find that you can convert heat to work, and work to heat, and that you will always lose some waste heat to the surroundings, but when you add it all up (in suitable units and paying careful attention to the signs of the quantities involved), the total amount of heat transferred and work done never changes.  If you put in 100 Watts worth of heat, you won't get more than 100 Watts worth of work out.  In fact, you'll get less.  The difference will be wasted heating the surroundings.

This is significant enough when it comes to heat engines, but that's only the beginning.  Heat isn't the only thing you can convert into work and vice versa.  You can use electricity to move things, and moving things to make electricity.  Chemical reactions can produce or absorb heat or produce electrical currents, or be driven by them.   You can spin up a rotating flywheel and then, say, connect it to a generator, or to a winch.

Many fascinating experiments were done, and the picture became clearer and clearer: Heat, electricity, the motion of an object, the potential for a chemical reaction, the stretch in a spring, the potential of a mass raised to a given height, among other quantities, can all be converted to each other, and if you measure carefully, you always find the total amount to be the same.

This leads to the conclusion that all of these are different forms of the same thing -- energy -- and that this thing is conserved, that is, never created or destroyed, only converted to different forms.


As far-reaching and powerful as this concept is, there were two other important shifts to come.  One was to take conservation of energy not as an empirical result of measuring the behavior of steam engines and chemical reactions, but as a fundamental law of the universe itself, something that could be used to evaluate new theories that had no direct connection to thermodynamics.

If you have a theory of how galaxies form over millions of years, or how electrons behave in an atom, and it predicts that energy isn't conserved, you're probably not going to get far.  That doesn't mean that all the cool scientist kids will point and laugh (though a certain amount of that has been known to happen).  It means that sooner or later your theory will hit a snag you hadn't thought of and sooner or later the numbers won't match up with reality*.  When this happens over and over and over, people start talking about fundamental laws.


The second major shift in the understanding of energy came with the quantum theory, that great upsetter of scientific apple carts everywhere.  At a macroscopic scale, energy still behaves something like a substance, like the caloric originally used to explain heat transfer.  In Count Rumford's cannon-boring experiment, mechanical energy is being converted into heat energy.  Heat itself is not a substance, but one could imagine that energy is, just one that can change forms and lacks many of the qualities -- color, mass, shape, and so forth -- that one often associates with a substance.

In the quantum view, though, saying that energy is conserved doesn't assume some substance or pseudo-substance that's never created or destroyed.  Saying that energy is conserved is saying that the laws describing the universe are time-symmetric, meaning that they behave the same at all times.  This is a consequence of Noether's theorem (after the mathematician Emmy Noether), one of the deepest results in mathematical physics, which relates conservation in general to symmetries in the laws describing a system.  Time symmetry implies conservation of energy.  Directional symmetry -- the laws work the same no matter which way you point your x, y and axes -- implies conservation of angular momentum.

Both of these are very abstract.  In the quantum world you can't really speak of a particle rotating on an axis, yet you can measure something that behaves like angular momentum, and which is conserved just like the momentum of spinning things is in the macroscopic world.  Just the same, energy in the quantum world has more to do with the rates at which the mathematical functions describing particles vary over space and time, but because of how the laws are structured it's conserved and, once you follow through all the implications, energy as we experience it on our scale is as well.

This is all a long way from electricity bills and the engines that drove the industrial revolution, but the connections are all there.  Putting them together is one of the great stories in human thought.

* I suppose I can't avoid at least mentioning virtual particles here.  From an informal description, of particles being created and destroyed spontaneously, it would seem that they violate conservation of energy (considering matter as a form of energy).  They don't, though.  Exactly why they don't is beyond my depth and touches on deeper questions of just how one should interpret quantum physics, but one informal way of putting it is that virtual particles are never around for long enough to be detectable.  Heisenberg uncertainty is often mentioned as well.

Friday, May 26, 2017

The value of the thing ...

... is what it will bring.

I've now seen several headlines along the lines of "NASA to explore  $10,000 quadrillion metal asteroid"

What does this even mean?

Two things, really:
  • NASA is planning a mission to the nickel-iron asteroid 16 Psyche, which is true
  • That asteroid contains $10 quintillion worth of metal, which is, um ...
I mean, on the one hand it's a simple calculation: Psyche contains X tons of nickel at $Y/ton, and likewise for iron.  Total value: $10 quintillion or, for whatever reason, $10,000 quadrillion.

Except that's about 100,000 times the world's GDP, so maybe we're missing something?

Suppose we could magically bring all the nickel and iron in Psyche to earth.  That's a ball about 200km across, so we'd have to be a bit careful, but say we break it down into a few million 1km heaps distributed strategically around the world.  How much is that really worth?

You might think "Yay, free iron and nickel!" but that's not quite right.  Even scrap iron, which has already been refined and packaged into usable pieces, costs something to buy, something to transport and something more to put to use, unless it happens to be in just the form you need.  More realistically, it would mean no more iron mining, which is great unless you happen to be in the iron/nickel mining business.  That's not nothing -- world iron production looks to be around $300 billion and nickel maybe more like $20 billion.  But it's not a trillion dollars, much less a quadrillion or quintillion.

Or look at it another way: We've got a rock out in space that's worth as much as the entire world economy would produce in 100,000 years at current rates.  The total budget of NASA is around $20 billion, with ESA JAXA and the Russian space agency accounting for a few billion more.  Surely it would be worth it to throw the world's entire space budget into mining that rock.

Except, the question isn't whether there's a bunch of valuable metal to be mined. The question is whether it's worth mining.  It currently costs about $20,000 per kilogram to get a payload to low earth orbit.  It's anybody's guess what it would cost to actually mine a given amount of metal in the asteroid belt and bring it back to Earth safely -- though if you're transporting a hunk of metal I suppose you just have to make sure that it doesn't hit anything on the way in.  But bulk nickel from Earth runs more like $10/kg and iron is cheaper yet, so ... maybe not.


I don't really want to pick on NASA for trying to drum up a little interest in its latest mission -- though it's probably worth mentioning that the past couple of decades of unmanned missions by NASA and the other space agencies have been spectacularly successful in exploring the solar system and in an ideal world that would speak for itself.  If there's a point here, it's that it's a good idea not to take numbers, especially eye-catching dollar amounts, at face value without asking what they actually mean.

Thursday, April 6, 2017

Big vocabulary, or just big words?

The other day I was reading an article that used a couple of words I hadn't seen in a while, say anodyne or encomium.  I more-or-less remembered what they meant, and it was reasonably clear from context what they meant, but I still ended up looking them up.  I had two feelings about this: on the one hand, did the author really have to drag those out?  Why not just use Plain English?  On the other hand, they were correctly used, and apt, so what's the big deal?  I'm sure I've thrown out a word or two here that I could replaced with something more familiar, maybe with a little rewording.

But I'm not here to critique style.  What stuck in my head about this incident was how conspicuous an unusual word can be (and besides that, unusual words tend to stick out).  The article itself was probably a thousand words or so, maybe more, but it was those two that changed the whole reading experience.

This wasn't just because of the extra time it took to look the words up and make sure I knew what they meant.  That's a speed bump these days, reading an online article with search bar and dictionary app at the ready, maybe an extra minute, if that.  Even if I hadn't had a dictionary handy, I could have gotten the good out of the article without knowing exactly what those words meant.

The real issue lies deeper in human perception: We (and living things with recognizable brains in general) are finely tuned to notice discrepancies.  In a field of green grass it's the shape of that predator, or that prey,  or that particularly tasty plant, or whatever, that stands out.  In an article of a thousand words, it's the unusual ones that stand out.

I could go on and on, but it's worth particularly noting how important this is in social environments.  We can spot an unfamiliar accent in seconds.  We can spot someone dressed differently, or with different features than we usually see, well before we're even aware that we have.  The other night I was watching a TV show with a foreign actor playing an American, and everything was just fine until they said "not" with a British "o".   It didn't ruin the whole show -- this was a single vowel, not Dick Van Dyke in Mary Poppins -- but it was noticeable enough I still recall it out of an hour of tense drama.

(I have to say that dialect coaching has gotten a lot better over the past couple of decades.  Time was, movie stars talked like movie stars, with a kind of over-enunciated diction that didn't sound like anyone in real life, and if a character was meant to sound foreign, pretty much anything would do.   This is doubtless because in the early days of "talking pictures" the medium was still transitioning from the stage, a theatre actor was used to projecting up to the cheap seats and a fake accent was as good as a fake beard since everything was a hand-painted set and there probably weren't that many people in the audience who knew what a true Elbonian accent sounded like anyway.  Today pretty much every part of that is different, and we expect realism -- Billy Bob Thornton's all-too-valid complaint about "that Southern accent that no one in the South actually speaks with" notwithstanding.)

Where was I?

I've argued before that we often seem to care most about distinctions when they matter least. Vocabulary is largely another example of that.  Unless  you're reading Finnegan's Wake or something equally chewy, you're probably OK just skimming over anything you don't know and looking it up later.  Even that blowhard commentator with the two-dollar words is trying to get a point across and isn't going to let the vocabulary get completely in the way.

As a corollary to that, you don't need to know very many unusual words in order to stand out.  If you know a few dozen and use them appropriately, you'll almost certainly draw attention (if you learn a few dozen and use them inappropriately you'll also draw attention, but probably not the kind you want).  This can happen naturally if you run across a rare-ish but useful word or two in your reading from time to time and hold onto it for future use.  There's something nice about, say, cogent that is hard to reword cleanly, the distinction between terse and concise is sometimes worth making, and so on.

Contrast that with the average human vocabulary.  This is a hard thing to measure, but if you've heard something on the order of "uneducated people have a vocabulary of 2000 words while educated people know 20,000", rest assured that's complete bunk.  If we're measuring vocabulary, we have to measure "listemes", that is, things that you just have to learn by rote because you can't work them out from their parts.

This includes all kinds of things:
  • proper names of people and places
  • distinct senses of words, particularly small words like out and by, which can have quite a few, depending on how you count.
  • idioms large and small, like in touch or look up (in its non-literal senses) to classics like red herring, two-dollar word and that's the way the cookie crumbles.
  • Cultural references, which are kind of like names and kind of like idioms
  • Fine points that we don't generally think of as idioms, but are idiomatic nonetheless, like fried egg meaning a particular way of frying an egg, as distinct from scrambling an egg or -- for whatever reason -- trying to fry a whole egg in a pan without removing the shell
I'm not trying to give a full taxonomy of things-that-you-just-have-to-learn, but I hope that gives the general idea.   The main point is that there are lots and lots of these, the categories they might fall into are somewhat arbitrary, and how many you know doesn't have a great deal to do with how many literary classics you've read.

I'm not really familiar with the research on this, but my understanding is that the average person knows somewhere in the hundreds of thousands of listemes, and a large portion of them are commonly understood.  On top of those, we can add a smaller portion of jargon, slang or sesquipedalianisms.  That part, people will notice.  But it's a relatively small part.

Monday, March 20, 2017

Did Dory jump the shark?

I was fortunate enough to attend SIGGRAPH 86 and see the premier of Luxo Jr.   If you haven't seen it, I'd highly recommend you do.  It's only two minutes long.

Luxo Jr. was an eye-opener to me for a number of reasons.  First, and this may be hard to believe now, it was a technical milestone.  At the time, the field of computer graphics was in the process of moving from 3D wireframes like this to something more realistic, and Pixar did a lot of the heavy lifting in that move.

There were a number of problems to be solved at the time.  Some of them had to do with how to render an image of a mathematical model, for example:
  • How to draw exactly what should be visible (hidden-line and hidden-surface removal).  If your model has a cube, an image of that model should only show the faces nearer to you and not the ones on the back -- or anything that's covered by nearer objects in the scene.
  • How to show more realistic textures than just flat polygons.  At first blush you might think that, say, a house is just a few flat walls with windows cut out.  But those walls won't just be flat surfaces.  There might be brick, or siding.  Even a concrete or stucco surface will have little irregularities.  Drawing flat surfaces with uniform colors will convey the overall design, but it won't look like the real thing.
  • How to deal with atmospheric effects.  In real life, there might be smoke or mist in the air.  Even on a clear day distant objects will have more muted colors than nearby ones.
  • How to deal with shiny objects.  Even in the best case, the math for figuring out how bright a particular point on a surface should be is harder for shiny surfaces.  At worst, you have to deal with reflections of other objects, and reflections of reflections, and so on, something like this.
  • How to deal with transparent and translucent objects (which might also be shiny).  Again, this ranges from harder math for the shading to figuring out how the rest of the scene appears when distorted by a curved surface.
  • How to deal with shadows.  If one part of your model is between a light source and another part, that other part will, naturally enough, be darker.
  • A whole slew of subtler optical effects -- color bleeding, depth of field, motion blur, caustics and probably several others I don't remember.  I recall one presenter at a conference half-joking that the whole field had devolved into finding a new optical subtlety and writing a paper about how to render it.
Even if you knew how to render a model accurately, there were thorny questions about modeling:
  • Real scenes contain a whole lot of objects.  Look around next time you're outside -- or inside an average house, office, store or whatever.  A realistic rendering will have to account, somehow, for every blade of grass, every leaf, every feather of every bird, every rock on a gravel path, and so forth.  You don't necessarily have to create a separate object for every detail, but somehow you have to be prepared to render either a green grassy texture or blades of grass, depending on how closely you're looking.  Keep in mind that at that time a typical mobile phone of today would have seemed like a supercomputer (That may seem like hyperbole, but it's not.  The ubiquitous SPARCtation 2, for example, ran at 40MHz with 128MB of RAM)
  • Objects move.  In reality, they obey the laws of physics.  In an entertainment video, they might move in all sorts of non-physical ways, but anything that's supposed to look lifelike had better move more or less like a real-live thing.  Modeling the movement of a piece of clothing, or a full head of hair, or the surface of the ocean, or the flames in a fire, were each good for multiple published papers.
There were (and, I think still are for the most part) two approaches to problems like these:
  • Grinding out exact solutions to the optics (for rendering) and physics (for modeling)
  • Finding Stuff That Works.
At the time, ray-tracing was the state of the art for bashing out the optics, though that would soon be superseded by radiosity -- which had the distinction of being even slower than ray-tracing -- and more sophisticated numerical approaches.  Jim Kajiya laid out a general form for the problem to be solved and demoed an image that used Monte Carlo simulation to produce what he called "a great simulation of film grain" (see the end of this PDF of the paper).  It was a technical tour de force, solving a good chunk of the rendering problems above with one integral equation, using techniques that had been used to model the atomic bomb a generation earlier, among other things.  It was not, however, a very impressive demo unless you knew exactly what to look for.


Pixar took the other, entirely different approach*.  They handled hidden surfaces through what came to be known as "polygon pushing" -- reducing everything to a model with flat sides that was close enough to the real thing.  Flatter parts of surfaces could get by with fewer polygons than curvier parts.  You could then sort those polygons to see which was closest to the eye at any particular point.  Fast sorting algorithms had been around for decades.  Sorting in three dimensions is harder, but it's still possible to do it relatively quickly, even on what was fairly ordinary hardware.

They handled shadows through "shadow mapping", essentially calculating where shadows would fall on a surface and making that a property of the surface.  You could figure out where the shadows would fall by looking at the scene from the point of view of the light source, using the same sorting algorithm as for hidden surfaces.  You only had to re-do the shadow map when things moved, and much of a typical movie scene is background or otherwise not moving.

They handled textures with texture mapping and bump mapping, which treated the surfaces as flat but then modified the color or local orientation used in the actual shading calculations based on what exact part of the surface you were looking at.  That's how the wooden floor in Luxo Jr. was done.

They also developed algorithms for modeling the movement of the lamps and their cords, but I'm less familiar with that.  Overall they built up a library of rendering techniques, modeling techniques and models, some general-purpose and powerful, some specialized to particular tasks.  Just as important, they built a framework to plug it all into harmoniously.

Kajiya's paper was a great example of the scientific approach, and it ended up underpinning a chunk of important work.  It offered only an approximate solution, out of necessity, to the actual problem of putting pixels on the screen but it rigorously defined the exact problem to solve.

Pixar did engineering.  They figured out what mattered and what didn't for the purposes of producing an image that would fool the eye in an entertaining video -- basically which shortcuts people would and wouldn't notice -- and applied their resources to solving the problems that mattered.  They also developed software for managing a server farm doing the rendering and all kinds of tools to support the animators in making their magic.


I suppose I should take a moment to push back against a couple of stereotypes.  It's tempting to write off "the scientific approach" as "of no practical value" or the engineering approach as "just a bunch of hacks".  From what I can tell, though, it's hard to write a useful scientific paper in CS without knowing how to code, and it's hard to come up with a good practical hack without understanding what the full solution looks like.  Both have been done, but most people who've made a difference have a healthy dose of both practical and theoretical knowledge and tend to move back and forth on the deep insight/cheap hack scale as the occasion demands or the mood strikes.


But all this technical discussion leaves out what made, and makes, Pixar truly special.  The Pixar folks didn't just have formidable technical chops and great engineering sense.  They told stories.

This was a conscious decision from the outset.  John Lasseter and the rest of the team paid a lot of attention to the generation of animators before them, particularly the Disney studio.

If you're drawing every single frame of a picture by hand, even if you're using techniques like cel animation to re-use background drawings, you have to make every line count.  The people who we now call the "traditional animators" developed a set of techniques, for example squash and stretch, to illustrate motion without detailing every single movement.  They studied facial expressions in order to make their characters emote in a way we instinctively understand.  They watched how people and animals moved in order to capture the essence of lifelike motion.  They noticed that cute baby animals had (relatively) bigger heads than their adult counterparts, and made countless other observations that went into their work.

If you're just trying to figure out how to shade a model of a teapot by the conference submission deadline you probably won't pay much attention to these things, but the Pixar team did because their goal, from the beginning, was to tell stories with animation.  This is crystal clear from the very start.  The story in Luxo Jr. is pretty simple, but it's clearly a story, with characters with real emotions, even if those characters are metal desk lamps.  In fact, that's the magic: Inanimate, computer-generated desk lamps brought to life -- literally animated.

Watching it at the time was one of those "I didn't realize you could do that" moments, not so much from the technical point of view, though it's technically quite good as well, but because after antiseptic wireframe video games and shiny special effects and endless discussions of ray-tracing vs. polygon pushing it didn't seem like storytelling had much at all to do with the field.


My co-workers and I went to dinner at a steakhouse in Dallas afterward.  I remember talking about what portion of the real-life scene there could be modeled and rendered realistically with the resources available.  Having seen a few papers presented on techniques for rendering transparent objects with curved surfaces I claimed that the wine glasses could be handled OK (not a foregone conclusion at that point).  My boss dipped his thumb in steak juice and smudged it on the glass.  "Render that".  I muttered something about transparency mapping and such, and I might have been right, but the point was made.

With the tools we have these days, that smudge would be a minor obstacle.  Computer-generated scenes still often have that too-clean look to them, but that's more a matter of choice.  Computer imagery can handle grit and grime, but it's often easier to model without it.  If it makes sense for the setting or character, it's there, but otherwise it's usually not.  Also, I suspect, it's easier for an audience to make sense of a scene if the animated main characters look somewhat unnaturally clean and shiny while the trees off in the distance look realistic.


Which brings us to Finding Dory.

In my opinion it's not a bad film, but there's something missing.  Technically, it continues Pixar's upward trend in awesomeness.  The modeling for Hank the Septopus is so seamless you forget all about the huge amount of work that must have gone into it, from the motions of the tentacles to studying enough octopus behavior to make Hank a move like a realistic cephalopod, to knowing enough old-school animation technique to make him expressive within those parameters.  And there's plenty more where that came from.

There are a number of acceptable breaks from reality, starting with talking animals, and on to reading animals, truck-driving animals, aquatic animals spending unlikely amounts of time out of water, and even a plot-convenient echolocation ability that apparently doesn't use ultrasound and works through air as well as water -- not to mention navigating around bends in pipes while still conveying that there are bends at all.  That's all fine.  I mean, if you're OK with talking underwater animals, hard-boiled skepticism is pretty much out the window to start with.

The problem, unfortunately, is the storytelling.

I had to stop here for a bit, partly because, even if I'm a bit of a curmudgeon, I don't really relish the thought of criticizing Dory, Nemo and the gang.  Curmudgeons can still be fans.  Mostly though, I realized that if I wanted to go there, I should at least have a specific reason to go there, and it took me a little while to pinpoint that reason.

In Finding Nemo, one of the best moments, and probably the biggest emotional payoff, is when Dory, the cognitively impaired blue tang who at first seems to have been there for comic relief and to play the role of the wacky, plot-complicating sidekick, realizes "I look at you and…I’m home."**  The setting is as spare as can be, just two characters alone against a plain backdrop, one of them not even speaking, and that's what makes it work: the characters, and their slow realization of what's happened.

Dory didn't see it coming because, well, she's Dory and she had only fleeting hints that she was lost in the first place.  Marlin didn't see it coming because he'd been consumed by his quest to redeem his guilt and remorse over losing Nemo.  The audience didn't see it coming because the rest of the story was zipping along at Pixar's usual frenetic-but-impeccably-timed pace and keeping us engaged with a steady parade of engaging characters.  It also doesn't hurt that Ellen DeGeneres delivers the speech perfectly.

And that's where Finding Dory's trouble begins.  It's just going to be really hard to top a moment like that.  It's probably not a good idea to even try.  If Dory's backstory stays a backstory we can carry it with us however we like.  Probably better to leave that magic alone.  But at the same time, you can't blame Pixar for trying anyway.  "How are you going to top that?" has driven a lot of creative people to a lot of really good work.  I can't imagine there wasn't a little voice in the back of someone's head saying "Challenge accepted."

Meanwhile, rendering and modeling technology march on.  Realistic waves crashing on a beach?  We can do that.  Schools of fish circling in a cylindrical tank?  No problem.  Northern California vegetation in a light mist?  That's the morning commute.  How about some Toy Story-style kids wreaking havoc as they plunge their hands into the touch pool, kicking up clouds of sand?  Done.  It's not that Pixar has ever been shy about pushing the technical envelope.  It just seems a bit -- visible.  Technique is hardly ever meant to be visible.

And of course, the mouse must be fed.  When Dory dodged under Destiny the shark at the last second, I couldn't help thinking "That'll feature somewhere in a Disney ride".  And it's not hard to guess which characters were likely to make for hot-selling plush toys.  Nemo, Marlin, Crush the sea turtle and whoever else have to be there because sequel.

It's not that commercial tie-ins and franchise characters are bad per se.  Those server farms don't run themselves (well, at least not yet).  It's just that, like the technical mastery, the commercial machinery is not supposed to actually jump out at you.


In the end, Finding Dory's weakness boils down to fundamentals: the external constraints are muscling in on the plot, and the plot is driving the characters, when it should be the other way around.  In Luxo Jr., there's hardly any plot at all.  The whole point is to use the technology -- really just a bunch of crunching of a bunch of numbers describing colors, geometric shapes and such -- to show us believable characters.  Character wins, maybe not every single time, but almost always.  That's especially true if you're Pixar, which is why between Luxo Jr. and Finding Dory, that two-minute short is the better film of the two.

Is this the end of Pixar as we know it?  Is it all merchandising and sequels from here on out?  Well, three of the four upcoming Pixar projects with titles are sequels (Cars 3, The Incredibles 2 and Toy Story 4)  Let's hope that Lasseter's pledge that "If we have a great story, we'll do a sequel" holds.  I haven't seen Monsters University or Toy Story 3 (I think), but as I recall Pixar handled Toy Story 2 pretty deftly, sequel though it was.

Really it's impressive that they haven't stumbled any more than they have, all things considered.  But this one definitely feels like a stumble.



* I'm writing most of this from memory, so I'm only mostly confident it's mostly right.  Corrections are welcome.
**Just put Dory home in the search bar, 14 years after Finding Nemo came out, and there it is.

[I still see it on the first page of hits, but there's a lot of Finding Dory mixed in with it now.  Not sure what to make of that --D.H. Mar 2020[

Friday, March 10, 2017

Science on a shoestring

On the other blog I would occasionally put out short notices of neat hacks (as always, "hack" in the "solving problems ingeniously" sense).  I recently ran across one that didn't have much to do with the web, so I thought I'd carry that tradition over to this blog.


Muons are subatomic particles similar to electrons but much heavier.  They are generally produced in high-energy interactions in particle accelerators or from cosmic rays slamming into the atmosphere.  Muons at rest take about 2 microseconds to decay, actually a pretty long time for an unstable particle.  Muons from cosmic ray collisions are moving fast enough that they take measurably longer to decay (in our reference frame), which is one of the many pieces of supporting evidence for special relativity.

The GRAPES-3 detector at Ooty in Tamil Nadu, India detects just such decays using an array of detectors set into a hill 2200m (7200 ft) above sea level.  The detectors themselves are made largely from recycled materials, particularly square metal pipes formerly used in construction projects in Japan.  The total annual budget for the project is under $400,000, but the team has already produced significant results.  Auntie has more details on the construction of the instruments here.

There are a couple of narratives that are often spun around stories like this.  One is a sort of condescending "Isn't that cute?" with maybe a reference to the Professor on Gilligan's Island building a radio out of coconuts.  Another is "Look what people can do without huge budgets.  Why do we need all these multi-billion-dollar projects anyway?"

I'd rather not tell either of those.  What I see here is highly skilled scientists making use of the resources they have available to produce significant results.  Their counterparts at CERN or whatever are making use of different resources to produce different significant results.  Both are moving the ball forward.  There have been plenty of neat hacks at CERN, including something called "HTTP",  but today I wanted to call out GRAPES-3, mainly because it's just plain cool.

Friday, March 3, 2017

Reworking the Drake equation

In speculating about life on other worlds (here and here for example) the Drake Equation provides a useful framework.  This equation multiplies a number of factors to arrive at the number of civilizations in the Milky Way that would be technologically capable of communicating with us.

When it was first formulated, most if not all of the factors had such wide error bars that it's hard to argue that any meaningful number could come out of it.  An answer of the form "2.5 million, but maybe zero and maybe several billion or anything in between", while honest, is not a particularly useful result.  For much of the time the Drake Equation has been around, it's been useful more as a  framework for reasoning about the possibility of alien civilizations (and, in my opinion, a reasonable one) than as a way of producing a meaningful number.

Recently, though, a couple of the error ranges have tightened considerably.  Let's look at the factors in question:
  • the average rate of star formation in our galaxy.  This is currently estimated at 1.5 - 3 stars per year
  • the fraction of formed stars that have planets. This is quite likely near 100%
  • the average number of planets per star that can potentially support life.  There is some dispute over this.  You can find numbers from 0.5 to 4 or 5, and even outside that range.  My personal guess is toward the high end. 
  • the fraction of those planets that actually develop life.  At this point we can only extrapolate from life on Earth, a minimal and biased sample.  It's noteworthy that life now seems to have begun shortly (in geological terms) after suitable conditions arose.
  • the fraction of planets bearing life on which intelligent, civilized life has developed.  Developing intelligent life as we understand it took considerably longer: billions of years.  Again extrapolating from our one known example, this implies that a large fraction of life-bearing planets haven't been around long enough to develop intelligent life.
  • the fraction of these civilizations that have developed technologies that release detectable signals into space.  Still extrapolating, this fraction may be pretty high.   On geological scales, humanity developed radio pretty much instantaneously, suggesting it was nearly inevitable.
  • the length of time, L, over which such civilizations release detectable signals.  I've argued that this is probably quite short (see the links above and the discussion below for a bit more detail).
Looking at the units in those factors, we have
  • civilizations = (stars/time) * (a bunch of fractions that amount to civilizations/star) * time
which is perfectly valid.  However, I'm not sure it's the best match for the problem that we're trying to solve.  I've argued previously that timing is important.  The last factor (length of time a civilization produces detectable signals) takes that into account, but the other time factor, in the rate of star formation, seems less relevant.  There are billions of stars in the galaxy.  At a rate of a couple of stars per year that's not going to change meaningfully over human timescales.

So let's try the same general idea but with different units:
  • expected signal = planets * (expected signal / planet)
First, shift the focus from stars to planets.  For our purposes here that includes objects like planet-sized moons of gas giants.  This cuts out the estimation of star formation and planets per star, since we can now observe planets (in some cases even directly) and get a pretty good count of them.  Or at least we're now guessing about planets directly, instead of guessing about stars and planets.

Then, let's pull back a bit from the details of how a planet would produce a signal of intelligent life, and focus on the signal itself, by estimating how strong a signal we can expect from a given planet.   This consolidates the estimates of life evolving, civilization evolving, civilization developing technology and the duration of any signal into a single factor.

The "expected" means we're looking at weighted probabilities.  To take a familiar example, if you roll a six-sided die and I pay you $10 per pip that comes up, you should expect to get $35 on average and you shouldn't pay more than that to play the game.  This really only holds up if you expect to play the game a number of times.  If you only roll the dice once, you could always just get a bad roll (or a good one).

Likewise, if we say that a planet is producing a signal of a given expected strength, we're saying that's the average strength over all the possibilities for that planet -- maybe it's young with only one-celled life, maybe it's harboring a civilization that's producing radio signals, etc.  We're not claiming that it's actually producing a signal of that strength.  We can get away with this, more or less, because we'll be adding up expectations over a reasonably large number of planets.

Looking at expected signal accounts for a couple of factors.  What a planet emits in the radio spectrum will vary over time.  The raw strength will vary.  Earth has gone from watts to at least gigawatts in the past century or so.  The signal to noise ratio will also vary.  As we make better use of encryption, compression and such, our signal looks more like noise.  Signal strength also accounts for distance.  A radio signal falls off as the square of the distance. 

A given planet will have a particular profile of signal strength over time.  Ours is zero for most of our history, rises significantly as humans develop radio and (I've argued), will drop off significantly as we come to use radio more efficiently and use broadcast less and less.

There are two sources of uncertainty in what strength of signal we would expect to detect, knowing how far away a planet is and how much background noise there is:  We don't know what the signal strength profile for a given planet is, and we don't know where we are in that profile, that is, just how old the planet is at the moment.

For the first uncertainty, the best we can currently do is compare to our experience on earth.  My best guess is that we should expect a very brief blip (brief on planetary scales).  If we expect a blip on the order of hundred years and a planetary age on the order of billions of years, this reduces the expected signal -- again, "expected" in the probabilistic sense -- at any given time to a very low level.  This would be true even if planets occasionally send out strong, targeted transmissions, as ours does.

In the absence of anything better, we can account for the second uncertainty by averaging the signal strength over the expected age of the planet.  That is, we assume the planet could be at any point in its history with equal probability.  In real life, we may be able to do better by looking at factors like the age of the star and the amount of dust around it.

Strictly speaking we should be talking about intervals rather than instants, since listening for a million years is more likely to turn something up than listening for a hundred, but human timescales are tiny enough that this doesn't really affect our calculations of what we should expect with current or near-future technology over our lifetimes.  Either way, we can still define expected signal.

We also need to account for the distribution of planets in space.  If stars were uniformly distributed in space and background noise didn't matter, this would cancel out the effect of decreasing signal strength, since the number of stars at a given distance would increase as the square of the distance.

But they're not.  If they were then the nighttime sky would also be uniformly bright in all directions.  The Milky way is only about a thousand light years thick.  After about half that distance the number of stars increases much more slowly than the square of the distance.  This means we're really looking at a weighted sum of expectations rather than just multiplying planets by expectation per planet, but that doesn't greatly change the overall analysis.

Finally, we should take background noise into account.  As the strength of a signal (actual, not expected strength) drops toward zero, our ability to detect it doesn't drop in tandem.  Once the signal becomes weaker than the general background noise in that part of the sky, our chances of detecting it are already very near zero.  This correction should be applied to the signal profile before averaging over time.

My engineering intuition tells me that the upshot is that we can neglect planets more than a relatively short distance away, say tens of light-years.  At some point background noise will wash everything out.  That's more or less the limit for having a meaningful conversation anyway, since it takes a year for a radio signal to travel a light-year.

So where does that leave us?

Estimating the probability of a detectable signal from a planet requires knowing
  • The distribution of planets as a function of distance.  Our knowledge of this has sharpened dramatically over the past couple of decades.
  • The effect of distance on the strength of a signal we detect.  This is fairly well understood.
  • The background noise for any particular location in the sky.  This is directly observable.
  • The expected strength of the signal emitted by a planet, averaged over its lifetime.  This is where the uncertainty is concentrated.
Essentially we've consolidated all the various fractions of the Drake equation into a single factor and characterized it in terms of signal strength over time (which we then average over time unless we can think of something better).

When searching for life, "signal" doesn't necessarily mean "radio signal".  Soon we will be able to search for signatures such as high levels of oxygen in the atmosphere, which suggest that there is life of a similar form to ours, though not necessarily intelligent, technological or whatever.  This signal would have a much different profile from radio.  In our case it would rapidly jump from zero to full strength relatively early in our history and stay there for billions of years.  It may also be a stronger signal than radio leakage in the sense that we can feasibly detect it from further away.

If we take our experience on earth as a basis, this implies it's quite likely that we'll detect life on other planets, but unlikely that we'll detect radio signals (and probably other smoking-gun signs of civilization as we know it).  Looking for signatures of life in general is probably going to be more informative in any case.  If we don't find any radio signals from other planets, which seems more and more likely, it could just be because even planets with intelligent life don't tend to emit high signal-to-noise radio signals for long.  If we find chemical signatures indicating life on X% of planets with detectable atmospheres, that gives a strong estimate on the probability of life arising in general.  This is true whether X is 0, 100 or something in between.

[Technical note: Somewhat ironically, since I started out talking about unit analysis, the units here are less clear than they might be.  If we're talking about radio, then at any given moment a planet is emitting radio signals at a given power, say X Watts.  Power is energy per unit time.  Probably the most natural way of expressing what we actually detect over time is an amount of energy, say Y Joules -- power times time is energy.  We'd like that to stay the same whether we're talking about an actual measurement or a probabilistic estimate.  So the quantity we're trying to estimate for a given planet is power.

If we assume a particular profile of power over time, and we average it, we're summing up power over time to get total energy, then dividing by the total time span over which we think we might be looking -- the age of the planet -- to get power again.  Accounting for distance still gives power, that is energy we expect to receive per unit time.  Using units of power also accounts for the amount of time we spend looking.  If we look for 100 years we expect to detect 10 times as much signal (energy) as if we look for 10 years.  I tried to gloss over that in the main article on the grounds that the numbers are all likely to be too small to matter.  But it's better to think of a minuscule amount of power over a shorter or longer time than to try to assume everything's an instant.

I've made a few edits to the main article, mainly changing "signal strength" to "signal" in several places to try to reflect this.]

[And having gone through all that, and thought it over a bit more ... the really natural units to use here are bits and bits per second.  At the end of the day, we're trying to glean information from listening to the skies, and information is measured in bits.  This accounts for several troublesome factors:
  • We're trying to estimate detectable information from other planets.  This starts by estimating what information they transmit over time, as measured by an observer in the near vicinity (say, in low Earth orbit or on the Moon in our case)
  • I've argued that as we use compression and encryption more, our signal looks more like noise.  This is quantifiable in terms of bits and bit rates.
  • If a planet is far away or in a noisy area of the sky, we're less likely to detect a signal from it.  There are well-established formulas relating signal power, bandwidth and signal/noise ratios that can be used to translate an estimate of what radio signals a planet emits to an estimate of bits/second we could detect.
  • As above, integrating bits/time over time spent listening gives us the total information we would expect to detect, which is arguably the quantity of interest in the whole exercise.
  • So
    • bits detected = sum over time of the sum over planets of bits per second we expect to detect from each planet
    • leaving out the sums, which don't change the units: bits = (bits/second)/planet * planets * seconds
]


Monday, January 2, 2017

How natural is nature?

Physics has produced several amazingly elegant theories that reduce a huge variety of phenomena to a few basic causes and concepts.  Even if the basic concepts are just a wee bit math-heavy and the results can be a just a wee bit mind-bending, a great number of important discoveries in physics can be reduced to fairly short descriptions.
  • Thermodynamics uses a handful of laws to explain things like why perpetual motion can't happen, how engines work or why Play-Doh™ always ends up looking gray-brown if you mash it together long enough.
  • Newton's laws explain things like why the Moon goes around the Earth, how you can tell if a car in an accident was speeding or how to sink the 8-ball in the corner pocket.
  • Nöther's theorem demonstrates (in a way I've never quite completely grasped) a deep relation between symmetry and conservation -- if, for example, the equations describing motion don't care about direction then angular momentum is conserved and that figure skater spins faster and faster as the arms come in.
  • General relativity holds that, left to themselves, objects travel in a straight line, the simplest possible path.  It just doesn't always look that way because space-time isn't flat, but this is why, for example, Mercury's orbit moves just a bit every time around.
  • Quantum physics ... yeah.  Quantum physics.
It's not that quantum physics lacks elegance.  The idea that all matter and energy, basically everything we can measure, can be explained by equations similar in form to those that describe a vibrating string is pretty astounding if you think about it.  The Standard Model of quantum physics has built on this to make a large number of predictions, including predictions of new particles, that have been confirmed with outstanding accuracy.

You'd think this would be good news.  Instead, a certain uneasiness has developed around the Standard Model.  The basic framework is nice enough, but it can't completely describe what we know until you plug in several parameters.  There are 19 in all, ranging from me  (the mass of the electron, 511 keV), to θ23 (the "CKM 23-mixing angle", 2.4°) to the recently established mH, (the Higgs mass, tentatively 125.36±0.41 GeV).  There aren't just infinitely many other ways to tune the knobs, there are not one, not two but 19 knobs to tune.

Tweak a few of them the wrong way and stars can never form, or worse, no kind of solid matter can form at all.  We seem to be in some sort of special regime where the parameters just happen to have the right values for us to be here to observe them.  Even if you adopt the view that there may be infinitely other universes out there where the knobs aren't tuned right, so where else could we be (the "weak anthropic principle"), it's still all pretty unsatisfying.  Our universe is some point in a 19-dimensional space that's suitable for life forms like us to develop?  That's it?


Particle physicist Sabine Hossenfelder  argues in a piece called The LHC “nightmare scenario” has come true that yep, that's it, get over it.  As I read it she makes two points.  The smaller one is that the Large Hadron Collider which was instrumental in finding the Higgs boson has likely found all the particles it's going to find, and maybe it's time to stop trying to build bigger and bigger particle accelerators.

Fellow particle physicist Matt Strassler argues that there's no nightmare regardless of whether there are any other new particles.  The LHC has produced ridiculous amounts of data which won't be thoroughly examined for years, and it can easily produce more.  There might be, indeed probably are, interesting discoveries to be pulled out of that data now that it's pretty well established that the Higgs exists.

This seems reasonable, but it's more an argument against Hossenfelder's headline than the substance of the article.  Disputes over what experiments to do (and, more to the point, what experiments to fund) are by no means new.  Hossenfelder's and Strassler's are by no means the only views on the subject, and they may not even be particularly divergent, but in any case whether to keep building bigger particle smashers is of greatest concern to particle physicists and those who fund them.

Public policy and the sociology of science are worthy topics, but I won't be conjecturing any further about them here.  I'm more interested in Hossenfelder's larger point, which as I understand it is about what makes a good theory of physics.

When people started taking a close look at Newtonian mechanics, heat transfer and other fields they started to find anomalies under extreme conditions that eventually led to the discovery of relativity and quantum physics.  This is just part of a long history of progress in physics.  For example:
  • Ptolemy explained the motions of the planets with a system of cycles and epicycles centered around the Earth.
  • Copernicus explained those motions more simply with a system of cycles and epicycles centered around the sun.
  • Kepler did away with epicycles using the notion that the planets moved in ellipses, not circles
  • Newton explained elliptical orbits in terms of a universal gravitational force following an inverse square law
  • and Einstein explained gravitation as a property of space-time itself
(I'm always a bit leery about ascribing a particular landmark result to a particular person, as in "Ptolemy explained ...".  There is more to each of these than a single person making a single discovery even when we know a particular person had a particular key insight.  But this will do for now.)

In all these cases, the new theory didn't just explain everything the old theory did, albeit in a new way.  It either made sense of something that had seemed arbitrary in the old theory, explained new things the old theory couldn't, or both.  Copernicus and Kepler dealt with epicycles, first simplifying them and then doing away with them altogether.  Newton's mechanics explained why the planets followed elliptical orbits as described by Kepler's laws and not some other shape.  It also explained why the Moon doesn't actually follow an exactly elliptical orbit, why the daily tides rise and fall, and much more.

Einstein's theory of relativity did away with gravitation as a force.  Objects under the influence of gravity still follow Newton's first law, just in a more subtle form.  It also gave better predictions for the motions of the planets and made a number of new predictions that were later confirmed, such as the direction and frequency of light being affected by gravity and why the orbits of stars in a binary system containing a pulsar can be seen to be slowing.

It's not just that the new theories were more powerful than the old ones.  That's to be expected.  Otherwise why adopt them?  In all these cases, and many others, the new theory was also, in some sense, more elegant than the old.  Elegant, in this sense, largely means simpler.  Fewer epicycles.  One universal force.  No universal force at all.  There is also a sense of reducing seemingly unrelated things to different aspects of the same thing.  The tides and the motions of the planet are both just effects of gravity.  Space and time are just components of a the space-time continuum.



Which brings us back to the Standard Model.


So far no one has come up with a theory-breaking anomaly for the Standard Model analogous to the precession of Mercury's orbit, or some new phenomenon, say an unpredicted particle or force, that the Standard Model could have been expected to predict but didn't.  There are a few candidates, but even after decades of effort nothing has really panned out.  The experiments at the LHC found the Higgs, at an energy consistent with the Standard Model, and nothing, or at least nothing definitive, inconsistent with it.

So the Standard Model is it, right?  We've described the fundamental forces and elementary particles of the natural world.  There's plenty of work, probably an endless amount, to be done working out the ramifications of that, and how it all fits in with relativity, what exactly it means to "measure" a system described by a wavefunction, and on and on, but as to explaining the basis for particle physics, we're done.  Right?

As I understand it, Hossenfelder's answer to that would be "looks like we could be", but that answer doesn't sit well with everyone.  How can such an inelegant theory, with its 19 arbitrary parameters, be the final answer?  "They just do" can't be an adequate answer to "why do those parameters have the values they do?" can it? Hossenfelder would likely say "sure it can".

In the history of physics, power and elegance seem to go hand in hand.  Or at least, after enough anomalies with ad-hoc descriptions turn up, eventually someone comes up with a new framework where it all makes sense again.  The new theory is both more elegant and more powerful.  Some would even say more "natural" and claim that nature is itself elegant, and if it doesn't seem that way we must not understand it properly.

The Standard Model seems ready to be replaced with something better, except it doesn't seem to be producing the sort of "close, but not quite" results that led us from Newton to Einstein.  There may be more elegant theories around -- string theory gets a lot of attention in this regard -- but nothing, so far, clearly more powerful.  If there's a more "natural" theory, nature doesn't seem keen to lead us to it.


This feeling that the world has to be more elegant than our current theories may just be an occupational hazard of physicists, and not necessarily the majority at that.  Plenty of working particle physicists are content to "shut up and calculate" without worrying too much about what it all might "mean" or whether the universe has some deep hidden simplicity.

Many chemists would shake their heads at the whole business.   There are around a hundred elements one can do meaningful chemistry with, each with its own particular properties.  That's not going to change with a new theory of chemistry.  There is a theory, namely the Standard Model, which explains why those elements are the way they are, and quantum effects definitely come into play in chemistry, but from a chemist's point of view it doesn't matter how many parameters the Standard Model has.  It matters what the electrons are going to do in a particular situation.

In my own field there are several models that can define the behavior of computers, and we do refer to them (particularly state machines and stack machines) from time to time, but there is not and is never likely to be a unified theory of software engineering.  And yet the servers still run.  Mostly.

Even mathematics, which can almost be defined as the relentless pursuit of elegance, is full of quirky, inelegant results.  What's so special about manifolds in four dimensions?  Why are there 26 sporadic groups?  Why is the 3N+1 problem so hard?  And let's not even get started on the prime numbers.



Suppose that everything in the universe could be precisely described by three simple rules ... and a table of three quadrillion quadrillion seven-digit numbers.  Even storing such a table would be completely infeasible using today's technology, but suppose we meet up with an alien race with full access to it.  Our best physicists pose them questions, they consult the table and deliver a verifiable answer every time (how to reduce any measurable question and its answer to an invocation of three simple rules is an interesting question, but roll with it).  Would we say the aliens have a good theory?

On the one hand, of course they do.  The hallmark of a good theory is making testable predictions that hold up.  On the other hand, there's something less than satisfying about a planet-sized table of numbers, each essentially its own arbitrary parameter.  What happens if our aliens go away or decide that we're not worthy of True Knowledge?  Maybe we should start asking questions that will reveal the nature of the magic number table and, ideally, allow us to reduce it to something our puny minds and computers can handle.

A good theory doesn't just have to be true in the sense of making true predictions.  It also has to be comprehensible and usable.  To this end, a theory with a thousand fairly simple rules and three or three hundred parameters with values we just have to accept is far better than the one I just described.  But this is not saying anything about nature.  It says something about us.  Our "natural" theories are the ones that work best for us, not just in aligning with nature, but with our resources and the way our minds work.

From that point of view, 19 is not a prohibitive number of parameters the way three quadrillion quadrillion would be.  If that's really how it is, we can probably live with it.  But the distinction is of degree, not kind.  The problem is not with arbitrary parameters themselves, but with having an intractable number of them.  Consulting our hypothetical aliens with knowledge beyond our ability to process is really just another kind of experiment from our point of view.   Consulting the Standard Model with its human-friendly list of parameters is better, and it would be even if its predictions weren't quite as good as they are.  It's certainly better than a more "elegant" theory that doesn't fit experiment as well as it does.

Nature is what it is.  A theory is only "natural" if it fits with our nature in particular as well as nature at large.

[Re-reading this, I realize I neglected to say that, although the basic equation of the standard model is fairly compact -- you can get a T-shirt with the Standard Model Lagrangian on it -- actually finding solutions for all but the simplest conditions is generally far beyond our computing ability.  In one sense this is more than a bit like the aliens-with-the-numbers scenario, but instead of a hidden table of numbers we can't begin to access, we have an equation we can barely begin to compute.  Except maybe with quantum computers ... --D.H.]


Re-reading Hossenfelder's piece, I see one more subtle point.  The main argument doesn't seem to be that there can't possibly be an elegant theory unifying quantum physics with relativity, or even a better way of explaining the results of the Standard Model.  Rather, a search for "elegance" or a "natural" theory is no longer a good way -- if it ever was -- of deciding what particle experiments to run next.  If we do find such a unified theory, it's probably not going to be because we found a more elegant replacement for the Standard Model, or because we found an unexpected particle with a new, more powerful accelerator, but because we found something else entirely and a theory to explain it that happens to subsume the Standard Model.