Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Tuesday, April 28, 2026

What is the pound measuring?

How much does the Orion capsule (that is, the Crew Module) that splashed down on April 11 weigh? According to NASA's reference guide for Orion, 22,900 pounds.

The guide specifically lists "liftoff weight", and there are a couple of reasons for that. One is that the capsule has reaction control thrusters, which are small rocket engines that allow for fine-tuning the attitude of the craft and small-scale maneuvering, and their propellant is part of that liftoff weight. For this and other reasons, the capsule did not have the exact same contents when it splashed down as when it took off.

The other reason, of course, is that the weight of the capsule depends on where the capsule is in its trajectory. For most of the mission, that weight was essentially zero, since the capsule was coasting in freefall except at a few key points. Units of weight, like pounds, measure force, not mass. At least that's what I was taught in high school physics.

For most practical purposes, though, the pound is a unit of mass. If the door of a bank vault weighs a ton (2,000) pounds, you know it will be a little hard to move, even if it's perfectly mounted on bearings with very low friction so that when you push on it you're not trying to lift its mass. That inertia is due to its mass. If you weigh out a quantity of something, you're interested in how much of it you're getting, that is, the total mass. The force that it exerts on the scale is just a way to deterimine the mass.

You're measuring that mass by way of how much that mass weighs on Earth, but it's still mass that you're measuring. Except in specialized applications like calculating load limits or foot-pounds of torque, the amount of force something exerts under gravity is secondary to how much of it you have.

Yes, it matters that a 22lb bag of something is easier to lift than a 44lb bag, but it matters just the same that a 10kg bag is easier to lift than a 20kg bag. You don't need to know the amount of force involved (about 98 and 196 Newtons, respectively) to make that determination and no one is thinking "Hmm ... that 20kg bag will require 196 Newtons to lift" before trying to pick it up.

There are units, the pound-mass and pound-force, that make the distinction between mass and weight. The pound-mass is now defined as exactly 0.45359237 kg, and the pound-force is the weight of this mass under standard Earth gravity of 9.8m/s2.

No one uses this. Well, maybe not absolutely no one, but you won't find anything on a supermarket shelf that says it weighs, say, 1.5 lbm, because no one at the supermarket cares. If you're doing precise engineering or scientific work where the distinction matters, you're not using pounds, but kilograms and Newtons. This is just an example of the distinction I previously discussed between everyday units of measure, which can be pretty much anything, and precisely-defined scientific units of measure.

There are several reasons that SI (metric) units work better than imperial units for scientific work (and why, for example, the telemetry feed that NASA put up during Artemis II showed both SI and imperial units, with SI units first as I recall). One is the consistent use of powers of ten and standard prefixes like mega- and milli-. Another is that SI units have been standard for generations, so anything you're referencing in a scientific context is almost certainly using them. Another is the body of very careful definitions of what each unit means.

A less obvious reason is that SI units carefully make distinctions that we gloss over in everyday use, particularly the mass-weight distinction. During re-entry, when a capsule may be pulling on the order of 5g, it matters quite a bit that the forces on the body of the capsule are much higher than when the capsule is on the launch pad. You want to be talking about Newtons of force and not kilograms of mass when you do those calculations. Using pound interchangeably for pound-mass and pound-force in everyday speech makes good sense when you're buying groceries. Trying to use mass and force interchangeably in mechanical engineering is a recipe for disaster.

To make the distinction completely clear, the Newton is defined as a kilogram-meter per second squared, with no reference to Earth's gravity. A pound-mass weighs a pound-force under standard gravity because we don't really care about the distinction when using pounds. A kilogram weighs about 9.8 Newtons, which helps keep the distinction clear when it matters.

NASA is happy to quote the weight of Orion in pounds and show its speed in miles per hour because the US audience is used to those units. Trying to point out that actually the mass is about 10.4 tonnes and the weight varies is just going to get in the way unless you're specifically talking about the effects of acceleration or microgravity. Using pounds interchangeably for mass and weight is only incorrect if you're doing engineering or science, but then you shouldn't be using pounds at all.

Saturday, August 10, 2024

On myths and theories

 Generally when people say something is a "myth", they mean it's not true:

"Are all bats blind?"

"No, that's just a myth."

There's nothing wrong that that, of course, but there's a richer, older, meaning of myth: A story we tell to explain something in the world.  In that sense, a myth is a story saying something like "This is the way it is because so-and-so did thus-and-such" (many constellations have stories like this associated with them) or "So-and-so did this so that thus-and-such" (the story of Prometheus bringing fire to humanity is a famous example).

The word theory is also used in two senses.  Generally, people use it to mean something that might be true but isn't proven.

"I personally think that the Loch Ness monster is actually an unusually large catfish, but that's just a theory."

In science, though, a theory is a coherent explanation of some set of phenomena, which can be tested experimentally.  There are a couple of related senses of theory, for example mathematical sense of theory, as in group theory, meaning a comprehensive framework that brings together a set of results and sets the direction for future research.  While mathematical theories ultimately rest on mathematical proofs and not measurements of the physical world, the goal is still to understand and explain.

For example, Newton's theory of universal gravitation explains a wide variety of phenomena, including apples falling from trees, the daily tides of the sea and the motions of planets in their orbits, by positing that any two massive bodies exert an attractive force on each other, and that this force depends only on the masses of the bodies and the distance between their centers of gravity (more precisely, it's the product of the two masses, divided by the square of the distance, times a constant that's the same everywhere in the universe).

Newton's theory actually gives measurably incorrect results once you start measuring the right things carefully enough.  For example, it gets Mercury's orbit wrong by a little bit, even after you account for the effects of the other planets (particularly Jupiter), and it doesn't explain gravitational lensing (an image will be distorted by the presence of mass between the observer and what is seen). 

Newtonian gravity is still taught anyway, since effects like these don't matter in most cases and it's much easier to multiply masses and divide by distance squared than to deal with the tensor calculus that General Relativity requires.

My point here is that, as with myths, the ability to explain is more important than some notion of objective truth.  As far as we currently understand it, Einstein's theory of gravity, General Relativity, is "true", while Newtonian gravity is "false", but Newton's version is still in wide use because it works just as well as an explanation, since in most cases it gives the same results for all intents and purposes.

Myths and theories both aim to explain, but there are a couple of key differences.  First, myths are stories.  Theories, even though they're sometimes referred to as stories, aren't stories in the usual sense.  There is no protagonist, or antagonist, or any characters at all.  Neither Newton's nor Einsteins theory of gravity starts out "Long ago, Gravity was looking at the sun in empty space, and thought 'I should make the planets go around it'" or anything like that.

Second, and perhaps more important, theories are not just explanations of things we already know, but the basis for predictions about things we don't know yet.  In the famous photographic experiments done during the eclipse of 1919, general relativity predicted that stars would appear in a different position in the photographs, due to the Sun's gravity distorting space, than the Newtonian version would predict (which was that they would be in the same place they'd be seen when the Sun wasn't between them and the Earth).  There's some dispute as to whether the actual photographs could be measured precisely enough to demonstrate that, but there's no dispute that the effect is real, thanks to plenty of other examples.

Myths make no claim of prediction.  If a particular myth says that a particular constellation is there because of some particular actions by some particular characters, it says nothing about what other constellations there might be.  The story of Prometheus bringing fire to humanity doesn't predict steam engines or cell phones.

It's exactly this power of prediction that gives scientific theories their value.  It's beside the point to say that some particular scientific theory is "just a theory".  Either it gives testable predictions that are borne out by actual measurements, or it doesn't.

Friday, August 9, 2024

Wicked gravity

Every once in a while in my news feed I run across an article about colonizing other planets, Mars in particular.  The most recent one was about an idea that might make it possible to raise the surface temperature by 10C (18F) in a matter of months.  That would be enough to melt water in some places, which would be important to those of us who need liquid water to drink and to irrigate crops.

All you have to do is mine the right raw materials and synthesize about two Empire State Buildings worth of a particular form of aerosol particle, and blast it into the atmosphere.  You'd have to keep doing this, at some rate, indefinitely since the particles will eventually settle out.

The authors of this idea don't claim that this would make Mars inhabitable, only that it would be a first step.  This is fortunate, since there are a few other practical obstacles, even if the particle-blasting part could be made to work:

  • The mean surface temperature of Mars is -47C (-53F) as opposed to 14C (52F) for Earth.  The resulting -37C (-35F) would not exactly be balmy.
  • Atmospheric pressure at the lowest point on Mars is around 14 mbar, compared to about 310 mbar at the top of Mount Everest.  Even if the atmosphere of Mars were 100% oxygen, the partial pressure would still be around 20% of what it is atop Everest, and there's a reason they call that the Death Zone.  In practice, you'd at least want some water vapor in the mix.
  • But of course, the atmosphere on Mars is not 100% oxygen (and even if it were, it wouldn't be for long, since oxygen is highly reactive -- exactly why we need it to breathe).  It's actually 0.1% oxygen.  There is oxygen in the atmosphere, but it's locked up in carbon dioxide, which makes up about 95% of the atmosphere.
It's at least technically feasible to build small, sealed outposts on the surface of Mars with adequate oxygen and liquid water, at a temperature where people could walk around comfortably, all using local materials.  Terraforming the whole planet is Not ... Going ... To ... Happen.

But let's assume it does.  Somehow, we figure out how to crack oxygen out of surface rocks (there's plenty of iron oxide around; again, there's carbon dioxide in the atmosphere, but nowhere near enough of it) and pump it into the atmosphere at a truly massive scale, far beyond any industrial process that's ever happened on Earth.  Mars's atmosphere has a mass of about 2.5x1016 kg, and that would need to increase by a factor of at least five, essentially all of it oxygen, for even the deepest point in Mars to have the same breathability as the peak of Everest.

By comparison, total emissions of carbon dioxide since 1850 are around  2.4×1015 kg and current emissions are around 4×1013 kg per year.  In other words, if we could pump oxygen into Mars's atmosphere at the same rate we're pumping carbon dioxide into Earth's atmosphere, it would take about three centuries before the lowest point on Mars had breathable air -- assuming all that oxygen stayed put instead of, say, recombining with the iron (or whatever) it had been split off from or escaping into space.

This is just scratching the surface of the practical difficulties involved in trying to terraform a planet.  Planets are big, yo[citation needed].

But then, not always big enough.  Broadly speaking, there's a reason that there's lots of hydrogen in Jupiter's atmosphere (about 85%, another 14% helium), while Mars's is mostly carbon dioxide and the Moon has essentially no atmosphere.  Jupiter's gravity is strong enough to keep light molecules like molecular hydrogen from escaping on their own or being carried away by the solar wind.  Mars's isn't.  It can hold onto heavier molecules like carbon dioxide OK, though still with some loss over time, but lighter molecules aren't going to stick around.

Earth is somewhere in the middle.  We don't have any loose hydrogen to speak of because it reacts with oxygen (because life), but we also don't have much loose helium because it escapes.

Blasting oxygen into Mars's atmosphere would work for a while.  Probably for a long while, in human terms (to be fair, atmospheric escape on Mars is measured in kg per second, or thousands of tons per year, much smaller than the in-blasting rate would be).  In the end, though, trying to terraform Mars means taking oxygen out of surface minerals and sending it into space, with a stopover in the atmosphere.

But there's another wildcard when it comes to establishing a long-term presence on a planet like Mars.  Let's put aside the idea of terraforming the atmosphere and stick to enclosed, radiation-shielded, heated spaces with artificially dense air.

The surface gravity of Mars is about 40% of that on Earth.  What does that mean?  We have no idea.  We have some idea of how microgravity (also known as zero-g) affects people.  Though fewer than a thousand people have ever been to space, some have spent long enough to study the effects.  They're not great.  They include loss of muscle and bone, a weakened immune system, decreased production of red blood cells and lots of other, less serious issues.

Obviously, none of this is fatal, there are ways to mitigate most of the effects, and some of them, like decreased muscle mass, may not matter if you're going to spend your whole life in space rather than coming back to earth after a few months (no one has ever spent more than about 14 months in space).  But then, that's a problem, too.  No one has spent years in microgravity.  No one has ever been born in microgravity or grown up in it.  We can guess what might happen, but it's a guess.

No one, ever, has spent any significant time in 40% of Earth gravity.  The closest is that two dozen people have been to the Moon (16% of Earth gravity), staying at most just over three days.  We know even less about the effects of Mars gravity on humans than we do about microgravity, which is only a little bit.

Maybe people would be just fine.  Maybe 40% is enough to trigger the same responses as happen normally under full Earth gravity.  Maybe it leads to a slow, miserable death as organ systems gradually shut down.  Maybe babies can be born and grow to adulthood just as well with 40% gravity as 100%.  Leaving aside the ethics of finding that out, maybe it just won't work.  Maybe a child raised under 40% gravity is subject to a host of barely-manageable ailments.  Maybe they do just great and enjoy a childhood of truly epic dunks at the 4-meter basketball hoop on the dome's playground.

Whatever the answer is, there's absolutely nothing a hypothetical Mars colony could do about it.  You can corral a bit of atmosphere into a sealed space and adjust it to be breathable.  You can heat a small corner of the new world to human-friendly temperatures.  You can separate usable soil out of the salty, toxic surfaces and grow food in the reduced light (the Sun is about 43% as bright on Mars).  You can project scenes of a lush, green landscape on the walls.

No matter what you do, the gravity is going to be what it is, and whoever's living there will have to live with it however they can.

Friday, January 12, 2024

On knowing a lot about something and something about a lot of things

The physicist Richard Feynman told a story about being on a panel of experts from a variety of academic fields.  The full details are in one of the Surely you're joking books I read many years ago.  I'm paraphrasing from memory here because lazy.  The gist is that the panel was asked to look at someone's paper that pulled together ideas from a variety of fields and was generating a lot of buzz.  Just the sort of thing you'd want an interdisciplinary panel of experts to look at.

All the experts on the panel had a similar reaction: Overall, it looks very interesting, but the stuff in my area needs quite a bit of work -- this bit is a little bit off, they're mis-applying these terms and these parts are just wrong.  But there are some really interesting ideas and this is definitely worth further attention.

In Feynman's telling, at least, he was the one to offer a different take: If every expert is saying the part they know about is bad, that says it's just bad all the way through.  It doesn't really matter what an expert thinks of the area outside their expertise.


Relying on people's subjective impressions is risky.  What we need here is some way to objectively determine the value of a paper that crosses areas of knowledge.  Here's one way to do it: Have everyone rate the paper in each area on a scale of 0 - 100 and then pull together the numbers.

Let's say we have five people on the panel, specializing in music theory, physics, Thai cuisine, medieval literature and athletics, and someone has written a paper pulling together ideas from these fields into an exciting new synthesis.  Their ratings might be:

Music Physics Thai food Medi. lit Athletics Overall
Music theorist 25 75 80 65 85 66
Physicist 70 15 80 60 60 57
Thai chef 65 85 5 70 70 59
Medievalist 90 70 80 25 85 70
Athlete 85 90 95 90 30 78
Overall 67 67 68 62 66 66

Overall, the panel rates the paper 66 out of 100.  We don't have enough context here to know whether 66 is a good score or a mediocre score, but it certainly doesn't look horrible.  The highest score is in Thai cuisine, and the highest score there was from the athletics expert, so maybe the author has discovered some interesting contribution to Thai food by way of athletics.

But hang on a minute.  The highest overall score is in Thai cuisine, but the lowest rating in that category from any expert is the 5 from the Thai chef.  Let's ask each of the experts how much they know about their fields and those outside their home turf:

Music Physics Thai food Medi. lit Athletics
Music theorist 95 5 15 10 5
Physicist 20 100 10 5 5
Thai chef 5 10 100 10 15
Medievalist 10 5 10 95 10
Athlete 10 15 5 10 95

Everyone feels confident in their own field, as you might expect, and they don't feel particularly confident outside their own field, which also makes sense. There's also quite a bit more variation outside the home fields, which makes a certain amount of sense as well.  Maybe the physicist happens to have taken a couple of courses in music theory.  Maybe the athlete has only had Thai food once.  You can expect someone to have studied extensively in their field, but who knows what they've done outside it.

We should take this into account when looking at the ratings.  A Thai chef saying that the paper is weak in Thai cuisine means more than an athlete saying it's great.  If we take a weighted average by multiplying each rating by the panelist's confidence, adding those up and dividing by the total weight (that is, the total of the confidence numbers), we get a considerably different picture:

Music Physics Thai food Medi. lit Athletics Overall
Weighted result 40 33 27 38 42 36

Overall, the paper rates 36 out of 100 rather than 66.  Its weakest area is Thai cuisine, and even its strongest area, athletics, is well below the previous score of 66.

This seems much more plausible.  The person who knows Thai food best rated it low, and now we're counting that ten times more heavily than the physicist's rating and twenty times more heavily than the judge who said they knew least about it.

I think there are a few lessons to be drawn here.  First, it's important to take context into account.  The medievalist's rating means a lot if it's about Medieval literature and not much if it's about physics, unless they also happen to have a background there.  Second, just putting numbers on something doesn't make it any more or less rigorous.  The 66 rating and the 36 rating are both numbers, but one means a lot more than the other.

Third, when it comes specifically to averages, a weighted average can be a useful tool for expressing how much any particular data point should count for.  Just be sure to assign the weights independently from the numbers you're weighting.  Asking the panelists ahead of time how much they know about each field makes sense.  Looking at rating numbers and then deciding how much to weight them is a classic example of data fiddling.

Finally, it's worth keeping in mind that people often give the benefit of the doubt to something that sounds plausible when they don't have anything better to go on.  As I understand it, this was the case in Feynman's example.  In that case, giving the paper to a panel of experts from different fields gave the author much more room to hide than if they'd, say, submitted a shortened version of the paper for each field.

The answer is not necessarily to actively distrust anything from outside one's own expertise, but it's important not to automatically trust something you don't know about just because it seems reasonable.  The better evaluation isn't "I don't believe it" but "I really can't say".

I'll leave it up to the reader how any of this might apply to, say, generative AI, LLMs and chatbots.

Friday, July 6, 2018

Are we alone in the face of uncertainty?

I keep seeing articles on the Drake equation and the Fermi Paradox on my news feed, and since I tend to click through and read them, I keep getting more of them.  And since I find at least some of the ideas interesting, I keep blogging about them.  So there will probably be a few more posts on this topic.  Here's one.

One of the key features of the Drake equation is how little we know, even now, about most of the factors.  Along these lines, a recent (preprint) paper by Anders Sandberg, Eric Drexler and Toby Ord claims to "dissolve" the Fermi Paradox (with so many other stars out there why haven't we heard from them?), claiming to find "a substantial ex ante probability of there being no other intelligent life in our observable universe".

As far as I can make out, "ex ante" (from before) means something like "before we gather any further evidence by trying to look for life".  In other words, there's no particular reason to believe there should be other intelligent life in the universe, so we shouldn't be surprised that we haven't found any.

I'm not completely confident that I understand the analysis correctly, but to the extent I do, I believe it goes like this (you can probably skip the bullet points if math makes your head hurt -- honestly, some of this makes my head hurt):
  • We have very little knowledge of the some of the factors in the Drake equation, particularly fl (probability of life on a planet that might support life) fi (probability of a planet with life developing intelligent life) and L (the length of time a civilization produces a detectable signal)
  • Estimates of those range over orders of magnitude.
    • Estimates for L range from 50 years to a billion or even 10 billion years.
    • The authors do some modeling and come up with a range of uncertainty of 50 orders of magnitude for fl.  That is, it might be close to 1 (that is, close to 100% certain), or it might be more like 1 in 100,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000.  Likewise they take fi to range over three orders of magnitude, from near 1 to 1 in 1,000.
  • Rather than assigning a single number to every term, as most authors do, it makes more sense to assign a probability distribution.  That is, instead of saying "the probability of life arising on a suitable planet is 90%", or 0.01% or whatever, assign probability for each possible value (the actual math is a bit more subtle, but that should do for our purposes).  Maybe the most likely probability of life developing intelligence is 1 in 20, but there's a possibility, though not as likely, that it's actually 1 in 10 or 1 in 100, so take that into account with a probability distribution..
  • (bear in mind that the numbers were looking at are themselves probabilities, so we're assigning a probability that the probability is a given number -- this is the part that makes my head hurt a bit)
  • Since we're looking very wide ranges of values, a reasonable distribution is the "log normal" distribution -- basically "the number of digits fits a bell curve".
  • These distributions have very long tails, meaning that if, say, 1 in a thousand is a likely value for the chance of life evolving into intelligent life, then (depending on the exact parameters) 1 in a million may be reasonably likely, 1 in a billion not too unlikely and 1 in trillion is not out of the question.
  • The factors in the Drake equation multiply, following the rules of probability, so it's quite possible that the aggregate result is very small.
    • For example if it's reasonably likely that fl is 1 in a trillion and fi is 1 in a million, then we can't ignore the chance that the product of the two is 1 in a quintillion.
    • Numbers like that would make it unlikely that there is any life in our galaxy's few hundred billion stars and that ours just happened to get lucky.
  • Putting it all together, they estimate that there's a significant chance that we're alone in the observable universe.

I'm not sure how much of this I buy.

There are two levels of probability here.  The terms in the Drake equation represent what has actually happened in the universe.  An omniscient observer that knew the entire history of every planet in the universe (and exactly what was meant by "life" and "intelligent") could count the number of planets, the number that had developed life and so forth and calculate the exact values of each factor in the equation.

The probability distributions in the paper, as I understand it, represent our ignorance of these numbers.  For all we know, the portion of "habitable" planets with intelligent life is near 100%, or near 1 in a quintillion or even lower.  If that's the case, then the paper is exploring to what extent our current knowledge is compatible with there being no other life in the universe.  The conclusion is that the two are fairly compatible -- if you start with what (very little) we know about the likelihood of life and so forth, there's a decent chance that the low estimates are right, or even too optimistic, and there's no one but us.

Why?  Because low probabilities are more plausible than we think, and multiplying probabilities increases that effect.  Again, the math is a bit subtle, but if you have a long chain of contingencies, any one of them failing breaks the whole chain.  If you have several unlikely links in the chain, the chances of the chain breaking are even better.


The conclusion -- that for all we know life might be extremely rare -- seems fine.  It's the methodology that makes me a bit queasy.

I've always found the Drake equation a bit long-winded.  Yes, the probability of intelligent life evolving on a planet is the probability of life evolving at all multiplied by the probability of life evolving into intelligent life, but does that really help?

On the one hand, it seems reasonable to separate the two.  As far as we know it took billions of years to go from one to the other, so clearly they're two different things.

But we don't really know the extent of our uncertainty about these things.  If you ask for an estimate of any quantity like this, or do your own estimate based on various factors, you'll likely* end up with something in the wide range of values people consider plausible enough to publish (I'm hoping to say more on this theme in a future post).  No one is going to say "zero ... absolutely no chance" in a published paper, so it's a matter of deriving a plausible really small number consistent given our near-complete ignorance of the real number -- no matter what that particular number represents or how many other numbers it's going to be combined with.

You could almost certainly fit the results of surveying several good-faith attempts into a log-normal distribution.  Log-normal distributions are everywhere, particularly where the normal normal distribution doesn't fit because the quantity being measured has something exponential about it -- say, you're multiplying probabilities or talking about orders of magnitude.

If the question is "what is the probability of intelligent life evolving on a habitable planet?" without any hints as to how to calculate it, that is, one not-very-well-determined number rather than two, then the published estimates, using various methodologies, should range from a small fraction to fairly close to certainty depending on the assumptions used by the particular authors.  You could then plug these into a log normal distribution and get some representation of our uncertainty about the overall question, regardless of how it's broken down.

You could just as well ask "What is the probability of any self-replicating system arising on a habitable planet?", "What is the probability of a self-replicating system evolving into cellular life?"  "What is the probability of cellular life evolving into multicellular life?" and so forth, that is, breaking the problem down into several not-very-well-determined numbers.  My strong suspicion is that the distribution for any one of those sub-parts will look a lot like the distribution for the one-question version, or the parts of the two-question version, because they're basically the same kind of guess as any answer to the overall question.  The difference is just in how many guesses your methodology requires you to make.

In particular, I seriously doubt that anyone is going to cross-check that pulling together several estimates is going to yield the same distribution, even approximately, as what's implied by a single overall estimate.  Rather, the more pieces you break the problem into, the more likely really small numbers become, as seen in the paper.


I think this is consistent with the view that the paper is quantifying our uncertainty.  If the methodology for estimating the number of civilizations requires you to break your estimate into pieces, each itself with high uncertainty, you'll get an overall estimate with very high uncertainty.  The conclusion "we're likely to be alone" will lie within that extremely broad range, and may even take up a sizable chunk of it.  But again, I think this says much more about our uncertainty than about the actual answer.

I suspect that if you surveyed estimates of how likely intelligent life is using any and all methodologies*, the distribution would imply that we're not likely to be alone, even if intelligent life is very rare.  If you could find estimates of fine-grained questions like "what is the probability of multicellular life given cellular life?" you might well get a distribution that implied we're an incredibly unlikely fluke and really shouldn't be here at all.  In other words, I don't think the approach taken in the paper is likely to be robust in the face of differing methodologies.  If it's not, it's hard to draw any conclusions from it about the actual likelihood of life.

I'm not even sure, though, how feasible it would be to survey a broad sample of methodologies.  The Drake formulation dominates discussion, and that itself says something.  What estimates are available to survey depends on what methods people tend to use, and that in turn depends on what's likely to get published.  It's not like anyone somehow compiled a set of possible ways to estimate the likelihood of intelligent life and prospective authors each picked one at random.

The more I ponder this, the more I'm convinced that the paper is a statement about the Drake equation and our uncertainty in calculating the left hand side from the right.  It doesn't "dissolve" the Fermi paradox so much as demonstrate that we don't really know if there's a paradox or not.  The gist of the paradox is "If intelligent life is so likely, why haven't we heard from anyone?", but we really have no clear idea how likely intelligent life is.


* So I'm talking about probabilities of probabilities about probabilities?

Friday, March 10, 2017

Science on a shoestring

On the other blog I would occasionally put out short notices of neat hacks (as always, "hack" in the "solving problems ingeniously" sense).  I recently ran across one that didn't have much to do with the web, so I thought I'd carry that tradition over to this blog.


Muons are subatomic particles similar to electrons but much heavier.  They are generally produced in high-energy interactions in particle accelerators or from cosmic rays slamming into the atmosphere.  Muons at rest take about 2 microseconds to decay, actually a pretty long time for an unstable particle.  Muons from cosmic ray collisions are moving fast enough that they take measurably longer to decay (in our reference frame), which is one of the many pieces of supporting evidence for special relativity.

The GRAPES-3 detector at Ooty in Tamil Nadu, India detects just such decays using an array of detectors set into a hill 2200m (7200 ft) above sea level.  The detectors themselves are made largely from recycled materials, particularly square metal pipes formerly used in construction projects in Japan.  The total annual budget for the project is under $400,000, but the team has already produced significant results.  Auntie has more details on the construction of the instruments here.

There are a couple of narratives that are often spun around stories like this.  One is a sort of condescending "Isn't that cute?" with maybe a reference to the Professor on Gilligan's Island building a radio out of coconuts.  Another is "Look what people can do without huge budgets.  Why do we need all these multi-billion-dollar projects anyway?"

I'd rather not tell either of those.  What I see here is highly skilled scientists making use of the resources they have available to produce significant results.  Their counterparts at CERN or whatever are making use of different resources to produce different significant results.  Both are moving the ball forward.  There have been plenty of neat hacks at CERN, including something called "HTTP",  but today I wanted to call out GRAPES-3, mainly because it's just plain cool.

Monday, January 2, 2017

How natural is nature?

Physics has produced several amazingly elegant theories that reduce a huge variety of phenomena to a few basic causes and concepts.  Even if the basic concepts are just a wee bit math-heavy and the results can be a just a wee bit mind-bending, a great number of important discoveries in physics can be reduced to fairly short descriptions.
  • Thermodynamics uses a handful of laws to explain things like why perpetual motion can't happen, how engines work or why Play-Doh™ always ends up looking gray-brown if you mash it together long enough.
  • Newton's laws explain things like why the Moon goes around the Earth, how you can tell if a car in an accident was speeding or how to sink the 8-ball in the corner pocket.
  • Nöther's theorem demonstrates (in a way I've never quite completely grasped) a deep relation between symmetry and conservation -- if, for example, the equations describing motion don't care about direction then angular momentum is conserved and that figure skater spins faster and faster as the arms come in.
  • General relativity holds that, left to themselves, objects travel in a straight line, the simplest possible path.  It just doesn't always look that way because space-time isn't flat, but this is why, for example, Mercury's orbit moves just a bit every time around.
  • Quantum physics ... yeah.  Quantum physics.
It's not that quantum physics lacks elegance.  The idea that all matter and energy, basically everything we can measure, can be explained by equations similar in form to those that describe a vibrating string is pretty astounding if you think about it.  The Standard Model of quantum physics has built on this to make a large number of predictions, including predictions of new particles, that have been confirmed with outstanding accuracy.

You'd think this would be good news.  Instead, a certain uneasiness has developed around the Standard Model.  The basic framework is nice enough, but it can't completely describe what we know until you plug in several parameters.  There are 19 in all, ranging from me  (the mass of the electron, 511 keV), to θ23 (the "CKM 23-mixing angle", 2.4°) to the recently established mH, (the Higgs mass, tentatively 125.36±0.41 GeV).  There aren't just infinitely many other ways to tune the knobs, there are not one, not two but 19 knobs to tune.

Tweak a few of them the wrong way and stars can never form, or worse, no kind of solid matter can form at all.  We seem to be in some sort of special regime where the parameters just happen to have the right values for us to be here to observe them.  Even if you adopt the view that there may be infinitely other universes out there where the knobs aren't tuned right, so where else could we be (the "weak anthropic principle"), it's still all pretty unsatisfying.  Our universe is some point in a 19-dimensional space that's suitable for life forms like us to develop?  That's it?


Particle physicist Sabine Hossenfelder  argues in a piece called The LHC “nightmare scenario” has come true that yep, that's it, get over it.  As I read it she makes two points.  The smaller one is that the Large Hadron Collider which was instrumental in finding the Higgs boson has likely found all the particles it's going to find, and maybe it's time to stop trying to build bigger and bigger particle accelerators.

Fellow particle physicist Matt Strassler argues that there's no nightmare regardless of whether there are any other new particles.  The LHC has produced ridiculous amounts of data which won't be thoroughly examined for years, and it can easily produce more.  There might be, indeed probably are, interesting discoveries to be pulled out of that data now that it's pretty well established that the Higgs exists.

This seems reasonable, but it's more an argument against Hossenfelder's headline than the substance of the article.  Disputes over what experiments to do (and, more to the point, what experiments to fund) are by no means new.  Hossenfelder's and Strassler's are by no means the only views on the subject, and they may not even be particularly divergent, but in any case whether to keep building bigger particle smashers is of greatest concern to particle physicists and those who fund them.

Public policy and the sociology of science are worthy topics, but I won't be conjecturing any further about them here.  I'm more interested in Hossenfelder's larger point, which as I understand it is about what makes a good theory of physics.

When people started taking a close look at Newtonian mechanics, heat transfer and other fields they started to find anomalies under extreme conditions that eventually led to the discovery of relativity and quantum physics.  This is just part of a long history of progress in physics.  For example:
  • Ptolemy explained the motions of the planets with a system of cycles and epicycles centered around the Earth.
  • Copernicus explained those motions more simply with a system of cycles and epicycles centered around the sun.
  • Kepler did away with epicycles using the notion that the planets moved in ellipses, not circles
  • Newton explained elliptical orbits in terms of a universal gravitational force following an inverse square law
  • and Einstein explained gravitation as a property of space-time itself
(I'm always a bit leery about ascribing a particular landmark result to a particular person, as in "Ptolemy explained ...".  There is more to each of these than a single person making a single discovery even when we know a particular person had a particular key insight.  But this will do for now.)

In all these cases, the new theory didn't just explain everything the old theory did, albeit in a new way.  It either made sense of something that had seemed arbitrary in the old theory, explained new things the old theory couldn't, or both.  Copernicus and Kepler dealt with epicycles, first simplifying them and then doing away with them altogether.  Newton's mechanics explained why the planets followed elliptical orbits as described by Kepler's laws and not some other shape.  It also explained why the Moon doesn't actually follow an exactly elliptical orbit, why the daily tides rise and fall, and much more.

Einstein's theory of relativity did away with gravitation as a force.  Objects under the influence of gravity still follow Newton's first law, just in a more subtle form.  It also gave better predictions for the motions of the planets and made a number of new predictions that were later confirmed, such as the direction and frequency of light being affected by gravity and why the orbits of stars in a binary system containing a pulsar can be seen to be slowing.

It's not just that the new theories were more powerful than the old ones.  That's to be expected.  Otherwise why adopt them?  In all these cases, and many others, the new theory was also, in some sense, more elegant than the old.  Elegant, in this sense, largely means simpler.  Fewer epicycles.  One universal force.  No universal force at all.  There is also a sense of reducing seemingly unrelated things to different aspects of the same thing.  The tides and the motions of the planet are both just effects of gravity.  Space and time are just components of a the space-time continuum.



Which brings us back to the Standard Model.


So far no one has come up with a theory-breaking anomaly for the Standard Model analogous to the precession of Mercury's orbit, or some new phenomenon, say an unpredicted particle or force, that the Standard Model could have been expected to predict but didn't.  There are a few candidates, but even after decades of effort nothing has really panned out.  The experiments at the LHC found the Higgs, at an energy consistent with the Standard Model, and nothing, or at least nothing definitive, inconsistent with it.

So the Standard Model is it, right?  We've described the fundamental forces and elementary particles of the natural world.  There's plenty of work, probably an endless amount, to be done working out the ramifications of that, and how it all fits in with relativity, what exactly it means to "measure" a system described by a wavefunction, and on and on, but as to explaining the basis for particle physics, we're done.  Right?

As I understand it, Hossenfelder's answer to that would be "looks like we could be", but that answer doesn't sit well with everyone.  How can such an inelegant theory, with its 19 arbitrary parameters, be the final answer?  "They just do" can't be an adequate answer to "why do those parameters have the values they do?" can it? Hossenfelder would likely say "sure it can".

In the history of physics, power and elegance seem to go hand in hand.  Or at least, after enough anomalies with ad-hoc descriptions turn up, eventually someone comes up with a new framework where it all makes sense again.  The new theory is both more elegant and more powerful.  Some would even say more "natural" and claim that nature is itself elegant, and if it doesn't seem that way we must not understand it properly.

The Standard Model seems ready to be replaced with something better, except it doesn't seem to be producing the sort of "close, but not quite" results that led us from Newton to Einstein.  There may be more elegant theories around -- string theory gets a lot of attention in this regard -- but nothing, so far, clearly more powerful.  If there's a more "natural" theory, nature doesn't seem keen to lead us to it.


This feeling that the world has to be more elegant than our current theories may just be an occupational hazard of physicists, and not necessarily the majority at that.  Plenty of working particle physicists are content to "shut up and calculate" without worrying too much about what it all might "mean" or whether the universe has some deep hidden simplicity.

Many chemists would shake their heads at the whole business.   There are around a hundred elements one can do meaningful chemistry with, each with its own particular properties.  That's not going to change with a new theory of chemistry.  There is a theory, namely the Standard Model, which explains why those elements are the way they are, and quantum effects definitely come into play in chemistry, but from a chemist's point of view it doesn't matter how many parameters the Standard Model has.  It matters what the electrons are going to do in a particular situation.

In my own field there are several models that can define the behavior of computers, and we do refer to them (particularly state machines and stack machines) from time to time, but there is not and is never likely to be a unified theory of software engineering.  And yet the servers still run.  Mostly.

Even mathematics, which can almost be defined as the relentless pursuit of elegance, is full of quirky, inelegant results.  What's so special about manifolds in four dimensions?  Why are there 26 sporadic groups?  Why is the 3N+1 problem so hard?  And let's not even get started on the prime numbers.



Suppose that everything in the universe could be precisely described by three simple rules ... and a table of three quadrillion quadrillion seven-digit numbers.  Even storing such a table would be completely infeasible using today's technology, but suppose we meet up with an alien race with full access to it.  Our best physicists pose them questions, they consult the table and deliver a verifiable answer every time (how to reduce any measurable question and its answer to an invocation of three simple rules is an interesting question, but roll with it).  Would we say the aliens have a good theory?

On the one hand, of course they do.  The hallmark of a good theory is making testable predictions that hold up.  On the other hand, there's something less than satisfying about a planet-sized table of numbers, each essentially its own arbitrary parameter.  What happens if our aliens go away or decide that we're not worthy of True Knowledge?  Maybe we should start asking questions that will reveal the nature of the magic number table and, ideally, allow us to reduce it to something our puny minds and computers can handle.

A good theory doesn't just have to be true in the sense of making true predictions.  It also has to be comprehensible and usable.  To this end, a theory with a thousand fairly simple rules and three or three hundred parameters with values we just have to accept is far better than the one I just described.  But this is not saying anything about nature.  It says something about us.  Our "natural" theories are the ones that work best for us, not just in aligning with nature, but with our resources and the way our minds work.

From that point of view, 19 is not a prohibitive number of parameters the way three quadrillion quadrillion would be.  If that's really how it is, we can probably live with it.  But the distinction is of degree, not kind.  The problem is not with arbitrary parameters themselves, but with having an intractable number of them.  Consulting our hypothetical aliens with knowledge beyond our ability to process is really just another kind of experiment from our point of view.   Consulting the Standard Model with its human-friendly list of parameters is better, and it would be even if its predictions weren't quite as good as they are.  It's certainly better than a more "elegant" theory that doesn't fit experiment as well as it does.

Nature is what it is.  A theory is only "natural" if it fits with our nature in particular as well as nature at large.

[Re-reading this, I realize I neglected to say that, although the basic equation of the standard model is fairly compact -- you can get a T-shirt with the Standard Model Lagrangian on it -- actually finding solutions for all but the simplest conditions is generally far beyond our computing ability.  In one sense this is more than a bit like the aliens-with-the-numbers scenario, but instead of a hidden table of numbers we can't begin to access, we have an equation we can barely begin to compute.  Except maybe with quantum computers ... --D.H.]


Re-reading Hossenfelder's piece, I see one more subtle point.  The main argument doesn't seem to be that there can't possibly be an elegant theory unifying quantum physics with relativity, or even a better way of explaining the results of the Standard Model.  Rather, a search for "elegance" or a "natural" theory is no longer a good way -- if it ever was -- of deciding what particle experiments to run next.  If we do find such a unified theory, it's probably not going to be because we found a more elegant replacement for the Standard Model, or because we found an unexpected particle with a new, more powerful accelerator, but because we found something else entirely and a theory to explain it that happens to subsume the Standard Model.

Thursday, September 1, 2016

Can we prove a dog is happy?

The previous post talked about qualia, or subjective experiences, but why should we care?  This being a matter of philosophy, there are a variety of answers to that, starting with "Why care about anything?" but nonetheless, there seems to be something significant about the question.  At least from my own subjective point of view.

For one thing, it seems like one of those fundamental questions.  How can we come to a complete understanding of the universe without understanding how we experience it?  Perhaps more than that, there are ethical concerns.  If we wish to increase happiness or we do not wish to cause unnecessary suffering in the world, we should understand what happiness and suffering are.  Outward appearances will only tell us so much.  It would be good to have more reliable indicators, or at least to know how reliable the ones we have are.

The problem with subjective experiences, though, is that they are subjective.  I can be well convinced that my own subjective experience is real.  Sentio ergo sum -- I feel, therefore I am.  There are several reasons for me to believe that someone else's feelings are real: I can see their reactions, they can tell me, and we know that humans have, for the most part, essentially the same neural apparatus.

Nonetheless I cannot know for sure what another person's feelings are in the same way that you and I could both put the same object on a balance scale and agree on its mass.  Each of the common-sense indications I just gave can fail.  Someone may not react visibly to a feeling or experience, or I may not catch the reaction.  They may not be able to tell me for any number of reasons.  Different people can have different ranges of feeling -- what seems intense to me might seem like nothing special to you, or vice-versa.

From a purely philosophical point of view we don't know for sure that having the same kind of neural pathway means having the same kinds of experiences.  Perhaps the ability to experience requires both a certain type of pathway and something else intangible that not everyone has.   Even if there is no such intangible, we're still far from knowing what physical pieces are associated with experience, though we do have some clues.  Without knowing just what pathways gives rise to subjective experience we have no way to be sure everyone has it.

When we go beyond human experience to other species, which react differently, can't verbalize their experiences (or at least not in ways we can presently understand), and have clearly different neural circuitry, we have even less to go on.  We can presume that a dog wagging its tail and barking when its human returns is happy, but it's always possible that dogs have simply co-evolved with us for long enough that they are able to act happy when that would be to their advantage (most people with dogs would dispute this, I expect).

Artificial constructs are even more problematic.  If I build a robot that avoids walls even if you push it toward one, it's easy to say "it doesn't like walls" because it's acting like a sentient being that disliked walls would, but it seems a much bigger step to say "it avoids walls because it experiences negative emotions when it's near one", particularly when we can point to the exact code that causes it to avoid walls.

Even if the code for the control system is extremely complex or has gone through some sort of machine learning process to develop an avoidance of walls, so that we couldn't point to exactly what was making it avoid walls, it still seems hard to argue that the robot is feeling emotions.  If incomprehensible code were the basis of sentience, there would be a lot of sentient software around.

When it comes to what we generally refer to as inanimate objects, the best we can say is that we have no reason to believe that a rock feels pain if we smash it with a hammer.  Nothing in our understanding of how we feel pain seems to apply to something like a rock.  Even so, how can we really know?


But how do we know anything?  We have no way of knowing whether we really live in a universe where the laws of physics hold.  It's possible that tomorrow things dropped will fall up instead of down.  Some theories of cosmology assign a non-zero (but still exceedingly small) chance that we live in such a universe.

In the absence of certain knowledge all we can do is try to build a coherent framework and constantly test and adjust the assumptions it rests on, a process we call "science".  From a scientific point of view we can figure out what sort of neural structures correspond with the subjective experiences that people report.  We can assess whether other organisms have such structures and even whether a particular combination of hardware and software has something functionally equivalent.

We can tell whether something's reactions to various stimuli are consistent with it having such capabilities, based on what people have reported.  We can conclude from that that it's likely or unlikely that the organism or construct we're examining is experiencing feelings, but we can never know for sure, no matter what philosophical machinery we develop for understanding qualia.

But this is nothing new.  Recently it was announced that gravitational waves had finally been detected, stemming from the collision of two black holes over a billion years ago.  The chain of inferences that rests on is mind-boggling.  A more accurate statement would have been "In two separate places, specially constructed instruments registered a signal that indicated that test masses had moved, over a distance much less than the size of an atom, in a way that indicated that space-time had been distorted in a way consistent with the collision of two black holes over a billion light-years away.  We feel confident about this because we believe that science works in general and we're convinced by a large web of observations and theoretical conclusions that the observable universe is billions of years old and billions of light-years in extent, black holes exist and, consistent with a distinct but overlapping web of observations and theoretical conclusions, in certain cases they should produce detectable gravitational waves.  We have also done extensive measurements to convince ourselves that the detectors are in fact detecting gravitational waves and not just trucks driving by ..."

And that would be the short version.  The full version fills textbooks and takes entire careers to grasp even a small portion of.

If science can accept that, can it come to accept that a dog is happy?

Not exactly.  The sticking point here is not whether we can accept a long chain of inference like "People report feeling happy when certain neurons are firing in certain ways, they behave in certain ways when this is happening, dogs have analogous neural pathways, and these tend to fire when dogs are engaged in behavior analogous to that of happy people, and/or people report that the dogs seem happy."  That's not a problem, particularly not compared to the detection of gravitational waves.

The problem is that science depends fundamentally on objective, repeatable measurements of numbers.  Happiness is subjective, and happiness is not a number.  Science can get quite close to measuring happiness, but it's up to us to decide where to go from there -- just like with any other scientific result.

Wednesday, August 10, 2016

Qualia, or why do we experience anything at all?

Today I'd like to discuss a topic which has baffled (at least some) philosophers for quite some time and which I am even more ill-qualified to address than usual.  Since I'm giving general impressions from general ignorance I'll be citing a few well-known examples without attribution.  You can find a good summary here, or it least it seemed like a good one to me.  Rest assured I'm not claiming to be doing any original work here, just ... conjecturing.

The term qualia has come to encompass experiences, and in particular subjective experiences.  For example, what is it like to see the color red, or what is it to be a bat.  Such experiences seem to be subjective, in that the experience depends, at least in principle, on who's experiencing it.  To take a very old example, cliche but no less valid for being cliche, I have no obvious way of knowing whether you experience the color red in the way I do.  Perhaps you experience it the way I experience the color blue, and vice versa, or perhaps you experience it some completely different way.

For that matter, how do I know that you experience anything?  If you and I are at an intersection, stopped at a red light, I can see you react to the light turning green, but that doesn't mean that you had the same experience I did of seeing a red light and then a green light.  I assume that you experienced the sensation of something red and then something green, and that the color red seemed essentially the same way to you as it did to me, but how would I know?

Suppose you were actually in a self-driving car browsing the news on your phone.  You didn't see the light at all.  Rather, the car's cameras recorded the light changing and the car's control system caused the car to go when the light turned green.  I'm perfectly comfortable saying "The car saw the light change and drove through the intersection when it turned green", anthropomorphizing the car, but that doesn't mean I think the car experienced the colors red and green in anything like the way you or I would (or at least, I think you would).

Trying to account for distinctions like this in some objective way has been referred to as "the hard problem of consciousness", as opposed to easier, more empirical problems like "How does the brain record memories?" or "To what extent are we conscious of our own decisions?"

In some sense it's quite likely that all experiences are distinct.  If I see a red paint chip today and then again tomorrow, I will almost certainly have different associations each time.  The first time might put me in mind of a stop sign, or blood, or a red apple.  The second time I might be more focused on whether it's the same paint chip I saw yesterday.  Likewise, you will almost certainly have different associations than I will even if we're looking at the same chip.

And yet, we would probably all agree that we are experiencing seeing something red, and that it feels like something to have that experience.   Even if there's no emotional response, you're still having some sort of experience.  How do we account for that?

Suppose we could account for every firing of every neuron in the nervous system (including the optic nerve, which is actually doing quite a bit of processing before the signal even gets to the brain).  Have we accounted for the experience?  Suppose that after decades of research we compile an exhaustive list of experiences and how they correlate to brain activity.  We bring in a new subject and scan their neural activity.  Pointing at a display, we say "That pattern of firing always occurs in response to seeing the color red".  We can say "that person is experiencing the color red", but how, exactly, do we know that for sure?

It's not hard to imagine what kind of data would back this up.  We hook hundreds of subjects from all over the world and all walks of life up to our highly-advanced brain scanner, flash colors at them and note the results.  We may even ask them to describe what they're experiencing.  When we see the same patterns for our new subject it's a reasonable inference that their brain is processing the color red, and it's reasonable to expect that if we ask them what they're experiencing, their answer will involve the color red.

That's probably good enough for a cognitive scientist, but not a philosopher.  The philosopher may well insist that you don't know what the subject experienced, but only how they would answer a question.  They -- and for that matter any of your other subjects -- might just as well be philosophical zombies who exhibit all the expected behaviors and responses without actually experiencing anything.  We may know intuitively, but we can't prove that the test subjects aren't just like the self-driving car, only on a more elaborate level.


There are a couple of ways out of this.  One is to deny that qualia exist in any well-defined way.  From a logical point of view, this seems quite plausible.  We can talk about the abstract concept of redness, but in real life we don't experience redness in the abstract.  We experience a particular something red at a particular place and time.  That feels a particular way at that place and time, and quite possibly nothing has ever felt quite the same before or ever will.  Maybe we should just stick to our knitting and figure out what happens in real brains in response to real stimuli.  We can still generalize and define abstractions, but if we want an objective description of the world we have to start with objective data.

And yet, we still experience things, subjectively, each of us (or at least I'm pretty sure about me).

So how do we distinguish between a person at a stop light and a self-driving car?  Maybe we don't need to make a strong distinction.  Maybe we're ... not so different.

There's no particular reason, beyond our innate sense of specialness, to assume that only human beings can have experiences.  If we see a hungry dog, our intuition tells us the dog is experiencing hunger.  Our intuition is probably right.  The dog may not be having exactly the same kind of experience we do, but there's no reason to assume it's a philosophical zombie that only looks like it's experiencing hunger.

One way of handling this is to assert that along with the physical properties of the world -- mass, position, velocity and so forth -- there is an experiential component that's completely distinct but which we might still be able to reason about.  Perhaps we will even discover laws that govern it and develop a comprehensive theory of experience.

One objection to this approach is that it seems to imply panpsychism, the idea that everything has consciousness.  There are already schools of thought that believe exactly that, but the concept doesn't sit particularly well in materialist circles (materialist in the philosophical sense).

However, this seems misguided.  If consciousness in the sense of being able to experience qualia is a property in a way similar to mass being a property of things, that doesn't mean that everything has to have that property.  Just as photons are massless, there's no contradiction in saying a rock is unconscious.

Rather than stating that everything has consciousness, we are asserting that objects can have consciousness, and we are trying to investigate under what circumstances that happens.  However, we are explicitly punting on the question of how it has consciousness.  We are saying that when the conditions are right "it just does", just as when a particle interacts with the Higgs field it has mass* (I believe physics has a more detailed account of this than "it just does", but at some point even physics has to make some base assumptions).

From that point of view it's still reasonable to say that a rock has no feelings or consciousness, but a human does, a dog does and just possibly a self-driving car has some limited degree of consciousness as well.  Moreover we may be able to prove that in the scientific sense of having a coherent theory and data to support it.  If so, it seems this theory will look a lot like a purely material explanation of memory, attention and other aspects of consciousness, together with an assertion that when certain of these are present, the thing in which they are present experiences qualia.

What is it to be a self-driving car?  Probably not much, but perhaps something.

* [That's not a really rigorous way to phrase that, but I don't know well enough to give a better one --D.H.]

Saturday, April 9, 2016

Primitives

Non sunt multiplicanda entia sine necessitate.

This is one of several formulations of Occam's razor, though Wikipedia informs us that William of Ockham didn't come up with that particular one.  Whatever its origins, Occam's razor comes up again and again in what we like to call "rational inquiry".  In modern science, for example, it's generally expressed along the lines of "Prefer the simplest explanation that fits the known facts".


If you see a broken glass on the floor, it's possible that someone took the glass into a neighboring county, painstakingly broke it into shards, sent the shards overseas by mail, and then had a friend bring them back on a plane and carefully place them on the floor in a plausible arrangement, but most likely the glass just fell and broke.  Only if someone showed you, say, video of the whole wild goose chase might you begin to consider the more complex scenario.

This preference for simple explanations is a major driving force in science.  On the one hand, it motivates a search for simpler explanations of known facts, for example Kepler's idea that planets move around the Sun in ellipses, rather than following circular orbits with epicyclets as Copernicus had held.  On the other hand, new facts can upset the apple cart and lead to a simple explanation giving way to a more complicated revision, for example the discoveries about the behavior of particles and light that eventually led to quantum theory.

But let's go back to the Latin up at the top.  Literally, it means "Entities are not to be multiplied without necessity," and, untangling that a bit, "Don't use more things than you have to", or, to paraphrase Strunk, "Omit needless things".  In the scientific world, the things in question are assumptions, but the same principle applies elsewhere.



The mathematical idea of Boolean Algebra underlies much of the computer science that ultimately powers the machinery that brings you this post to read.  In fact, many programming languages have a data type called "boolean" or something similar.

In the usual Boolean algebra, a value is always either True or False.  You can combine boolean values with several operators, particularly AND, OR and NOT, just as you can combine numbers with operations like multiplication, addition and negation.  These boolean operators mean about what you might think they mean, as we can describe them with truth tables very similar to ordinary multiplication or addition tables:

ANDTrueFalse
TrueTrueFalse
FalseFalseFalse

ORTrueFalse
TrueTrueTrue
FalseTrueFalse

NOTTrueFalse

FalseTrue

In other words, A AND B is true exactly when both A and B are true, A OR B is true whenever at least one of the two is true, and NOT A is true exactly when A is false.  Again, about what you might expect.

You can do a lot with just these simple parts.  You can prove things like "A AND NOT A" is always False (something and its opposite can't both be true) and "A OR NOT A" is always True (the "law of the excluded middle": either A or its opposite is true, that is, A is either true or false).

You can break any truth table for any number of variables down into AND, OR and NOT.  For example, if you prefer to say that "or" means "one or the other, but not both", you can define a truth table for "exclusive or" (XOR):

XORTrueFalse
TrueFalseTrue
FalseTrueFalse

If you look at where the True entries are, you can read off what that means in  terms AND, OR and NOT: There's a True where A is True and B is False, that is, A AND NOT B, and one where B is True and A is False, that is, B AND NOT A.  XOR is true when one or the other of those cases hold. They can't both hold at the same time, so it's safe to use ordinary OR to express this: A XOR B = (A AND NOT B) OR (B AND NOT A).  The same procedure works for any truth table.

In a situation like this, where we're expressing one concept in terms of others that we take to be more basic, we call the basic concepts "primitive" and the ones built up from them "derived".  In this case, AND, OR and NOT are our primitives and we derive XOR (or any other boolean function we like) from them.

Now consider the boolean function NAND, which is true exactly when AND is false.  Its truth table looks like this:

NANDTrueFalse
TrueFalseTrue
FalseTrueTrue

This is just the table for AND with True entries changed to False and vice versa.  That is, A NAND B = NOT (A AND B).

What's A NAND A?  If A is True, then we get True NAND True, which is False.  If A is False, we get False NAND False, which is True.  That is, A NAND A = NOT A.  If we have NAND, we don't need NOT.  We could just as well use AND, OR and NAND instead of AND, OR and NOT.

Since NAND is just NOT AND, and we can use NAND to make NOT, we don't need AND, either.  A AND B = (A NAND B) NAND (A NAND B).  So we can get by with just NAND and OR.

As it turns out, we never needed OR to begin with.  Quite some time ago, Augustus De Morgan pointed out that A OR B = NOT (NOT A AND NOT B), that is, A or B (or both) are true if both of them are not false, a rule which sometimes comes in handy in making "if" statements in code simpler (the other version of the rule, with the AND and OR switched, is also valid).  Using NAND, we can recast that as A OR B = (NOT A NAND NOT B), and we can get rid of the NOT, leaving A OR B = ((A NAND A) NAND (B NAND B)).

Summing up, we can build any boolean function at all out of AND, OR and NOT, and we can build all three of those out of NAND, so we can build any boolean function at all from NAND alone.

For example A XOR B = (A AND NOT B) OR (B AND NOT A).  We can use DeMorgan's rules to change that to (NOT (NOT (A AND NOT B) AND NOT (B AND NOT A))), that is, (NOT (A AND NOT B)) NAND (NOT (B AND NOT A)), or more simply, (A NAND NOT B) NAND (B NAND NOT A).  We can then replace the NOTs to get (A NAND (B NAND B)) NAND (B NAND (A NAND A)).

Yes, it's ... that ... simple.  Feel free to plug in all four combinations of A and B to check.

As silly as this may seem, it has real applications.  A particular kind of transistor lets current flow from its source terminal to its drain terminal when the voltage on a third terminal, called the gate, is high.  Put a high voltage on the source and let current flow either through the transistor or through an output for the whole thing.  Tie the gate of the transistor to an input.  If the voltage on the input is high, current will flow through the transistor and not to the output.  If the voltage on the input is low, current will flow to the output and not through the transistor.  That is, the output voltage will be high when the input voltage is not.  The whole thing is called a NOT gate (or inverter).

If you put two transistors in a row, then current will only flow through both of them when the voltage on both of the inputs is high, meaning it will flow through through the output, and the output voltage will be high, unless the voltage on both of the inputs is high.  The whole thing is called a NAND gate, and as we saw above, you can build any boolean function you like out of NANDs *.

In fact, we have it a bit easier here because we can build NOT A directly instead of as A NAND A, and for that matter we can build a three-input NAND  -- NOT (A AND B AND C) -- easily as well, but even if we couldn't, being able to build a NAND would be enough.



There are other cases where we can build up a whole system from a single primitive.  Notably, any computer program (technically, anything a Turing machine can compute) can be expressed in terms of a single instruction or, alternatively, a single "combinator".   This includes any boolean function, any numerical function, an HTML parser for web pages, whatever.  Of course, there's a difference between being able to express a computation in theory and being able to run it on your laptop or tablet.  We'll come back to that.

Before we go on to what all of this might mean, it's worth noting that many significant areas of thought haven't been reduced to simple primitives.  For example, chemistry is built from around a hundred elements (there are currently 118 on the periodic table, but you can't do meaningful chemistry with all of them).  An atom of any element is composed of protons, neutrons and electrons in varying numbers.

The Standard Model recognizes electrons as elementary, that is, primitive, while protons and neutrons are composed of quarks.  In all, it holds that there are 17 particles that everything is composed of  -- six quarks, six leptons (including the electron), four gauge bosons and the Higgs.  So far, no one has found anything simpler that these might be built up of, but not for lack of trying.

In mathematics, you typically start with a handful of axioms -- statements you assume to be true without proof -- and build from there.  There has been extensive work in reducing this to a minimal foundation, but the current formulation of set theory together with model theory has several important basic pieces, not one single concept to rule them all.  And, in fact, there are several ways of describing both set theory and model theory, not one single definitive way.

In short, some things can be reduced to a single primitive, but most can't.  Even when you can reduce something to a single primitive, there are typically several ways to do it.  For boolean algebra, you can just as well use NOR as NAND.  In computing there are several universal operations with little to pick among them.



In theory, there is no difference between theory and practice. But, in practice, there is. (attributed to Jan v/d Snepscheut)


If you can reduce any boolean function to NAND, is Boolean algebra in some meaningful sense really just NAND?  Is computing really just the study of a single instruction or operator?  If we're trying to reduce the number of entities involved, following Occam, is it not better to study a single operation than many?

I think most working mathematicians and computer scientists would answer "No" to all of the above.  A general Boolean algebra is a set of objects and operations that follow certain rules.  We noted above that A AND A = A.  In set theory, the union of a set with itself is that set, and there are other examples.  We would like to capture that common behavior somehow, and we do it by defining rules that hold for anything that behaves in the way we're interested in, that is, axioms.  In the case of Boolean algebra there are five (chase the link if you're interested).

It just so happens that in the simple case of True, False and the operations on them, NAND and NOR can be used to build all the others.  That's nice, but not essential.  It's of interest in building circuits out of transistors, but even then there's no requirement to build everything from one type of gate if you don't have to.  As noted above, it takes fewer transistors to build a NOT directly, and that's how real circuits are built.

Even from the point of view of Occam's razor, it's not clear that reducing everything to NAND is a good idea.  Yes, you have only one operation to deal with, but you can define one truth table just as easily as any other.  If you want to use XOR, it's simpler to define the truth table for it than to define the truth table for NAND and then define XOR in terms of it.

In computing, if you have the machinery to rigorously define one instruction or operation, you can define as many as you like with the same machinery.  It may be interesting or even useful in some situations that you can define some operations in terms of others, but it doesn't make the system more useful.  In practice, I don't care how many instructions the processor has.  I care very much if there's an easy way to talk to the network or write to a file, things which are not even mentioned in theoretical discussions of computing (nor should they be in most situations).

So why even bother?  Is reducing a system to a single primitive just an interesting academic exercise?  Not necessarily.  If you're trying to prove properties about circuits or programming systems in general, it can be useful to divide and conquer.  First prove, once and for all, that any circuit or program can be reduced to a single primitive, then prove all sorts of useful properties about systems using that primitive.  Since you only have a single operator or whatever to deal with, your proofs will be shorter.  Since you've proved your operator is universal, they'll be just as powerful.  Essentially you've said that when it comes to general properties of a system, adding new operators or whatever doesn't do anything interesting.

You don't have to reduce everything all the way to a single primitive for this to be useful.  If you can only reduce a system to five primitives, doing proofs using those five is still easier than doing proofs on an equivalent system with twenty.


In general, there's a tension between keeping a system minimal and making it easy to use.  A minimal system is easier to build and it's easier to be confident that it works properly.  A larger system is easier to use, as long as there's not too much to take in.  Typically there's a sweet spot somewhere in the middle.

There are sixteen possible boolean operators on two variables, and you can build any of them up from NAND (or NOR), but usually we focus on three of them: AND, OR and NOT.  These are enough to build anything else in a way that's straightforward to understand.  They also correspond fairly closely to familiar concepts.   In some useful sense, they minimize the number of things you have to deal with in normal cases, and William of Ockham can rest easy.

It doesn't only matter how many things you have.  It matters which things.




* As usual, there are a few more wrinkles to this.  You have to tie the output of the last (or only) transistor to ground for current to flow, you need a resistor next to the input, the output also needs to be tied to ground eventually, and so forth.

You may have noted that current is flowing through the two transistors in the NAND gate precisely when both inputs are high, that is, the current is flowing through them (and not to the output) when one gate voltage is high AND the other is.  You might think it would be simpler to build an AND than a NAND.  However,  the current will only flow if the drain of the last transistor is connected to a low voltage.   That voltage will be low regardless of what's happening on the gates.  To see a difference in voltages, we have to look at the voltage at the top, which will vary depending on whether current is flowing through the transistors (that's probably not too clear, but it's the best I can do).

Sunday, September 13, 2015

More on the invisible oceans, and their roundness

Previously, in speculating about what it would take for an inhabitant of one of the several subsurface oceans thought to exist in our solar system to discover that their world was round, I said
Figuring out that the world is round would be a significant accomplishment.  The major cues the Greeks used -- ships sinking below the horizon, lunar eclipses, the position of the noontime sun at different latitudes -- would not be available.  The most obvious route left is to actually circumnavigate the world.  And figure out that you did it.
Fortunately, I was smart enough to leave myself some wiggle room: "I'm very reluctant to say 'such and such would be impossible because ...'"  After all, humanity is in a similar situation in trying to figure out the shape of our universe, though in our case circumnavigating doesn't seem to be even remotely close to an option.  Even so, we've had some apparent success.

On further reflection, there are at least two ways an intelligent species living in a world like Ganymede's subsurface ocean could figure out that the world is round.

First, and perhaps most likely, it's not necessary for a Ganymedean to circumnavigate the world, if something can.  In this case that something would be sound.  Under the right conditions, sound can travel thousands of kilometers underwater.  The circumference of Ganymede is around 10,000km, which may be too far.  Europa's is around 6,000km, which as I understand it is on the edge of what experiments in Earth's oceans have been able to detect (pretty impressive, that).

We've already postulated that sound would be one of the more important senses in such a world.  It doesn't seem impossible that someone would notice that especially loud noises tended to be followed a couple of hours later by similar noises from all directions.  This assumes that the the oceans are unobstructed, but if they're not, you have landmarks to measure by, which would make the original "circumnavigate and tell that you did it" less of a challenge.

Second, if some sort of light-production and light-detection evolve, then one of the cues the Greeks used is indeed available, at least in theory: one can see things disappear over the horizon.  To actually make use of this one would need to be able to see far enough to tell that the object was disappearing due to curvature and not behind some obstacle, or simply because it was too far away to see.  The exact details depend on how smooth the inner surface is.

Humans noticed the effect with ships because a calm sea is quite flat, that is to say, quite close to perfectly round.  If the inner surface of the ocean is rough, one might have to float at a considerable distance from it, and thus wait for the object to recede a considerable distance, to be sure of the effect.  On the other hand, floating a considerable distance from the inner surface would be much easier than floating the same distance from the surface of the earth.  For that matter, it would also be possible to note what's visible and what's not at different distances from the inner surface.

A little back-of-the-envelope estimation suggests that one would have to be able to see objects kilometers or tens of kilometers away.  By way of comparison, Earth's oceans are quite dark at a depth of one kilometer, so this seems like a longshot.  Nor does it help that it's possible to hear long distances, since sound doesn't necessarily propagate in a straight line (neither does light, but that's a different can of worms).


As I originally disclaimed, it's not a good idea to rule something out as impossible just because you can't think of a way to do it.  The inhabitants of a subsurface ocean would have thousands, if not millions, of years to figure things out, even if they wouldn't have the advantage of already knowing approximately what their world looks like.

Tuesday, August 18, 2015

What, if anything, is a grammar?

This is going to be one of the more technical posts.  As always, I hope it will be worth wading through the details.

There are hundreds, if not thousands, of computer languages.  You may have heard of ones like Java, C#, C, C++, ECMAScript (aka JavaScript, sort of), XML, HTML (and various other *ML, not to mention ML, which is totally different), Scheme, Python, Ruby, LISP, COBOL, FORTRAN etc., but there are many, many others.  They all have one thing in common: they all have precise grammars in a certain technical sense: A set of rules for generating all, and only, the legal programs in the given language.

This is true even if no one ever wrote the grammar down.  If you can actually use a language, there is an implementation that takes code and tells a computer to follow its instructions.  That implementation will accept some set of inputs and reject some other set of inputs, meaning it encodes the grammar of the language.

Fine print, because I can't help myself: I claim that if you don't have at least a formal grammar or an implementation, you don't really have a language.  You could define a language that, say, accepts all possible inputs, even if it doesn't do anything useful with them.  I'd go math major on you then and claim that the set of "other code" it rejects is empty, but a set nonetheless.  I'd also argue that if two different compilers are nominally for the same language but behave differently, not that that would ever happen, you have two variants of the same language, or technically,  two different grammars.

Formal grammars are great in the world of programming languages.  They tell implementers what to implement.  They tell programmers what's legal and what's not, so you don't get into (as many) fights with the compiler.  Writing a formal grammar helps designers flush out issues that they might not have thought of.  There are dissertations written on programming language grammars.  There are even tools that can take a grammar and produce a "parser", which is code that will tell you if a particular piece of text is legal, and if it is, how it breaks down into the parts the grammar specifies.

Here's a classic example of a formal grammar:
  • expr ::= term | term '+' term
  • term ::= factor | factor '*' factor
  • factor ::= variable | number | '(' expr ')'
In English, this would be
  • an expression is either a term by itself or a term, a plus sign, and another term
  • a term is either a factor by itself or a factor, a multiplication sign ('*') and another factor
  • a factor is either a variable, a number or an expression in parentheses
Strictly speaking, you'd have to give precise definitions of variable and number, but you get the idea.

Let's look at the expression (x + 4)*(6*y + 3).  Applying the grammar, we get
  • the whole thing is a term
  • that term is the product of two factors, namely (x + 4) and (6*y + 3)
  • the first factor is an expression, namely x+4, in parentheses
  • that expression is the sum of two terms, namely x and 4
  • x is a factor by itself, namely the variable x
  • 4 is a factor by itself, namely the number 4
  • 6*y + 3 is the sum of two terms, namely 6*y and 3
  • and so forth
One thing worth noting is that, because terms are defined as the product of factors, 6*y + 3 is automatically interpreted the way it should be, with the multiplication happening first.  If the first two rules had been switched, it would have been interpreted as having the addition happening first.  This sort of thing makes grammars particularly useful for designing computer languages, as they give precise rules for resolving ambiguities.  For the same reason they also make it slightly tricky to make sure they're defining what you think they're defining, but there are known tools and techniques for dealing with that, at least when it comes to computer languages.


The grammar I gave above is technically a particular kind of grammar, namely a context-free grammar, which is a particular kind of phrase structure grammar.  There are several other kind.  One way to break them down is the Chomsky hierarchy, which defines a series of types of grammar, each more powerful than the last in the sense that it can recognize more kinds of legal input.  A context-free grammar is somewhere in the middle.  It can define most of the rules for real programming languages, but typically you'll need to add rules like "you have to say what kind of variable x is before you can use it" that won't fit in the grammar itself.

At the top of the Chomsky hierarchy is a kind of grammar that is as powerful as a Turing machine, meaning that if you can't specify it with such a grammar, no computer can compute whether or not a given sentence is legal according to the grammar or not.  From a computing/AI point of view (and even from a theoretical point of view), it sure would be nice if real languages could be defined by grammars somewhere in the Chomsky hierarchy.  As a result, decades of research have been put into finding such grammars for natural languages.

Without great success.

At first glance, it seems like a context-free grammar, or something like it, would be great for the job.  The formal grammars in the Chomsky hierarchy are basically more rigorous versions of the grammars that were taught, literally, in grammar schools for centuries.  Here's a sketch of a simple formal grammar for English:
  • S ::= NP VP (a sentence is a noun phrase followed by a verb phrase)
  • NP ::= N | Adj NP (a noun phrase is a noun or an adjective followed by a noun phrase)
  • VP ::= V | VP Adv (a verb phrase is a verb or a verb phrase followed by an adverb)
Put this together with a list of nouns, verbs, adjectives and adverbs, and voila: English sentences.  For example, if ideas is a noun, colorless and green are adjectives, sleep is a verb and furiously is an adverb, then we can analyze Colorless green ideas sleep furiously with this grammar just like we could analyze (x + 4)*(6*y + 3) above. 

Granted, this isn't a very good grammar as it stands.  It doesn't know how to say A colorless green idea sleeps furiously or I think that colorless green ideas sleep furiously or any of a number of other perfectly good sentences, but it's not hard to imagine adding rules to cover cases like these.

Likewise, there are already well-studied notions of subject-verb agreement, dependent clauses and such.  Translating them into formal rules doesn't seem like too big a step.  If we can just add all the rules, we should be able to build a grammar that will describe "all and only" the grammatical sentences in English, or whatever other language we're analyzing.

Except, what do we mean by "grammatical"?

In the world of formal grammar, "grammatical" has a clear meaning: If the rules will generate a sentence, it's grammatical.  Otherwise it's not.

When it comes to people actually using language, however, things aren't quite so clear.  What's OK and what's not depends on context, whom you're talking to, who's talking and all manner of other factors.  Since people use language to communicate in the real world, over noisy channels and often with some uncertainty over what the speaker is trying to say and particularly over how the listener will interpret it, it's not surprising that we allow a certain amount of leeway.

You may have been taught that constructs like ain't gonna and it don't are ungrammatical, and you may not say them yourself, but if someone says them to you, you'll still know what they mean.  If a foreign-born speaker says something that makes more sense in their language than yours, something like I will now wash myself the hands, you can still be pretty certain what they're trying to say.  Even if someone says something that seems like nonsense, say That's a really call three, but there's a lot of noise, you'll probably assume they meant something else, maybe That's a really tall tree, if they're pointing at a tall tree.

In short, there probably isn't any such thing as "grammatical" vs. "ungrammatical" in practical use, particularly once you get away from the notion that "grammatical" means "like I was taught in school" -- and as far as that goes, different schools teach different rules and there is plenty of room for interpretation in any particular set of rules.  To capture such uncertainty, linguistics recognizes a distinction between competence (one's knowledge of language) and performance (how people actually use language), though not all linguists recognize a sharp distinction between the two.

When there is variation among individuals and over time, science tends to take a statistical view.  It's often possible to say very precisely what's going on statistically even if it's impossible to tell what will happen in any particular instance.  When there is communication in the presence of noise and uncertainty, information theory can be a useful tool.  It doesn't seem outrageous that a grammar aimed at describing real language would take a statistical approach.  Whatever approach it takes, it should at least provide a general account of what the analysis meant in information theoretic terms.

[I didn't elaborate on this in the original post, but as I understand it, Chomsky has strongly disputed both of these notions.  Peter Norvig quotes Chomsky as saying "But it must be recognized that the notion of probability of a sentence is an entirely useless one, under any known interpretation of this term."  This is certainly forceful, but not convincing, even after following the link to the original context of the quote, which contains such interesting assertions as "On empirical grounds, the probability of my producing some given sentence of English [...] is indistinguishable from the probability of my producing some given sentence of Japanese.  Introduction of the notion of "probability relative to a situation" changes nothing."  This must be missing something, somewhere.  If I'm a native English speaker and my hair is on fire, it's much more likely that I'm about to say "My hair is on fire!" than "閑けさや 岩にしみいる 蝉の声" or even "My, what lovely weather we're having".

The entire article by Norvig is well worth reading.  But I digress.]

Even if you buy into the idea of precise, hard-and-fast rules describing a language, there are several well-known reasons to think that there is more to the picture than what formal grammars describe.  For example, consider these two sentences:
  • The horse raced past the barn fell.
  • The letter sent to the president disappeared.
Formally, these sentences are essentially identical, but chances are you had more trouble understanding the first one, because raced can be either transitive, as in I raced the horse past the barn (in the sense of "I was riding the horse in a race"), or intransitive, as in The horse raced past the barn (the horse was moving quickly as it passed the barn -- there's also "I was in a race against the horse", but never mind).  On reading The horse raced, it's natural to think of the second sense, but the whole sentence only makes sense if you use the first sense of raced, but in passive voice.

Sentences like The horse raced past the barn fell are called "garden path" sentences, as they "lead you down the garden path" to an incorrect analysis and you then have to backtrack to figure them out.

A purely formal grammar describes only the structure of sentences, not any mechanism for parsing them, that is, recovering the structure of the sentence from its words.  There's good experimental data, though, to suggest that we parse The horse raced past the barn fell differently from The letter sent to the president disappeared.

A purely formal grammar doesn't place any practical limits on the sentence that it generates.  From a formal point of view,  The man whom the woman whom the dog that the cat that walked by the fence that was painted blue scratched barked at waved to was tall is a perfectly good sentence, and would still be a perfectly good sentence if we replaced the fence that was painted blue with the fence that went around the yard that surrounded the house that was painted blue, giving The man whom the woman whom the dog that the cat that walked by the fence that went around the yard that surrounded the house that was painted blue scratched barked at waved to was tall.

You could indeed diagram such a sentence, breaking it down into its parts, but I doubt very many people would say they understood it if it were spoken to them at a normal rate.  If you asked "Was it the house or the fence that was painted blue?", you might not get an answer better than guessing (I'm sure such experiments have been done, but I don't have any data handy).  Again, the mechanism used for understanding the sentence places limits on what people will say in real life.

In fact, while most languages, if not all, allow for dependent clauses like that was painted blue, they're used relatively rarely.  Nested clauses like the previous example are even rarer.  I recall recently standing in a fast food restaurant looking at the menu, signs, ads, warnings and so forth and realizing that a large portion of what I saw wasn't even composed of complete sentences, much less complex sentences with dependent clauses: Free side of fries with milkshake ... no climbing outside of play structure ... now hiring.  Only when I looked at the fine-print signs on the drink machine and elsewhere did I see mostly "grammatical" text, and even that was in legalese and clearly meant to be easily ignored.

In short, there are a number of practical difficulties in applying the concepts of formal grammar to actual language.  In the broadest sense, there isn't any theoretical difficulty.  Formal grammars can describe anything computable, and there's no reason to believe that natural language processing is uncomputable, even if we haven't had great success with this or that particular approach.  The problems come in when you try to apply that theory to actual language use.


Traditionally, the study of linguistics has been split into several sub-disciplines, including
  • phonology, the study of how we use sounds in language, for example why we variously use the sounds -s, -z and -ez to represent plurals and similar endings in English (e.g., catsbeds and dishes)
  • morphology, the study of the forms of words, for example how words is composed of word and a plural marker -S.
  • syntax, the study of how words fit together in sentences.  Formal grammars are a tool for analyzing syntax
  • semantics, the study of how the meaning of a sentence can be derived from its syntax and morphology
  • pragmatics, the study of how language is actually used -- are we telling someone to do something? asking for information? announcing our presence? shouting curses at the universe?
It's not a given that actual language respects these boundaries.  Fundamentally, we use language pragmatically, sometimes to convey fine shades of meaning but sometimes in very basic ways.  We tolerate significant variation in syntax and morphology.  It's somewhat remarkable that we recognize a distinction between words and parts of words at all.

In fact, it's somewhat remarkable that we recognize discrete words.  On a sonogram of normal speech, it's not at all obvious where words start and end.  Clearly our speech-processing system knows something that our visual system -- which is itself capable of many remarkable things -- can't pick out of a sonogram.

Nonetheless, we can in fact recognize words and parts of words.  We know that walk, walks, walking and walked are all forms of the same word and that -s, -ing, and -ed are word-parts that we can use on other words.  If someone says to you I just merped, you can easily reply What do you mean, "merp"? What's merping?  Who merps?

This ability to break down flowing speech into words and words (spoken or written) into parts appears to be universal across languages, so it's reasonable to say "Morphology is a thing" and take it as a starting point.  Syntax, then, is the study of how we put morphological pieces together, and a grammar is a set of rules for describing that process.

The trickier question is where to stop.  Consider these two sentences:
  • I eat pasta with cheese
  • I eat pasta with a fork
Certainly these mean different things.  In the first sentence, with modifies pasta, as in Do they have pasta with cheese on the menu here? or It was the pasta with cheese that set the tone for the evening.  In the second sentence, with modifies eat, as in I know how to eat with a fork or Eating with a fork is a true sign of sophistication.  The only difference is what follows with.  The indefinite article a doesn't provide any real clue.  You can also say I eat pasta with a fine Chianti.

The real difference is semantic, and you don't know which case you have until the very last word of the sentence.  A fork is an instrument used for eating.  Cheese isn't.  If the job of syntax is to say what words relate to what, it would appear that it needs some knowledge of semantics to do its job.

There are several possible ways around problems like this:
  • Introduce new categories of word to capture whatever semantic information is needed in understanding the syntax.  Along with being a noun, fork can also be an instrument or even an eating instrument.  We can then add syntactic rules to the effect that with <instrument> modifies verbs while with <non-instrument> modifies nouns.  Unfortunately, people don't seem to respect these categories very well.  If someone has just described a method of folding a slice of cheese and using it to scoop pasta, then cheese is an eating instrument.  If we're describing someone with bizarre eating habits, a fork could be just one more topping for one's pasta.
  • Pick one or the other option as the syntactic structure, and leave it up to semantics to do with it as it will.  In this case:
    • Say that, syntactically, with modifies pasta, or more generally the noun, but that this analysis may be modified by the semantic interpretation -- if with is associated with an instrument, it's really modifying the verb.
    • Say that, syntactically, with modifies eat, or more generally the verb, and interpret eat with cheese as meaning something like "the pasta is covered with cheese when you eat it".  This may not be as weird as it sounds.  What else do you eat with cheese? is a perfectly good question, though you can argue that in this case the object of eat is what, and with cheese modifies what.
  • Make syntax indeterminate.  Instead of saying that with ... modifies eat or pasta, say that  syntactically it modifies "eat or pasta, depending on the semantic interpretation".  This provides a somewhat cleaner separation between syntax and semantics, since the semantics only has to decide what's being modified without knowing the exact structure of the sentence it was in.
  • Say that with modifies eat pasta as a unit, and leave it up to semantics to slice things more finely.  This is much the same as the previous option.  Technically it gives a definite syntactic interpretation, but only by punting on the seemingly syntactic question of what modifies what.
  • Say that the distinction between syntax and semantics is artificial, and in real life we use both kinds of structure -- how words are arranged and what they mean -- simultaneously to make sense of language.  And while we're at it, the distinction between semantics and pragmatics may turn out not to be that sharp either.
There are two reasons one might want to separate syntax from semantics (and likewise for the other distinctions):
  • They're fundamentally separate.  Perhaps one portion of the brain's speech circuitry handles syntax and another semantics.
  • On the other hand, separating the two may just be a convenient way of describing the world.  One set of phenomena is more easily described with tools like grammars, while another set needs some other sort of tools.  Perhaps they're all aspects of the same thing deep down, but this is the best explanation we have so far.
As much as we've learned about language, it's striking, though not necessarily surprising, that fundamental questions like "What is the proper role of syntax as opposed to semantics?" remain open.  However, it seems hard to answer the question "What, if anything, is a grammar" without answering the deeper question.

[If you're finding yourself asking "But what about all the work being done in parsing natural languages?" or "But what about dependency grammars?", please see this followup]