Thursday, August 13, 2009

On putting faith in models

Here's a conversation I had with my brother via g-chat the other day. It's a great prototype of a conversation I've had many times in the last few months. The basic question is "how much faith can we place in mathematical models?" Most people seem skeptical; I'm more of a believer.

This particular exchange was unusual because 1) it was conveniently recorded, and 2) it went in some interesting and productive directions at the end. I'm posting it unedited, except for some spelling fixes and external links. Comments welcome.


10:44 AM Sam: this is probably trivial compared to the analysis you usually look at, but i thought you might like it nonetheless
me: I'll check it out


6 minutes
10:51 AM me: Interesting stuff
I hadn't seen the wine article before
10:52 AM I'd read the paper on war, but I hadn't seen the TED talk
10:54 AM This is exactly the kind of stuff I'm interested in doing
Sam: have you read when genius failed?
me: no
Sam: i can't remember if we talked about it already
it's about a hedgefund
10:55 AM and their story
it's the classic cautionary tale of putting too much faith in models
me: oh, yes
I haven't read it, but I know the gist
10:56 AM I would change the interpretation a little bit and say putting too much faith in a theory.
A model is a theory that happens to be expressed mathematically
Like any theory, if your assumptions are bad your conclusions will be bad
10:57 AM Sam: it was all mathematical
they determined 'actual' risk spreads based on reams of historical data and market conditions
10:58 AM and then tried to beat the market by playing the spreads dictated by their systems
me: right
as I understand it, the mistake they made was in the way they calculated risk
10:59 AM Sam: this is an interesting meta-argument
because you're pointing to the specific problem of their models
where i'm saying their mistake was over reliance on models
me: yes
11:00 AM this is a live debate in social science
I'm a pretty strong advocate for the quant side
11:01 AM Sam: hm
me: I'd say the key question is "is there anything substantively different about theories expressed in math verses theories expressed in English?"
Sam: this is probably a classic
me: classic?
11:02 AM Sam: academics vs business
11:03 AM me: hm
maybe
it's not so clean cut, though, because there are academics who reject the quant stuff and businessmen who embrace it
11:04 AM I think it has to do with the way people think about math
Is it a set of fixed processes for getting answers, or is it a language for expressing ideas?
11:05 AM If it's a language, then the fault for bad models lies is the assumptions expressed, not the language for expressing them.

5 minutes
11:10 AM Sam: accepting that an omniscient agent could express all ideas as formulas and believing that you can are different, right?
11:11 AM its that leap where you decide to stake a business and millions of dollars on the formula that puts you on one side of the line or the other, in my opinion
me: yes, fair points
11:13 AM It seems to me that you're introducing another aspect of theories, which is that they don't just have to expressed, they also have to be acted upon
and strange things can happen when you act on a theory without fully understanding it
It's kind of a Jurassic Park idea
11:14 AM Sam: ha
i guess ultimately the moral to when genius is the same as jurassic park
11:15 AM but trust funds drying up is less dramatic than dino-carnage
me: so maybe the problem with models isn't that they are more likely to be wrong, but that they invite careless extrapolation
Sam: i don't think it's careless
it's hubris
me: "I'm going to leverage a billion dollars 30 times"
11:16 AM "I'm going to bring back dinosaurs from the dead, focusing mainly on intelligent top predators"
Sam: it's the opposite of being careless- it's spending so much time working on a model that you believe that you can and have thought of everything
11:17 AM its validating that model again and again against the datasets you have without allowing for future events to be unknowable
11:18 AM me: I agree -- my only reservation would be that hubris is not specific to people who frame their models using math
Sam: of course
the opposite story is much more common
which is why when genius failed is a story worth telling
11:19 AM when hubris failed would be too common
11:21 AM me: so instead of "people fail because of math," it's "even people using math can fail because of hubris"
11:22 AM Sam: yeah, that's the gist
me: okay, I don't have to feel so defensive now
11:23 AM a lot of the people around me here are model-builders
some of the faculty are among the best in the world
I see both types
11:24 AM Some use models for the sake of transparency -- all the assumptions are laid out for criticism and improvement
Others use models for the sake of beating up on people who don't speak game theory
11:25 AM I'm pretty invested in the notion that models can help the process of collective learning, and it frustrates me when the arrogant ones give modeling a bad name

Thursday, August 6, 2009

Big news stories in the 2008 elections -- Looking for you input

I have a ~10 minute favor to ask. It has to do with timelines again*.

I've pulled a list of the top ~200 events from the run-up to last year's presidential elections. (Download it here as an .xlsx file, here as an .xls file, and here as a .txt file) Sometime when you weren't doing anything anyway (e.g. facebook), skim through the list and pick the top 20 or so events that you think were the biggest stories* in the campaign. If there are important stories that you think are missing, you can add them.

When you're done, post your results in the comments section. Here are the rules:
  • By "big news stories," I mean events that did at least one of three things: 1) generated a lot of media buzz, 2) affected public opinion, or 3) evoked a strong response from the campaigns themselves. Any combination of these things qualifies as a news story.
  • Don't do any background research and don't ask anybody else for their opinion. I'm just looking for a gut check on which events were the most important.
  • Don't worry if you aren't a big-time pundit. I'm not either, and they're mostly bluffing anyway.
  • And don't don't don't read other peoples' responses before you post your own -- this will be much more useful and interesting if everyone's ideas are independent.
I'm going to use these responses (anonymously) at a conference coming up in a few weeks. Once the conference is over I'll post some diagnostics and results here so you can see how your intuition stacks up against the wisdom of the crowd.

Like I said, this should only take about 10 minutes or so. Thanks for being part of a convenience sample of the willing!

Some background:
I'm working on a research project using automated content analysis to identify big news stories in archives of media content. The 2008 presidential election is my test case. Basically, I'm throwing a lot of text at the computer and using tricks from computational linguistics to tease out the news stories. It would be neat to be able to do this because 1) news stories are a big part of the way we think about politics, 2) this would make it possible to identify news stories in a replicable way on a grand scale, and 3) complicated algorithms are cool.

Your answers will help me construct a baseline to check how well the algorithm is doing. If the computer finds events that are broadly consistent with human intuition, that's a good sign that it's working. Thanks again for your help.


*PS on my previous post: Chronologic turned out harder than I initially thought! I'm working on ways to weed out some of the ridiculously obscure cards.)

Friday, July 24, 2009

Playtesters needed for a new and improved trivia game

Playing trivia games has always reminded me of listening to a badly scratched CD, or the worst parts of taking the SAT. All three experiences serve up lots of disconnected fragments of something that ought to be big and meaningful, but ends up just coming across as frustrating. Trivia games are history at its reductionist, one-fact-after-another worst.

There are two main symptoms of this fragmentation. First, for any given question, you either know the answer or you don't. Who was Speaker of the House in 1810? Beats me. And if you don't know, there's nothing else to discuss. You bubble in your guess and move on.

Second, there's no element of collaboration or coordination within teams. Trivia teams don't really work together. You just hope your teammates know the stuff that you don't. Practically speaking, this means that most players are not participating most of the time during most trivia games.

Onwards and Upwards
So instead of just griping, I decided to strike back at bad trivia games and do something about it.

I started by writing a web spider to crawl over wikipedia's list of historical anniversaries and pull out dates and descriptions for events in history. For example, according to wikipedia, on May 10, 1801, "the Barbary pirates of Tripoli declare[d] war on the United States of America."

I cached about 15,000 such events in an xml document. Next, I did some text processing to clean up unecessary tags, links, etc. and used XSL-FO scripting to format the events as printable cards. This is all technobabble -- I just want you to appreciate how hard/nifty it all was.

The upshot is that I've created a deck of random historical events, suitable for playing trivia games. They range from the well-known (July 4, 1776 -American Revolution: The Declaration of Independence is adopted by the Second Contintental Congress) to the hopelessly obscure (Nov. 14, 1923 - Kentaro Suzuki completes his ascent of Mount Iizuna). The cards are in .pdf format -- just print and cut along the dotted lines. Depending on your printer, you might want the .pdfs with separate fronts and backs, or you might want the every-other-page version.

Of course, cards alone do not make the game. Here's my attempt to improve on the obnoxious fragmentation of trivia games. The rules are based loosely on an older game called Chronology, but trust me, they're an improvement. I call the new game Chronologic.



Chronologic Rules:
Object: As a team, score the most points by placing events from history in sequence.

Setup: Divide into teams. A few (2-5) teams of a few (1-4) players are good. Choose a target score -- the score where the game will end. For me, 100 is a good target for a ~30 minute game. Decide on handicaps and time limits if necessary. Place the event cards text-side up somewhere where they are easy to get at.

Game play: Play proceeds in rounds. In each round, each team constructs and scores its own timeline.

Building timelines: During each round, your team will construct a timeline by drawing event cards and placing them in sequence one at a time, until you decide to pass. Don't look at the backs of any of the cards until you move to the socring phase! The first event is easy to place in sequence -- there are no other events, so you get it right by default. For every subsequent event, you must decide exactly where it falls in the sequence of events already on the table. If your team doesn't know and doesn't want to guess where a given event belongs, you can pass. Once you pass, you are done for the round.

Every team constructs its own timeline, but they do so simultaneously. So all teams draw their first card together. Once those cards are played, they draw their second cards at the same time, and so on. If your team has already passed during a given round, wait for the other teams to finish their timelines. When all teams have passed, proceed to scoring.

Scoring timelines: Once you finish building your timeline, you get to score it. Timelines are an all-or-nothing proposition. If all of the events in your timeline are in order, your team scores the number of events in the line, squared. If a single event is out of order, your team gets 0 points. So if you have 3 events all in order, you score 9 points. If you have 4 events all in order, you score 16. If you have 4 events but one of them is out of order, you score 0. If you're having trouble remembering your square numbers, they go 1, 4, 9, 16, 25, 36, 49, 64, 81, 100, and up from there.

After scoring the timeline, add each team's points to their running total. Discard all the cards in all of the timelines. Move on to the next round.

The last round: Once a team reaches the target score, the game goes into a final round. Every team except the one that reached the target gets one last chance to score some points. The team with the most points after the final round wins.



That's all, except for the obvious: discussion, synthesizing sidetracks, and fact-finding trips to wikipedia in the course of the game (after timelines are scored) are encouraged. The whole point is to make sense of how things connect.

I've had a lot of fun getting this game ready for play testing. I'm hoping you'll enjoy it, and I'm open to feedback and suggestions for improvements, extensions, and whatnot.

Wednesday, March 25, 2009

A political puzzle...

Following a hunch, I pulled together historical data on US political participation. Here's the result:


The turnout data is from data is from wikipedia. Non-presidential election years are averages for the previous and following election. Polarization data comes from the difference between House party medians in the first two DW-NOMINATE dimensions. Basically, polarization measures how much Representatives' votes broke along party lines.

Here's the puzzle: up until around 1900, the two lines track quite closely. The raw values are about the same, and they tend to move up and down together. But after 1920, the polarization and participation stop tracking each other. Why?

Saturday, March 21, 2009

Tracking campaign contributions

Here's a site trying to make campaign contributions more transparent:
"MAPLight.org, a groundbreaking public database, illuminates the connection between campaign donations and legislative votes in unprecedented ways. Elected officials collect large sums of money to run their campaigns, and they often pay back campaign contributors with special access and favorable laws. This common practice is contrary to the public interest, yet legal. MAPLight.org makes money/vote connections transparent, to help citizens hold their legislators accountable."
I haven't dug into the site in much detail, but it looks interesting.  Anybody know more?