Here’s a thought-provoking piece on how rural America could benefit from the outsourcing trend if we just get on the stick. There are two pieces to this; one is getting high-speed connections out into the hinterlands and the other is to provide the people with the training required to enable them to make a living as knowledge workers. (A third might be developing and promoting business models that are friendly to a relatively distributed rural workforce.) Naturally, online learning could play a big part in making this happen by bringing the educational mountain to Mohammed, so to speak.
Author: Michael Feldstein
-
Correction on the Origins of Informational Cascade Research
I was mistaken in an earlier post when I claimed that informational cascades research comes from the “heuristics and biases approach” in psychology. It definitely comes from behavioral economics.
Both behavioral economics and the heuristics and biases approach share common ancestry from the work of Herbert Simon. A genuine polymath, Simon won a Nobel Prize in economics for his theory of bounded rationality while also helping to give birth to modern cognitive psychology and making major contributions to the field of artificial intelligence. In the case of bounded rationality, Simon’s insight is that we cannot realistically expect that the human mind works perfectly rationally without regard to computational cost. Some rational calculations just take too much time and attention to be useful in real-world decision-making situations. Simon suggested that humans use mental short cuts that are good enough to get the job done most of the time and take far less time and energy than fully rational calculations.
Both behavioral economics and the heuristics and biases approach draw heavily on this insight. However, the specific work on informational cascades definitely came from the economists. The earliest references to them were in the following papers:
- Abhijit V. Banerjee. August 1992. A Simple Model of Herd Behavior. Quarterly Journal of Economics 107:3, 797-818
- Sushil Bikhchandani, David Hirshleifer, Ivo Welch. October 1992. A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades. Journal of Political Economy 100:5, 992-1026
- Ivo Welch. June 1992. Sequential Sales, Learning, and Cascades. Journal of Finance 47:2, 695-732
For more info on the informational cascades literature, see this bibliography. For more info on the heuristics and biases approach, consider reading Heuristics and Biases: The Psychology of Intuitive Judgment
-
Small World Indeed
This is apropos of nothing but it’s such a great story that I have to share it. Today in her ESL class, Kathy paired up her students for an icebreaker writing assignment. And, as often happens, there were two stragglers left over (one young woman and one young man, in this case, sitting on opposite sides of the room from each other) who didn’t feel comfortable taking the initiative to pair with anyone. So Kathy put them together.
Which was a good thing. As it turned out, even though the two didn’t know each other, they happen to have grown up in the same home town.
In China.
-
How to Recognize Competence in a Software Engineer (Or Any Other Professional)
(Warning: Techno-babble ahead. I’ll keep it limited to the first paragraph and translate it at the end.)
The other day I stumbled across this InfoWorldarticle that points to two possible informational cascades in the IT world. First, it touches on the “PHP doesn’t scale” meme. I’ve seen this point argued ad nauseam in many places, often without any evidence. It’s just something that “everybody knows.” And whenever everybody knows something, you can bet there’s an informational cascade involved. The second potential cascade it touches on is “three-tiered architecture is the best.” Remember when everyone remotely connected with IT would ramble on incessantly about “middleware”? That everything should be done using application servers? Well here’s an article about how Friendster canned their Java system for a PHP architecture where most of the logic lives in the database and got far better performance and scalability out of the deal.
(Non-techie translation: the company built a website that could handle much more traffic by ignoring two points of conventional wisdom in the geek world.)
This brings to mind the many painful experiences I have had trying to find programmers who could give me competent advice. It’s hard to identify real expertise in a domain that you don’t know a whole lot about. But one easy way to weed out the dumb ones was to ask a simple question: How do you know? The mediocre programmers can’t answer this question. “Oh yeah, Java is definitely the way to go.” “How do you know?” “Well, it’s much more scalable.” “But how do you know that it’s much more scalable?” “Uh…well…PHP is well known to have scalability problems. Java, on the other hand, is used everywhere.”
Bzzt. Wrong answer.
What I want to hear is, “Well, when I tried using Java for X, here’s what happened….” Or “I’ve read Y about PHP but I don’t have much personal experience with it.” Or “The argument I’ve heard against PHP is Z, which makes some sense to me because….” In other words, I’m looking for somebody who knows the difference between what economists would call private information (what you and I might call personal experience or other direct knowledge) and what “everybody knows.” A good software engineer knows that a point of “common knowledge” is just another hypothesis that needs to be tested. S/he also knows whether s/he has had any direct experiences that confirm or disconfirm the hypothesis. S/he knows the difference between what s/he knows and what s/he doesn’t know.
This, of course, is true of any good knowledge worker. Truly competent knowledge workers harbor a healthy skepticism of common knowledge. They use it when it’s convenient (and often it’s the only information we have to go on when we make our decisions) but they don’t trust it more than they have to.
-
Tracking Memes in the Wild, Part III
In my last two posts, I wrote about the limitations of one method for tracking memes and the promise of second method. That latter method, in brief, was to tag each meme-containing post with a unique text string that could enable you to use a search engine as an aggregator. But how can this be turned into a useful service?
The first challenge we have to tackle is a way to generate strings that are long enough to be unique even if our service is highly utilized yet short enough that they won’t fill up half of the screen real estate available for a given post. This one’s easy; it’s called a URL. Let’s imagine that Google is providing this service. You would enter the name of the meme for which you want to create a tracking tag and Google would spit back a unique tracking URL that you could link to the name of the meme in the post, or to some “meme tracker” graphic, or anything else you’d like.
Now, how would you let the service know that a new instance of the meme exists on the web? Not everybody has trackback, XML-RPC, and what have you. The lowest tech method I can think of (albeit somewhat more resource-intensive than, say, a RESTful API) would be through referrer logs. When a person posts a new instance of the meme on a page, s/he would just click on the URL to go to the tracking page. The service would compare the URL of the referring page to all known instances of the meme and, seeing that the URL is a new instance, catalogue it.
So far this is kinda nice but not yet what I’d consider to be super-cool. You can tag arbitrary content on the web, you can do it in a low-tech way to make it easy for everyone to do, and you can allow people who know about the service to submit their instances to the service for tracking very easily. But there are a couple of problems that this service doesn’t solved yet. How do you find instances that people haven’t tagged? How do you deal with overlapping meme labels?
The answer to these two problems is simple: Bayesian analysis. Once you have built up a corpus of tagged meme instances, you run a Bayesian filter similar to the adaptive systems used in good junk mail filters these days. The system can begin to recognize the common characteristics (i.e., words, phrases, and other text strings) to the various meme instances. It can tell you what those characteristics are. And it can construct a web search using those characteristics. It would even be fairly easy to put in a slider control, much like the agressiveness slider in spam filters, to be more or less choosy about how wide a net the meme search should cast. Likewise, it would be fairly easy to run a comparison of two meme labels to see if there is overlap in the corpi. A “meme mapping” tool could tell you, for example, that there is an 85% overlap in terminology between the “idea virus” meme and the “meme” meme and even allow you to merge the two (or map the overlap and differences).
There would only be one more piece that we’d need to make the tool complete. Remember, we want to study meme propagation which means we need to know how it spreads over time. So we’d need a tool that gathers a bit more meta-data about the meme instances and their relationship to each other on the network and over time. Specifically we’d need to know
- When each meme instance was published (to the best of our ability to determine)
- Whether that instance is linked to previous instances
- Whether the site in which the instance appears is otherwise linked to sites that contain previous instances
With this data, we could at least start making educated guesses about how particular memes spread, how fast they spread, the nature of the network in which they spread, and so on. This piece is beyond my extremely limited technical knowledge to plan out, but I would imagine that a relatively simple spider could gather all this info. And with this last piece in place, you’d have a fairly complete meme-tracking service that requires very little technical ability or infrastructure from the end users and relies on fairly basic and well-understood technologies.
Note to hackers: You missed my birthday this year, but there’s still plenty of time before Hannukah…
-
Tracking Memes in the Wild, Part II
In my last post I talked about a meme tracking experiment and bemoaned the fact that it provided no way to track arbitrary memes in the wild. Luckily, an e-Literate reader put me on the track to a workable idea in his comment on a previous post.
Martin Terre Blanche points us to a post on his own blog which describes an a trick performed by Stephen Downes (jeez, does this guy ever sleep?) to aggregate blog posts from different blogs about a conference. The trick was that all the posts would use a particular text string or “shibboleth” phrase to identify them as being about the same topic. as Martin correctly points out, any unique (or relatively rare) text string could be used by Google or any other search engine to aggregate posts on a topic, as long as those posts were properly tagged.
There are several things that are interesting about this approach:
- It can be applied to any arbitrary content.
- The participants would decide what counts as the “meme.” So you could really find out what people perceive to be the common idea independent of the particular text content in individual posts. For example, a person writing about “idea viruses” who never uses the word “meme” in a post could still label that post as belonging to the “meme”…er…meme and we could aggregate it along with this post.
- It would be pretty simple to implement this as a web service with existing technology.
This last point really has me thinking. In my next post (after I walk the doggies), I will describe my own vision of how a meme tracking service could work.
-
Tracking Memes in the Wild, Part I
The other day, I ran into this post on the Contentious weblog which, in turn, led me to this longer post about an experiment conducted by a PhD student. Basically, he created a survey that he asked people to fill out, post to their blogs, and then pass on, like chain email. He wanted to see how the “meme” propagated across the web, tracking where it showed up and how it changed or “mutated.”
While this is an interesting study, it has several very prominent limitations. First, I’m not sure if I would call a complete survey a “meme.” Here is how the author himself defines “memes:”
A meme can be defined as self-propagating unit of information similar to the biological concept of a gene, the term was first used by Richard Dawkins in his book The Selfish Gene). Familiar examples of Internet memes are well known: I kiss you, All your base are belong to us, The Nike sweatshop story and more recently p23s5 and the order of words meme. Memes are are often considered “idea viruses” that spread in communities, with many believing that various religions can be considered meme like.
Leaving aside the open question of how big a meme can be (e.g., can a whole religion really be put in the same category as, say, an urban legend?), there’s nothing really viral about a survey per se. What’s interesting about the idea of memes is not simply that they spread but that they spread by sticking in our heads. In other words, you shouldn’t need to do an elaborate copy and paste operation in order to propagate a meme; it should be inherently memorable. So I don’t know if what the author is testing really is meme propagation.
From a more practical standpoint (and this isn’t really the fault of the author, since he isn’t claiming that his method will have practical applications), what I really want to be able to do is track any arbitrary meme as it wings its way across the internet. The author’s method doesn’t do this for us. Luckily, there is an idea floating out there in the blogosphere that might just do the trick.
More on that in the next two posts.
