e-Literate

Present is Prologue

Tag: memes

  • Tuning Folksonomies

    A while back, I posted an idea for checking to see the degree to which two differently named memes overlap in content. Looking back, what I was really talking about was tuning a folksonomy. What we really want is a way to see how much overlap there is between two tags so that we can merge or split them when appropriate.

    It now looks like somebody has taken the first step in this direction. A group of folks (including Lilia Efimova of Mathemagenic fame) are building an application called Blogtrace that, among other things, compares two web pages to see the degree to which their key terms (i.e., ontologies) overlap. You can read a brief description and even see a screen grab on Anjo Anjeweirden’s blog.

    It seems to me that if you just extend this model a bit you could compare two folksonomy tags by spidering the content of each and performing the same textual analysis. I think that would be very useful. Going a bit further, it would be interesting if, after performing the analysis, it could spit out a list of URLs for each tag that have a low degree of overlap with the other tag. This could be helpful in understanding why the overlap is not perfect and even how the tags could be refactored to make them more useful.

    Anyway, blogtrace looks interesting. If you’re curious, Lilia has posted a schematic of how the application works.

  • WordPress Related Entries plugin

    A while back I hinted to the hackers in the crowd that I would love to have a tool for aggregating the posts related to a particular meme on my site. Well, it looks like somebody did it. Unfortunately, it was for the wrong blog tool.

    D’Arcy Norman called my attention to this WordPress plug-in that does pretty much what I wanted. Very cool.

    Overall, I’m very impressed with what I see at WordPress. We’ll have to see whether the pMachine hackers catch up and create an equivalent plug-in before the WordPress hackers finish developing their import tool from pMachine blogs….

  • Tracking Memes in the Wild, Part III

    In my last two posts, I wrote about the limitations of one method for tracking memes and the promise of second method. That latter method, in brief, was to tag each meme-containing post with a unique text string that could enable you to use a search engine as an aggregator. But how can this be turned into a useful service?

    The first challenge we have to tackle is a way to generate strings that are long enough to be unique even if our service is highly utilized yet short enough that they won’t fill up half of the screen real estate available for a given post. This one’s easy; it’s called a URL. Let’s imagine that Google is providing this service. You would enter the name of the meme for which you want to create a tracking tag and Google would spit back a unique tracking URL that you could link to the name of the meme in the post, or to some “meme tracker” graphic, or anything else you’d like.

    Now, how would you let the service know that a new instance of the meme exists on the web? Not everybody has trackback, XML-RPC, and what have you. The lowest tech method I can think of (albeit somewhat more resource-intensive than, say, a RESTful API) would be through referrer logs. When a person posts a new instance of the meme on a page, s/he would just click on the URL to go to the tracking page. The service would compare the URL of the referring page to all known instances of the meme and, seeing that the URL is a new instance, catalogue it.

    So far this is kinda nice but not yet what I’d consider to be super-cool. You can tag arbitrary content on the web, you can do it in a low-tech way to make it easy for everyone to do, and you can allow people who know about the service to submit their instances to the service for tracking very easily. But there are a couple of problems that this service doesn’t solved yet. How do you find instances that people haven’t tagged? How do you deal with overlapping meme labels?

    The answer to these two problems is simple: Bayesian analysis. Once you have built up a corpus of tagged meme instances, you run a Bayesian filter similar to the adaptive systems used in good junk mail filters these days. The system can begin to recognize the common characteristics (i.e., words, phrases, and other text strings) to the various meme instances. It can tell you what those characteristics are. And it can construct a web search using those characteristics. It would even be fairly easy to put in a slider control, much like the agressiveness slider in spam filters, to be more or less choosy about how wide a net the meme search should cast. Likewise, it would be fairly easy to run a comparison of two meme labels to see if there is overlap in the corpi. A “meme mapping” tool could tell you, for example, that there is an 85% overlap in terminology between the “idea virus” meme and the “meme” meme and even allow you to merge the two (or map the overlap and differences).

    There would only be one more piece that we’d need to make the tool complete. Remember, we want to study meme propagation which means we need to know how it spreads over time. So we’d need a tool that gathers a bit more meta-data about the meme instances and their relationship to each other on the network and over time. Specifically we’d need to know

    • When each meme instance was published (to the best of our ability to determine)
    • Whether that instance is linked to previous instances
    • Whether the site in which the instance appears is otherwise linked to sites that contain previous instances

    With this data, we could at least start making educated guesses about how particular memes spread, how fast they spread, the nature of the network in which they spread, and so on. This piece is beyond my extremely limited technical knowledge to plan out, but I would imagine that a relatively simple spider could gather all this info. And with this last piece in place, you’d have a fairly complete meme-tracking service that requires very little technical ability or infrastructure from the end users and relies on fairly basic and well-understood technologies.

    Note to hackers: You missed my birthday this year, but there’s still plenty of time before Hannukah…

  • Tracking Memes in the Wild, Part II

    In my last post I talked about a meme tracking experiment and bemoaned the fact that it provided no way to track arbitrary memes in the wild. Luckily, an e-Literate reader put me on the track to a workable idea in his comment on a previous post.

    Martin Terre Blanche points us to a post on his own blog which describes an a trick performed by Stephen Downes (jeez, does this guy ever sleep?) to aggregate blog posts from different blogs about a conference. The trick was that all the posts would use a particular text string or “shibboleth” phrase to identify them as being about the same topic. as Martin correctly points out, any unique (or relatively rare) text string could be used by Google or any other search engine to aggregate posts on a topic, as long as those posts were properly tagged.

    There are several things that are interesting about this approach:

    • It can be applied to any arbitrary content.
    • The participants would decide what counts as the “meme.” So you could really find out what people perceive to be the common idea independent of the particular text content in individual posts. For example, a person writing about “idea viruses” who never uses the word “meme” in a post could still label that post as belonging to the “meme”…er…meme and we could aggregate it along with this post.
    • It would be pretty simple to implement this as a web service with existing technology.

    This last point really has me thinking. In my next post (after I walk the doggies), I will describe my own vision of how a meme tracking service could work.

  • Tracking Memes in the Wild, Part I

    The other day, I ran into this post on the Contentious weblog which, in turn, led me to this longer post about an experiment conducted by a PhD student. Basically, he created a survey that he asked people to fill out, post to their blogs, and then pass on, like chain email. He wanted to see how the “meme” propagated across the web, tracking where it showed up and how it changed or “mutated.”

    While this is an interesting study, it has several very prominent limitations. First, I’m not sure if I would call a complete survey a “meme.” Here is how the author himself defines “memes:”

    A meme can be defined as self-propagating unit of information similar to the biological concept of a gene, the term was first used by Richard Dawkins in his book The Selfish Gene). Familiar examples of Internet memes are well known: I kiss you, All your base are belong to us, The Nike sweatshop story and more recently p23s5 and the order of words meme. Memes are are often considered “idea viruses” that spread in communities, with many believing that various religions can be considered meme like.

    Leaving aside the open question of how big a meme can be (e.g., can a whole religion really be put in the same category as, say, an urban legend?), there’s nothing really viral about a survey per se. What’s interesting about the idea of memes is not simply that they spread but that they spread by sticking in our heads. In other words, you shouldn’t need to do an elaborate copy and paste operation in order to propagate a meme; it should be inherently memorable. So I don’t know if what the author is testing really is meme propagation.

    From a more practical standpoint (and this isn’t really the fault of the author, since he isn’t claiming that his method will have practical applications), what I really want to be able to do is track any arbitrary meme as it wings its way across the internet. The author’s method doesn’t do this for us. Luckily, there is an idea floating out there in the blogosphere that might just do the trick.

    More on that in the next two posts.

  • Idea Viruses as Informational Cascades

    In a previous post, I suggested that so-called “idea viruses” might be thought of as either causes of or manifestations of informational cascades. I now think manifestation is the right characterization rather than cause. I am persuaded by this fascinating and frightening piece of research [PDF] showing that legislators tend to try to garner support for their positions by creating informational cascades within the legislative body. It’s not quite a classic idea virus in the sense that it isn’t a new idea that each recipient “catches,” but it does reflect a contagion of an opinion or position.

  • Book Recommendation: The Selfish Gene

    It may seem odd, given the focus of this blog, to recommend a book on evolutionary biology. But Richard Dawkins’ book The Selfish Gene lays a solid foundation for helping to understand developments in the aggregation sciences.

    Dawkins’ main thesis is that evolution is driven by the survival of the fittest genes, not organisms and not species. If it seems odd to think of dumb genes “competing” or acting “selfishly,” as the title suggests, then you’re catching on to what’s provocative about this book. If you have trouble wrapping your head around the ideas of complex adaptive systems and emergence, then take a step back and start by trying to understand selfish genes. If you can make the latter conceptual leap then the former becomes much easier to get. Dawkins helps with the learning process by giving many, many concrete examples and carefully reasoned, easy-to-follow arguments.

    If you borrow the book rather than buying it, make sure that you get the 1989 edition or later; the newer edition has two new chapters that are essential. In particular, Chapter 12, “Nice guys finish first,” is about the evolution of cooperation and has quite a bit about game theory in it. Again, if you understand game theory and its impact on a population, then other ideas in aggregation science will come a lot easier.

    One last point of interest is that Dawkins’ book is the origin of the term “meme”. If you’ve heard the term “idea virus,” then you have a sense of what memes are. (And by the way, thinking about memes is one good way to start getting a sense of why informational cascades are a potential problem in the kind of networked conversational learning that happens, for example, in a blogging community. You could argue that an “idea virus” is either a cause or a manifestation of an informational cascade.) To see an example of somebody using the “meme”…er…meme productively, see this oldie but goodie post from the Truth Laid Bear weblog.