Showing posts with label crowdsourcing. Show all posts
Showing posts with label crowdsourcing. Show all posts

Sunday, April 1, 2012

Fish, Phones and Crowdsourcing

Originally posted 8/26/2010

I began my career many years ago developing and teaching software system development methods.  One of the principles my colleagues and I stressed is still valid today in the context of social networking but often ignored: data is only as good as it’s most frequent usage.

Why does frequency of use effect data quality?  Data that is frequently used is apt to be good data, since people are looking at it regularly and using it for some meaningful real-world purpose:  If an error is present in the data to start with, or if the data changes, someone notices it quickly and it gets corrected.  Data that is infrequently used is typically of lower quality for the same reason: It isn’t used as often, so when errors or changes occur they take longer to notice and fix.  Consider phone numbers: if you build a web site and require visitors to enter phone numbers, how many of them are valid to start with? How many initially valid ones are still be good after a year? After five years? Unless you are regularly using those numbers to call people, the quality declines with time.

Social networking sites often seem to forget that data, like fish, can go bad.  And I’m not referring to the fact that lots of sites try to collect profile information that will never be used (like asking for my phone number when you have no intention of ever calling me).  I’m referring more to crowdsourcing: social networking sites set up to collect user generated content (UGC) intended to help solve some kind of problem.  While some UGC is long-lived (“how do I turn on/off a setting in a popular application,” for instance), much UGC has a much shorter shelf life (“where’s the cheapest place to buy gasoline this week,” for instance).  Everyone has had experience wandering into a forum, discussion group, or blog where there has been no activity for many, many months.  And if you are like me, when I wander into such I site I tend to wander right back out again. 

Like stale fish in your refrigerator, stale UGC on your social networking site can spoil other sections by driving away visitors. So a word of advice to organizations setting up crowdsourcing applications: keep the freshness of data in mind.  Throw out new interesting topics frequently to keep content fresh and retire old topics when they’ve stopped attracting new content.  Move old content to archives so that users don’t accidentally confuse fresh content with older content.

And while you are at it, stop asking for my phone number.

Friday, March 30, 2012

The Flip Side of Social Networking Success

Originally published 8/11/2010

An interesting story came to light a few days ago. Presumably you are aware of the social news site Digg.  Well, it seems that a group of conservative members of the site conspired (apparently for months if not years) to drive its member-voted top content towards the right end of the political spectrum. By utilizing multiple accounts (against Digg’s terms of use) and by communicating with each other outside the community (through an invitation-only Yahoo group called the Digg Patriots), this group was massively successful in voting up news items that furthered their conservative agenda, and burying content that conflicted with it. They were also successful in provoking more liberal members of the site to the point where those members became belligerent and insulting enough to be banned from the site.

Now shills and plants were around for a long time before social networking, and some would argue that this group did nothing wrong. After all, Digg is all about crowdsourcing the top news stories (ok, not just news, but whatever bright and shiny web content that attracts the attention of the community). But this is an interesting phenomenon: These (apparently) weren’t paid actors intent on driving people to do more business with a company (or do less with a competitor); this was a group that basically took it upon themselves to hijack a community to further their political agenda. Why did they pick Digg? One would assume because of Digg’s popularity: according to a 2009 study by Social Computing Journal, it’s the #9 top social networking site.  After all, why hijack a site hardly anyone looks at? For social networking it seems, success is a two edged sword: A large group of sufficiently motivated users might make your social networking site an incredible success, but such a group might also choose to bend or break it to suit their agenda. And sadly, the more successful your social networking site the more likely it is that it will attract this type of activity.

Saturday, March 26, 2011

Crowdsourcing College Basketball

I saw an interesting article a few days ago about using crowdsourcing to make your picks for the big college basketball tournament currently underway (I don’t think I’m allowed to use any phrases that contain the words “madness,” “final,” “March,” or “four” or the initials “N,” “C,” “A,” or “A” without paying somebody a royalty, so I won’t). 

The article turned out not to be what I thought it might be about.

When I first saw the headline, I thought it might somehow be a crowdsourcing project to find out what a perfect tournament bracket would look like going into the tournament (after all, filling out a perfect bracket after the tournament is a trivial exercise).  That seemed like a mildly interesting idea but ultimately a futile one: we don’t award titles based on how people *think* teams are going to do, but how they *actually* do.  As the old saying goes, “that’s why you play the games.” So what do you do with a crowdsourced bracket where the crowd picks the winners?

Why else fill out a bracket?  The article presented a strategy for betting (ok, not “betting” betting, but more like “winning the office pool” betting, since “betting” is illegal in a lot of places).  The idea, originally put forward by author Mark McClusky last year, is to figure out what that ideal bracket is—the one with all the favorites identified—and use some basic statistical analysis to figure out how to bet (ahem, position) against it.  In other words, if the idea of the office pool is to pick more winners than everyone else, it would be good to pick some teams that have a reasonable chance of winning, but are not the favorites.  That way, if you are right, your picks vault ahead of the “crowd” that picked the “likely” winners.

Of course, in a year where there are no upsets (which isn’t very likely) you’ll undoubtedly lose.  All the players who pick the favorites will all tie for the top spot and have to split the pool.  Kind of like playing Hurley’s lottery numbers from LOST and winning. (You would think not very many would have played those numbers in January 2011, since LOST went off the air last year. You would be wrong. Over nine thousand people hit the New York Mega Millions jackpot with those numbers this January: they all split the total pot, of course, and received a whopping $150)

Of course, using crowdsourcing to set odds is not new: if you’ve ever been to a horse race or dog race, that’s how the odds are calculated.  The horse (or dog) that most people have picked to win is the favorite, with the long odds on the horse (or dog) that the fewest people have picked to win.  (Which, by the way, is why most movies or TV shows that show someone winning a fortune by making a big bet on an underdog are bogus: if you bet a million dollars on a million-to-one long shot, you just made that long shot into the new favorite, and *if* it wins you won’t get much more than your bet back.)

In any case, I thought it was an interesting way to use the wisdom of the crowds to bet (er, position) against the crowd.  A high-risk, high-reward strategy to be sure, but an interesting one.

And for the record, this year I’m picking Kansas University to win the national championship (to be fair, I pick them every year). I didn’t attend KU, but I lived near Lawrence for nearly 30 years and know a lot of people that did. So Rock Chalk! Jayhawk! Go KU!

Thursday, December 30, 2010

Social Networking Your Car

An interesting article about crowdsourcing your car hit the wires this week.  Nissan’s new Leaf electric car (which began shipping recently) includes a wireless check-in feature that you can use to compete with other Leaf owners for energy efficiency awards.

It’s a pretty interesting idea: using the Carwings system (an OnStar-like system that originated in Japan) the Leaf wirelessly transmits your energy usage statistics to their servers.  The system then compares your usage to other Leaf owners in your region and awards trophies to the winners, which then display in the Leaf’s dashboard control center. 

In addition, the system uses data from drivers of other Carwings users (not just Leaf drivers) to aggregate and display near real-time traffic predictions, further enabling the Leaf driver to make intelligent decisions about energy economy.

This friendly competition seems an excellent way to encourage Leaf drivers to maximize their energy efficiency, but I wonder if it’s a good idea for the rest of us.  Let me explain: The concept of “hypermiling” (and it amuses me more than it should to include a hyperlink here) started a couple of years ago when drivers—particularly drivers of hybrid cars—began to wonder just how many miles-per-gallon they could actually get from their vehicle.  Unfortunately, some of the techniques they began to use were illegal and/or dangerous: things like tailgating to “draft” behind large trucks, turning the car on and off to coast in neutral, coasting through stop signs, and driving even more slowly in freeway slow lanes.

If we have hypermilers (including Leaf drivers) increasingly trying to use some of these techniques to optimize their miles, it may in fact cause those of us who aren’t hypermiling to burn more energy trying to avoid them or get around them.  The net effect is that if a hypermiler saves one ounce of fuel, but causes other people to burn more than one ounce of fuel, it’s a net loss for the energy system, not a net gain. (I do have to note that there are lots of ethical hypermilers out there that utilize *safe* and *legal* methods to squeeze those extra miles out of their cars, and I highly commend them.)

I’ve always been a proponent of giving people lots of information so they can make informed decisions.  The problem with the Leaf contest, in my humble opinion, is that without some education as to safe and legal ways to drive efficiently (and not inconvenience or endanger other drivers) is going to cause problems for other drivers.  Leaf drivers who’ve never heard of hypermiling or know what aspects of it are safe and legal may be encouraged by the competition to rediscover driving techniques that have already been tried by earlier hypermilers and discarded for safety reasons.  People seem to be very good at rediscovering inventive wrong things to do, and making them aware of safe ways to achieve energy efficiency should be a mandatory prerequisite to owning such a car.  Especially a car that encourages competition.

And if they really wanted this to be a fair and accurate competition as to who is using energy most efficiently, they should include mileage statistics from bicyclists and people working from home.