Showing posts with label Web Analytics. Show all posts
Showing posts with label Web Analytics. Show all posts

Saturday, August 28, 2010

Mad Skills for Big Data

Big Data is a big deal these days, so it was with great interest that we welcomed Brian Dolan to the SDForum Business Intelligence SIG August meeting to speak on "MAD Skills: New Analysis Practices for Big Data". MAD is an acronym for Magnetic Agile Deep, and as Brian explained, these skills are all important in handling big data. Brian is a mathematician who came to Fox Interactive Media as Lead Analyst. There he had to help the marketing group with deciding how to price and serve advertisements to users. As they had tens of millions of users that they often knew quite a lot about, and served billions of advertisements per day, this was a big data problem. They used a 40 node Greenplum parallel database system and also had access to a 105 node map reduce cluster.

The presentation started with the three skills. Magnetic, means drawing the analyst in by giving them a free reign over their data and access to use their own methods. At Fox, Brian grappled with a button down DBA to establish his own his own private sandbox where he could access and manipulate his own data. There he could bring in his own data sets, both internal and external. Over time the analysts group established a set of mathematical operations that could be run in parallel over the data in the database system speeding up their analyses by orders of magnitude.

Agile means analytics that adjust react and learn from your business. Brian talked about the virtuous cycle of analytics, where the analyst first acquires new data to be analyzed, then runs analytics to improve performance and finally the analytics causes business practices to suit. He talked through the issues at each step in the cycle and led us through a case study of audience forecasting at Fox which illustrated problems with sampling and scaling results.

Deep analytics is about producing more than reports. In fact Brian pointed out that even data mining can concentrate on finding a single answer to a single problem where big analytics has the need to solve millions of problems at the same time. For example, he suggested that statistical density methods may be better at dealing with big analytics than other more focused techniques. Another problem with deep analysis of big data is that, given the volume of data, it is possible to find data that supports almost any conclusion. Brian used the parable of the Zen Tea Cup to illustrate the issue. The analyst needs to be to approach their analysis without preconceived notions or they will just find exactly what they are looking for.

Of all the topics that came up during the presentation, the one the caused most frissons with the audience was dirty data. Brian's experience has been that cleaning data can lose valuable information and that a good analyst can easily handle dirty data as a part of their analysis. When pressed by an audience member he said "well 'clean' only means that it fits your expectation". As an analyst is looking for the nuggets that do not meet obvious expectations, sanitizing data can lose those very nuggets. The recent trend to load data and then do the cleaning transformations in the database means that the original data is in the database as well as the cleaned data. If that original data is saved, the analyst can do their analysis with either data as they please.

Mad Skills also refers to the ability to do amazing and unexpected things, especially in motocross motor bike riding. Brian's personal sensibilities were more forged in punk rock, so you could say that he showed us the "kick out the jams" approach to analytics. You can get the presentation from the BI SIG web site. The original MAD Skills paper was presented at the 2009 VLDB conference and a version of it is available online.

Sunday, April 04, 2010

Web Analytics 2.0

Web analytics is changing fast as we discovered at the March meeting of the SDForum Business Intelligence SIG. Avinash Kaushik, Analytics Evangelist for Google, spoke on 'Web Analytics 2.0: Rethinking Decision Making in a "2.0" World'. Avinash started off by telling us how he became well known as a web analytics guru. A few years ago he started writing his blog "Occam's Razor". It soon gathered a large readership and a publisher approached him to write a book. His first book "Web Analytics: An Hour a Day" was a distillation of his blog posts. The book is a best seller even although much of its contents is available for free on the web. His second book "Web Analytics 2.0" came out recently.

Avinash is an excellent communicator with a strong personal style. One aspect of that style, quite obvious from his blog posts, is the urge to create lists of ideas. For his presentation, Avinash offered us a list of simple ideas on web sites metrics and analytics. Here are some of the ideas that he presented.

The first idea is simple and direct - Don't Suck. The suckage of a web page can be measured by a metric called Bounce. This is a relatively new metric that we had not previously heard discussed at the Business Intelligence SIG. Bounce measures the users whose experience of the web page and site is, as Avinash put it "I came. I puked. I left." and he showed us some pretty pukey pages that might back up this behavior. A typical analysis is to look at the pages with the highest bounce rate, determine why they cause that behavior and what can be done about it.

His next idea is Segment or Die. Analytics is about aggregating data to make sense of large datasets, however over-aggregation results in a single number and nothing to compare it with. Segmenting the data gives us a number of data items that we can compare. Avinash showed us a simple example where he took a hospital web site and classified the content into 8 segments and then compared the amount of content against the number of page views in each segment. It was immediately apparent where the effort should go into adding and improving content.

Analyzing your web logs only tells a part of the story, you also have to worry about what the analytics cannot tell you. Ex Defense Secretary Donald Rumsfeld is infamous for having talked about "the known knowns, the known unknowns, and the unknown unknowns". The unknown unknowns are the things that you don't even know that you don't know and therefore the thing you should be most worried about. You can start to get a handle on what you do not know by looking at your performance relative to your competitors. This is known as Benchmarking, or in the case of a deep study as Competitive Intelligence. For an example of what can be done, a recent post on the Occam's Razor blog discusses 8 sources for Competitive Intelligence data.

Most web site analytics only looks for the one big conversion from a web site, however there are many other small conversions that are tracked and worth evaluating because there may be hidden value lurking in the long tail. For example, recently Avinash wanted to know what his blog was worth so that he could defend taking time away from the family to write it. After determining a value for each reader, he started adding up all the other micro-conversion like people who subscribe to the RSS feed and advertisements for his books and the non-profit organizations that he supports. Overall he came up with a figure of about $26000 per month. Now Avinash does not make a penny from his blog, so this is notional money that adds to his personal brand value, but that value seems to make the effort of writing the blog well worth the time spent.

The next idea is Fail Faster. By this Avinash means do lots of different experiment, many of which will fail, to find out what works. He led us through an example from the Obama Presidential campaign. President Obama raised huge amounts of money from many small donations on his web site. The initial web page worked well. The experiments were to try some variations on the theme. Pages with video, and stirring video at that, did very badly. A simple picture of Obama with his family did a little better than the initial picture, so that one was chosen.

Avinash showed us this example to make a number of points. The Obama analytics team was tiny. Often the best work is done by a small agile team that has the freedom to experiment. The team used free tools. Avinash believes that that good people are much more important than good tools. His suggestion for dividing up the analytics budget is to spend 90% on people and 10% on tools. Sometimes, a web site design feature starts from a HiPPO (Highest Paid Persons Opinion), which can be destructively bad, and difficult to get around because in all organizations the highest paid persons opinion is taken very seriously. The best way to counter a HiPPO is to show that other ideas work better through the results of experiments that produce hard evidence.

While some may think that web analytics is a mostly solved problem, Avinash believes we are just starting to figure out what can be done, and that there is plenty of room for more innovation. I will continue to read Occam's Razor to find out where he takes us next.

Wednesday, January 21, 2009

Numbers not Napkins

Dave McClure, Master of 500 Hats, spoke the SDForum Business Intelligence SIG January meeting on "Numbers not Napkins: Simple Metrics & Business Models for Startups". This is an adult version of a talk that he has given several times in different places on "Startup Metrics for Pirates". The presentation is available online at SlideShare, one of several companies that Dave advises.

The back story is that Web 2.0 start-ups need good metrics to see how well they are doing and to identify the things that they need to improve. Dave groups metrics into 5 buckets: Acquisition, Activation, Retention, Referral and Revenue; or AARRR hence Metrics for Pirates. On the web, once you have decided what you want to know, it is easy to collect good data. So the issue how do you decide what you want to know?

Dave proposes that you can do all this with 3 simple one page documents. The first document is a one page business plan. The plan is a table with columns for each type of visitor to the web site and a row for actions that they take to use the web site. A good plan will have at most 2 or 3 different types of users and 3 to 5 rows drawn from the {
acquisition, activation, retention, referral, revenue} set. Over time the business plan will evolve. For example, at the very beginning acquisition and activation may dominate while later on revenue will take a bigger part of the picture. Dave spent some time discussing business plans and giving examples from companies that he is involved in.

The one page business plan leads into the second document, a set of metrics that measure "conversions", that is the number of visitors who perform an action described in the business plan. Also important to measure is how the visitor got there. This leads to understanding the channels that bring visitors to your web site and defines the third one page document, the marketing plan. For each channel, you want to measure at a minimum the volume, cost and conversion rate of visitors.

Dave's has three mantras that he repeated throughout his talk. Firstly, keep it simple so that you can understand what is going on and quickly react to the things that your metrics tell you. Secondly, keep the data actionable so that you can act on what you find out. Thirdly, keep iterating, test new ideas, optimize based on the insight from your tests.

All in all it was a great talk, that got much of its weight from the many examples that Dave could discuss with first hand knowledge.

Sunday, November 16, 2008

Map-Reduce versus Relational Database

In the previous post I said that Map-Reduce is just the old database concept of aggregation rewritten for extremely large scale data. To understand this, lets look at the example in that post, and see how it would be implemented in a relational database. The problem is to take the World Wide Web and for each web page count the number of different domains that reference that page in links on the other web pages.

As a starting point, let us consider using the same data structure, a two column table where the first column contains the URL of each web page and the second column contains the contents of that web page. However, this immediately presents a problem. Data in a relational database should be normalized, and the first rule of normalization is that each data item should be atomic. While there is some argument as to exactly what atomic means, everyone would agree that the contents of a web page with multiple links to other web pages is not an an atomic data item, particularly if we are interested in those links.

The obvious relational data structure for this data is a join table with two columns. One column, called PAGE_URL, contains the URL of the page. The other column, called LINK_URL, contains URLS of links on the corresponding page. There is one row in this table (called WWW_LINKS) for every link in the World Wide Web. Given this structure we can write the following SQL query to solve the problem in the example (presuming a function called getdomain that returns the domain from a URL):

SELECT LINK_URL, count(distinct getdomain(PAGE_URL))
FROM WWW_LINKS
GROUP BY LINK_URL

The point of this example is to show that Map-Reduce and SQL aggregate functions both address the same kind of data manipulation. My belief is that most Map-Reduce problems can be similarly expressed by database aggregation. However there are differences. Map-Reduce is obviously more flexible and puts less constraint on how the data is represented.

I strongly believe that every programmer should understand the principals of data normalization and why it is useful, but I am willing to be flexible when it comes to practicalities. In this example, if the WWW_LINKS table is a useful structure that is used in a number of different queries, then it is worth building. However if the only reason for building the table is to do one aggregation on it, the Map-Reduce solution is better.

Tuesday, November 11, 2008

Understanding Map-Reduce

Map-Reduce is the hoopy new data management function. Google produced the seminal implementation. Start-ups are jumping on the gravy train. The old guard decry it. What is it? In my opinion it is just the old database concept of aggregation rewritten for extremely large scale data as I will explain in another post. But firstly we need to understand what Map-Reduce does, and I have yet to find a good clear explanation, so here goes mine.

Map Reduce is an application for performing analysis on very large data sets. I will give a brief explanation of what Map Reduce does conceptually and then give an example. The Map Reduce application takes three inputs. The first input is a map (note lower case). A map is a data structure, sometimes called a dictionary. A tuple is a pair of values and a map is a set of tuples. The first value in a tuple is called the key and the second is called the value. Each key in a map is unique. The second input to Map-Reduce is a Map function (note upper case) . The Map function takes as input a tuple, (k1, v1) and produces a list of tuples (list(k2, v2)) from data in its input. Note that the list may be empty or contain only one value. The third input is a Reduce function. The Reduce function takes a tuple where the value is a list of values and returns a tuple. In practice it reduces the list of values to a single value.

The Map Reduce application takes the input map and applies the Map function to each tuple in that map. We can think of it creating an intermediate result that is a single large list from the lists produced by each application of the Map function:
{ Map(k1, v1) } -> { list(k2, v2) }
Then for each unique key in the intermediate result list it groups all the corresponding values into a list associated with the key value:
{ list(k2, v2) } -> { (k2, list(v2)) }
Finally it goes through this structure and applies the Reduce function to the value list in each element:
{ Reduce(k2, list(v2)) } -> { (k2, v3) }
The output of Map Reduce is a map.

Now for an example. In this application we are going to take the World Wide Web and for each web page count the number of other domains that reference that page. A domain is the part of a URL between the first two sets of slashes. For example, the domain of this web page is "www.bandb.blogspot.com". A web page is uniquely identified by its URL, so a URL is a good key for for a map. The data input is a map of the entire web. The key for each map element is the URL of the page, and the value is the corresponding web page. Now I know that this is a data structure on a scale that is difficult to imagine, however this is the kind of data that Google has to process to organize the worlds information.

The Map function takes the URL, web page pair and adds an element to its output list for every URL that it finds on the web page. The key in the output list is the URL found on the web page and the value is the domain from the key value in the input URL. So for example, on this page, our Map function finds the link to the Google paper on Map Reduce and adds to its list of outputs the tuple ("research.google.com/archive/mapreduce.html", "www.bandb.blogspot.com"). Map-Reduce reorganizes its intermediate data so that for each URL it collects all the domains that reference that page and stores them as a list. The Reduce function goes through the list of domains and counts the number of different domain values that it finds. The result of Map-Reduce is a map where the key is a URL and the value is a number, the number of other domains on the web that reference that page.

While this example is invented, Google reports that they use a set of 5 to 10 such Map-Reduce steps to generate their web index. The point of Map Reduce is that a user can write a couple of simple functions and have them applied to data on a vast scale.

Wednesday, September 17, 2008

SaaS Data Integration

Data integration is the problem of gathering data, perhaps from many different application for the purpose of doing some analysis of the data as a whole. Mike Pittaro, Co-Founder of SnapLogic spoke to the SDForum Business Intelligence SIG September meeting on "Enhancing SaaS Applications Through Data Integration with SnapLogic".

The big players in data integration are Informatica and Ascential (now IBM Information Integration) who sell large, expensive and complex products. Because of the cost, these products are often not used, particularly for one off projects which are common. Mike helped found SnapLogic in 2005 to bring a new perspective to data integration. SnapLogic is an open source framework and therefore both affordable and extensible by its users.

He showed us the complexity of data integration. It involves dealing with many different access protocols, multiple ways of getting the data and each type of data has its own metadata format to describe the data. This he contrasted with the World Wide Web where huge amounts of data are pulled back and forth every day, without interoperability problems. There are almost 200 million web sites, and billions of users, yet World Wide Web is completely decentralized, with heterogeneous model that allows for different operating system, servers, client software applications and frameworks, and yet they are all compatible and interoperable.

The World Wide Web is based on open standards and protocols and an architectural principal called REST, which stands for REpresentational State Transfer. REST plays with data resources, in standardized representations and each resource identified by a unique identifier like a URL.

SnapLogic builds on this by turning data sources into standard web resources. With SnapLogic you configure a server to extracts data from a datasource like a file or database and transform the data into the form you want. The server presents the datasource as a standard web resource with a URL. These servers are the blocks for building a data integration application.

Sunday, March 30, 2008

Building Better Products Through Experimentation

Experimentation is the theme of the SDForum Business Intelligence SIG so far this year. The March meeting featured Deepak Nadig, a Principal Architect at eBay, talking about "Building Better Products Through Experimentation". Experimentation is an important technique for Business Intelligence, although its first uses were with medicine. In 1747, James Lind, a British naval surgeon performed a controlled experiment to find a cure for scurvy. In his book "Supercrunchers", Ian Ayres describes how the Food and Drug Administration has used experimentation since the 1940s to determine whether a medical treatment is efficacious.

While eBay has always used experimentation test and fine tune its web pages, in recent years the process has been formalized. While anyone can propose an experiment, product managers are the group of people who are most likely to do so. Deepak took us through the eBay process and discussed issues with using experimentation. Because they have the infrastructure, simple experiments can be set up within a matter of days. eBay usually runs an experiment for at least a week so that it is exposed to a full cycle of user behavior. Simple experiments to test a small feature typically run for a week or so, larger experiments may run for a month or two and some critical tests run continuously.

For example, eBay is interested in whether it is a good idea to place advertising on their pages. On the one hand it brings in extra revenue in the short term, on the other hand, it might cannibalize revenue in the long term. Experimentation has shown that advertising is a good thing in some situations, however its use is being monitored by some long term experiments to ensure that it remains beneficial.

Deepak took us through some of the issues that with experimentation. One issue is concurrency, how many experiments can be carried out at the same time. As eBay has a high traffic web site, they can get good results with experiments on a small proportion of the users, at most a few percent. As each experiment uses a small percentage of the users, several experiments can be run in parallel. Another issue is establishing a signal to noise ratio for experiments to ensure that experiments are working and giving valid results. eBay has done some AB experiments where A and B are exactly the same to establish whether their experimental technique has any biases.

Thursday, March 13, 2008

Customer Relationship Intelligence

There is a curious thing about the organization of a typical company. While there is one Vice President in charge of Finance and one Vice President in charge of Operations there can be up to three Vice Presidents facing the customer: a Marketing Vice President, a Sales Vice President, and a Service Vice President. On the one hand, the multiplicity of Vice Presidents and their attendant organizations is a testament to the importance of the customer. On the other hand, multiple organizations mean that no one is in charge of the customer relationship and thus no one takes responsibility for it.

We see this in the metrics that are normally used to measure and reward customer-facing employees. Marketing measure themselves on how well they find leads regardless of whether sales uses the leads. Sales measure themselves on the efficiency of the sales people in making sales regardless of whether the customer is satisfied. Service, left to pick up the pieces of an overpromised sale, measure themselves on how quickly they answer the phone. Every one is measuring their own actions and no one is measuring the customer.

Linda Sharp addresses this conundrum head on in her new book Customer Relationship Intelligence. As Linda explains, a customer relationship is built upon a series of interactions between a business and its customer. For example, the interactions starts with acquiring a lead, perhaps through an email or mass mailing response or a clickthrough on a web site. Next, more interactions qualify the lead as a potential customer. Making the sale requires further interactions leading up to the closing. After the sale there are yet more interactions to deliver and install the product and service to keep it working. Linda's thesis is that each interaction builds the relationship and that by recording all the interactions and giving them both a value and a cost, the business builds a quantified measure of the value of its customer relationships and how much it has spent to build them.

Having a value for a customer relationship completely changes the perspective of that relationship. It gives marketing, sales and service an incentive to work together to build the value in the relationship rather than working at cross purposes to build their own empires. Moreover, knowing the cost of having built the relationship suggests the value in continuing the relationship after the sale is made. In the book, Linda takes the whole of the second chapter to discuss customer retention and why that is where the real profit is.

The rest of the book is logically laid out. Chapter Three “A Comprehensive, Consistent Framework” creates a unified model of a customer relationship throughout its entire lifecycle from the first contact by marketing through sales and service to partnership. This lays a firm bedrock for Chapter Four, “The Missing Metric: Relationship Value” which explains the customer relationship metric, the idea that by measuring the interactions that make the relationship we can give a value to the relationship.

The next two chapters discuss how the metric can be used to drive customer relationship strategy and tactics. The discussion of tactics lays the foundation for Chapter Seven, which shows how the metric is used in the execution of customer relationships. Chapters Six and Seven contain enough concrete examples of how the data can be collected and used to give to give us a feeling of the metric’s practicality. Chapter Eight compares the customer relationship metric with other metrics and explores the many ways in which it can be used. Finally, Chapter Nine summarizes the value of the Customer Relationship Intelligence approach.

Linda backs up her argument with some wonderful metaphors. One example is the contrast between data mining and the data farming approach that she proposes with her Relationship Value metric. For data mining, we gather a large pile of data and then use advanced mathematical algorithms to determine which parts of the pile may contain some useful nuggets of information. This is like the hunter-gatherer stage of information management. When we advance into the data farming stage, we know what customer relationship metric is important and collect that data directly.

As the metaphor suggests, we are still in the early days of understanding and developing customer relationship metrics. Until now, these metrics have concentrated on measuring our own performance to see how well we are doing. Linda Sharp’s Relationship Value metric turns this on its head with a new metric that measures our whole relationship with customers. Read the book to discover a new and unified way of thinking about and measuring your customers.

Sunday, September 23, 2007

The Evolution of Web Analytics

We have not had a talk on Web Analytics for many years at the SDForum Business Intelligence SIG, so it was great to hear Stephen Oachs, founder and CTO of VisiStat, speak on "The Evolution of Web Analytics" at our September meeting. VisiStat is a two year old start-up that provides a web site performance measurement and analytics service (Software as a Service model) in the Small and Medium Business market.

As Stephen told us, first generation web analytics was about collecting data from web logs, integrating that with data from other sources and presenting historical results to IT specialists. The current generation, which he called web site performance management, collects data by page tagging, which entails adding a small snippet of JavaScript to each page. In practice the code snippet is added to a common page header or footer so it only needs to be added once to cover all pages in a site.

Page tagging collects more information than can be extracted from web server logs and it does not require difficult integration to make sense of the data. With better analysis software, the results of page tagging are ready to show directly to end users like the marketing and sales people who are responsible for the contents of the web site. Also, with page tagging we can see the data in real time, which allows the following of a user as they browse around the web site.

Real time access to information opens new doors. Stephen told us about a specific case where a bank became aware that it was subject to a phishing attack on its customers when the bank noticed an unusual change in the patterns of access to their web site. Similarly, click fraud can be detected by unusually high bounce rates from specific a key word. If detected in time the click fraud may be subverted by changing the price for the specific keyword. Finally, real time data provides web site availability monitoring, an additional service that for example, VisiStat offers for free.

The evening ended with a demo of the VisiStat product by Tina Bean, VisiStat Director of Sales and Marketing. The demo involved logging into real live customer web sites. You can see much of the same thing by visiting the VisiStat web site and looking at their live demo. VisiStat is a powerful tool for understanding how a web site is being used. At the same time we could see that it is designed from the end user perspective, so that typical small business user can use it effectively without needing support from an IT department or consultant. Most impressive is the fact that all the power of VisiStat is available to a small web sites for as little as $20 a month.

Sunday, May 13, 2007

People Search Redux

Reading through my last post on People Search, I realize that I did not quite join up the dots. Here is how people search works. It is something I have written about before, and the Search SIG meeting did reveal some new angles.

Firstly, the search engine spiders the web collecting people related information. Next comes the difficult part, arranging the information into profiles where there is one profile per person. This is most difficult for common names like Richard Taylor. Then there are other little variations like nickname, for example, Dick, Rich, Ricky for Richard, spelling variations (Shakespear) and middle names or initials that may or may not be present.

A good profile linked to an identified user is a valuable thing. For example, it can be used to direct advertising to the desired demographic, making the advertising more valuable. As I have noted before this kind of information is most valuable to large internet companies like Yahoo, Google who effectively direct a large part of online advertising.

A profile is much more valuable if the person has taken control of their profile and effectively verified it. So the final step for the people search companies is to create enough awareness that people feel compelled to take control of their own profile. I have run across ZoomInfo profiles that have been verified, so they have started to do this for their specialized audience. Wink and Spock will have to try much harder. I looked at my Wink profile and I was not impressed. I have seen a scarily accurate profile of myself online, and Wink did not come close.

At the meeting, DJ Cline opined that people might be willing to pay money to have their profiles taken down. The panel of search company CEOs disagreed that this was a good model and told us that they would try to talk someone out of demanding that their profile is removed. I think that what this means is that a good profile is actually of more value to others than it is to the target of the profile. Quite apart from that is the thought that blackmail is a "difficult" business model.

Michael Arrington several times expressed the opinion that that Spock would be sued for what they were doing, particularly as one of the example profiles they showed was Bill Clinton with the tag "sex-scandal". I was concerned with a possibility that a profile could be hijacked, as the hijacker could then play tricks to embarrass the target of the profile. A high profile lawsuit or profile hijacking with a lot of attendant publicity could be the catalyst that brings people search to the public attention given what people search gains from an event that makes everyone go out and claim their profile.

However this gets us back to blackmail as a business model. It is one thing for Joe Average to create his own MySpace page. It is quite another thing if Joe Average feels that he has to go out and claim a profile page that someone else has put together without even asking him, just so that he can defend his own good name. People search has had a long history of privacy concerns and it will continue to do so.