Thursday, January 16, 2014
What does it take to be a Data Scientist?
I was looking at some of the job ads for a Data Scientist and found the desired skills/qualifications section quite interesting. The following is a list of companies and the educational qualification, statistical/DM skills, and software knowledge required for a Data Scientist.
SAS and R seems to be very popular. Samsung wants a PhD only. And the ad for CITI looks really amusing - I think they are looking for an army not a person :)
APPLE
Edu - Bachelor’s degree in Computer Science, Statistics, Mathematics or a related field. Master’s degree preferred
St/DM - Prior experience with AB testing (or other experimental design), statistical data analysis, model creation and refinement on big data sets
SW - Working knowledge of SQL (Hadoop/Hive preferred), SAS, R, Python, or Java
CITI
Edu - PhD/Post Doc from a renown institution in any advanced quantitative modeling oriented discipline including but not limited to Machine learning, Statistics, Marketing Science, Operations Research, Econometrics, Stochastic Finance, Distributed and parallel computing, Digital media analytics, etc.
St/DM -
1. Advanced statistical methods including complex multi-variate statistical methods, discrete choice modeling, conjoint based analysis
2. Machine learning including Bayesian methods, reinforcement learning, Neural networks, Support vector machines, Hidden Markov Models, relevance vector machines, Probabilistic/ Evidential Reasoning
3. Operations Research (Queuing, Markov Models, DEA, Integer Programming, Dynamic programming, Stochastic Programming, Game theory)
4. Macroeconomic modeling, Leading indicator analysis, Long term and near term Forecasting, Time series based methods, Bayesian multivariate regression methods, ARCH/GARCH/VAR models and other advanced regression methods, Mathematical economics, System dynamics, Stochastic control, Nonlinear dynamic models, etc. Prior practical industrial scale modeling exposure is a must.
5. Advanced quantitative methods relevant to modeling consumer experience in the digital world. Experience in web log mining for visitor segmentation, visitor behavior modeling, common path analysis, conversion analysis, abandonment analysis, promotion analytics, buzz analysis, sentiment analysis, social networking analysis etc is a must.
6. Latent class models, Multivariate logit/probit/tobit models, Multinomial logistic models, Marketing mix modeling, Hidden Markov models, Conjoint methods, Market research and optimization methods. Prior experience in customer mindset modeling, customer loyalty, customer choice, brand equity, advertisements/promotion mix, etc is a must.
7. Parallelizing existing traditional or modern (machine learning) based algorithms, Randomized algorithms, Simulations and Simulation based methods including Markov Chain Monte Carlo, parallel and distributed simulations, next gen optimization methods, etc. Knowledge of Hadoop/grid based programming for large scale problem solving is a must.
SW - Proven ability in model building and application experience in data mining techniques and tools (SAS and/or other modeling packages like R, Matlab, Mathematica, ILOG etc)
DELL
Edu - Master’s Degree/ Bachelors with 5 years’ work experience in business analytics, reporting, and process design
St/DM - Strong understanding and implementation of predictive / analytical modeling techniques, theories, principles, and practices Specific experience in more than one of: statistical modeling, machine learning and text mining techniques
SW - Excellent knowledge of data mining / predictive modeling tools such as SAS, R, or SPSS
Edu - M.S. or Ph.D. in a relevant technical field, or 4+ years experience in a relevant role
St/DM - Extensive experience solving analytical problems using quantitative approaches
SW -
1. Fluency with at least one scripting language such as Python or PHP
2. Familiarity with relational databases and SQL
3. Expert knowledge of an analysis tool such as R, Matlab, or SAS
4. Experience working with large data sets, experience working with distributed computing tools a plus (Map/Reduce, Hadoop, Hive, etc.)
Edu - Masters, Phd, or equivalent experience in a quantitative field (computer science, physics, mathematics, bioinformatics, etc.)
St/DM - Strong background in Machine Learning, Statistics, Information Retrieval, or Graph Analysis
SW -
1. Some experience working with large datasets, preferably using tools like Hadoop, MapReduce, Pig, or Hive
2. Experience programming in an object oriented language (Java, C++, etc)
3. Knowledge of scripting languages like Ruby or Python, familiarity with web frameworks a plus
4. Comfortable with data analysis & visualization using tools like R, Matlab, or SciPy
MICROSOFT
Edu - Bachelor’s degree in Mathematics, Statistics or Computer Science with strong statistical background
St/DM - Demonstrated statistical analysis skills in a business environment
SW - Knowledge of structured (SQL) and non-structured (Log files) data bases
YAHOO
Edu - PhD in a quantitative discipline (Applied Mathematics, Statistics, Computer Science, Operations Research, or related field)
St/DM - Experience utilizing both qualitative analysis (e.g. content analysis, phenomenology, hypothesis testing) and quantitative analysis techniques (e.g. clustering, regression, pattern recognition, descriptive and inferential statistics)
SW - R preferred
SAMSUNG
Edu - PhD Computer Science Candidate Only
St/DM - Strong background in machine learning
SW - Strong programming experience in Java/C++
WALMART
Edu - Bachelors degree in Statistics, Economics, Analytics, Mathematics, Computer Science, Information Technology or related field and 10 years experience in an analytics related field OR Masters' degree in Statistics, Economics, Analytics, Mathematics, Computer Science, Information Technology, or related field and 8 years experience in an analytics related field
St/DM - Certificate in business analytics, data mining, or statistical analysis
SW -
1. 7+ years experience with statistical programming languages (for example, SAS, R)
2. 7+ years experience with SQL and relational databases (for example, DB2, Oracle, SQL Server)
Tuesday, January 14, 2014
Analytics 3.0
This slide is from the ebook available at http://iianalytics.com/a3/
The full article on Analytics 3.0 is at http://hbr.org/2013/12/analytics-30/ar/1
Thursday, July 25, 2013
All these research studies!
Reading about claims/findings from many of these research studies often makes me shake my head in disbelief. The latest one I saw was titled "Scientists suggest beer after a workout". Below are some of the claims/statements from that article:
My questions:
Friday, March 16, 2012
Wednesday, January 25, 2012
You want more drama? Use less data!
Sunday, January 8, 2012
Sampling or Weights
---
Tuesday, October 4, 2011
Discount/Variety stores productivity
The table has the following columns:
1. Sales
2. YoY change in Sales
3. # of Stores
4. Average Store Size
5. Sales per Store
6. YoY change in Sales/Store
7. Sales per sqft
8. YoY change in Sales/sqft
So, which companies are doing good?
Total Sales is very much dependent on the number of stores and the size of the stores. YoY change in Sales would have been useful if there was information on the change in the number of stores, YoY.
Sales per store is again strongly associated with size of the store. But does bigger store size always mean higher sales?
Walmart has the biggest stores, with an average size of about 162,000 sqft. But Costco, with an average store size of 145,000 sqft has an impressive sales per store - $58 million more than a Walmart store. A Costco store is about 0.9 times the size of a Walmart store but its sales is 1.9 times that of a Walmart store.
Another interesting comparison is Target and Sam's Club. Almost the same store size, but sales at a Sam's Club store is almost twice that of a Target store.
You will see a very similar story if you use Sales per sqft instead of Sales per store. But what makes these analyses incomplete is the absence of average price/item. The comparison between Costco and Walmart is obvious - prices will be much lower at Walmart. But what about Target & Sam's Club? Prices at Sam's Club may still be lower but can that alone explain a Target store having twice the sales of a Sam's Club? Other factors that can be the reason behind a Target store's higher sales will be - optimized layouts (more efficient use of space, more items per unit area), higher traffic, larger proportion of high ticket items, etc..
And how are these companies growing?
Looks good - almost all of these companies are seeing positive trends in both total sales and sales per sqft, YoY. What about Kmart? Sales per sqft is growing YoY but total sales YoY is down. The first thing that comes to my mind - Kmart must have closed some of its non-performing stores. Any other reasons you can think of?
Wednesday, June 29, 2011
How to make an impact with Analytics
Sunday, June 5, 2011
Keep it simple
Is analytics all about actionable insights? Go and check out some of the websites of companies that offers analytics as a service. Chances are extremely high that you will come across either “actionable” or “actionable insights”.
But the truth is a lot of the work done at Analytics companies is not actionable at all. A lot of the daily or regular requirements will be the “good to know” numbers or information, or what many analysts will derogatorily refer to as Reporting work.
To be honest, when I just started my Analytics career I had the same biased or uninformed opinions. Predictive Analytics or Modeling was the only “cool” thing in Analytics. As they say, much water has flowed under the bridge and below are some of the things I have learned along the way, over the years.
All numbers and insights need not be actionable
When working with a new department or analyzing a new customer base, most of the analyses will start with understanding the business or the customers. The simple reports or the exploratory data analysis is going to be a very useful and important guide for future business plans and strategies.
If the results of your analysis can answer a business question, that’s good. And sometimes, it can be very good.
Averages are sometimes the best
I learned this all over again recently. The client wanted to see how different their 2 groups of customers were. As the initial discussion was focused on the purchase behavior or purchase life cycles of these two groups, I jumped into analyzing the customers’ monthly transactions since their acquisition dates - trying to see if these two groups have different buying patterns across their tenures.
When their overall sales didn’t throw up any surprises, I went into sales within specific product categories. By the end of the week, I found a few interesting patterns. But during the second meeting with the client next week, as I was going through the slides one by one (about 7-8 slides) – both my client and I realized that though the buying patterns were very similar, there was a big gap between the lines. Instead of analyzing how the behavior spiked or dipped or flattened out month on month, the biggest and most important analysis would have been a single slide on the averages of the two customer groups over their 1 year tenure. The differences in their average sales, basket size, number of trips, etc. was clearly seen – one single slide, one table – that was what I should have done when my hypothesis or what we all wanted to see was proved wrong by the data.
Understand the drivers – is modeling really required?
When clients say they want to understand the drivers of customer attrition or response, the first thing many analysts will do is to develop a model. But wait a minute, have some patience and ask a few more questions. A model has to be built if the client has a marketing plan or strategy because you will need to score customers for targeting.
What if all the client wants is to understand the drivers? Maybe all she needs is to identify and understand who these attriters or responders are. And to answer that, you don’t need to develop a model that will take a lot of time and money.
All you need is EDA. Means, frequencies and cross tabs based on the target variable (example, responded or not) will reveal things like – 70% of the responders visited the store in the last 3 months, 80% of the responders live within 5 miles of the store, 65% of the responders use a Credit Card, etc. And this will very much answer your client’s questions.
Signing off with:
Karma police, arrest this man
He talks in maths
He buzzes like a fridge
He's like a detuned radio
-- Karma Police by Radiohead
Wednesday, February 23, 2011
The Keyword Tree
Lately, I am getting very interested in data visualization and text mining. Just got the evaluation version of Tibco Spotfire and I would have to say this - it rocks! Beautiful high-quality visualizations, and a good number of features.
Was also playing around with the Keyword Tree from Juice Analytics. A brief description of this tool:
What search words and phrases are driving traffic to my site?
In the word tree visualization, you'll see a frequently used search term at the center. To the right and left, it shows the search terms that are most often used in combination with that word. The words are sized by their frequency of use and colored by bounce rate (or 0% new visitors or average time on site).
I then linked it to my Google Analytics account, and to this blog and here's what I got. Simply BEAUTIFUL!

Wednesday, February 9, 2011
The Looking Glass
(Analysis time period: Jan 2011)
1. Visits vs % New Visits

2. Average time vs Pages/Visit

3. Where are you from?

In January, I have a lot of new visitors - a LOT :) The largest number of visitors/readers are from the US and India, but my readers from Russia seem to be the most ardent - almost 4 pages per visit. I cannot write "Thank you" in russian but google tells me how it sounds like. So to my readers from Russia - spasibo! :)
Friday, December 17, 2010
What's behind your Tree?
Almost all the time, I use a combination of RFM, Decision Tree or Logistic Regression techniques for sorting, profiling and/or scoring customers (hopefully, I can post a separate detailed blog on this).
The best thing about a decision tree is that it has very less assumptions or requirements on the data unlike, let’s say, logistic regression. Another thing is that everyone can understand it! Depending on the software you use, there are a number of different Tree algorithms available with the most common being CHAID, CART and C5.
CART can handle only binary splits (produce splits of two child nodes). It uses a measure of impurity called Gini for splitting the nodes. This is a measure of dispersion that depends on the distribution of the outcome variables. Its values range from 1 (worst) to 0 (best). You get a 0 when all records of a node are falling under a single category level (e.g. all 10,000 customers in a terminal node are responders). This is a purely theoretical example, by the way!
In C5, splits are based on the ratio of the information gain. C5 prunes the tree by examining the error rate at each node and assuming that the true error rate is actually substantially worse. If N records arrive at a node, and E of them are classified incorrectly, then the error rate at that node is E/N.
Information gain can also be simply defined as –
Information (Parent Node) – Information (after splitting on a particular variable)
CHAID is an efficient decision tree technique based on the Chi-Square test of independence of 2 categorical fields. CHAID makes use of the Chi-square test in several ways—first to merge classes that do not have significantly different effects on the target variable; then to choose a best split; and finally to decide whether it is worth performing any additional splits on a node.
CHAID and C5 can handle multiple splits unlike CART. And as far as my own experiences go, I prefer CHAID over C5 as C5 tends to produce very bushy trees.
References:
Data Mining Techniques: Michael J.A. Berry & Gordon S. Linoff
Data Mining Techniques (Inside Customer Segmentation): Konstantinos Tsiptsis & Antonios Chorianopoulos
Wednesday, October 20, 2010
So you thought...?
1. Linear/Pearson Correlation: The most misunderstood term as far as i know. Before doing anything else, check if the 2 variables share a linear relation. Correlation values without a linear pattern is meaningless. And also be aware that in many softwares (including MS Excel), the default is pearson correlation, for which a linear relation between the two variables is a requirement.
2. Significance Test: Many many people into Analytics (?) will never ever understand this or will never try to understand this. Just because you see 2 groups doesn't mean that you can do a significance test. Know something or everything about sampling and designs before talking about significance test.
3. Lift and Cumulative Gains Charts: They are different, period. Don't confuse one with another.


Lift - Without a model, we get 30% of the responders by contacting 30% of the customers. Using a model, we get 60% of responders. The lift is 60/30 = 2 times.
Cumulative Gains - Using the model, if we contact 30% of the customers we get 60% of all responders.
4. Clustering/Segmentation and Profiling: Let's make this simple. Clustering/Segmenting will answer - Can my customer base be broken up into distinct groups based on certain attributes/characteristics? Customers within a group will be very similar to one another while customers across groups will be different.
Profiling will answer - Who are my best customers? What do they purchase? How often? What is their ethnicity, their household size and income, etc.? In many cases, profiling usually follows clustering/segmentation. Who are the customers in Group 1?
Signing off with:
"There must be some kind of way out of here,"
Said the joker to the thief
"There's too much confusion,
I can get no relief"
-- All along the watchtower by Jimi Hendrix
Sunday, July 4, 2010
Sunday, April 18, 2010
The dark side of Statistically Significant
In statistics, a result is called statistically significant if it is unlikely to have occurred by chance.
The use of the word significance in statistics is different from the standard one, which suggests that something is important or meaningful. For example, a study that included tens of thousands of participants might be able to say with very great confidence that people of one state are more intelligent than people of another state by 1/20 of an IQ point. This result would be statistically significant, but the difference is small enough to be utterly unimportant. Many researchers urge that tests of significance should always be accompanied by effect-size statistics, which approximate the size and thus the practical importance of the difference.
Statistically significant is something I come across everyday in my line of work, and to be honest, the most abused and misunderstood term. There are people who assume that if something comes out statistically significant, all their questions are answered and their problems are solved.
An article by Tom Siegfried on Science News throws up interesting facts and assumptions about the Statistical Significance test. I have changed the 2nd part to use a channel effectiveness example instead of clinical trials in the original article.
1. The Hunger Hypothesis
The amount of evidence required to accept that an event is unlikely to have arisen by chance is known as the significance level or critical p-value. In other words, a p-value of .05 means that there is only a 5 % chance of obtaining the observed (or more extreme) result by chance.
So does this mean that you are 95% certain that the observed difference between groups, or sets of samples, is real and could not have arisen by chance? It is incorrect, however, to transpose that finding into a 95 percent probability that the null hypothesis is false. “The P value is calculated under the assumption that the null hypothesis is true,” writes biostatistician Steven Goodman. “It therefore cannot simultaneously be a probability that the null hypothesis is false.”
That interpretation commits an egregious logical error (technical term: “transposed conditional”): confusing the odds of getting a result (if a hypothesis is true) with the odds favoring the hypothesis if you observe that result.
Consider this simplified example. Suppose a certain dog is known to bark constantly when hungry. But when well-fed, the dog barks less than 5 percent of the time. So if you assume for the null hypothesis that the dog is not hungry, the probability of observing the dog barking (given that hypothesis) is less than 5 percent. If you then actually do observe the dog barking, what is the likelihood that the null hypothesis is incorrect and the dog is in fact hungry?
That probability cannot be computed with the information given. The dog barks 100% of the time when hungry, and less than 5% of the time when not hungry. A well-fed dog may seldom bark, but observing the rare bark does not imply that the dog is hungry.
2. Statistical significance is not always statistically significant.
The effectiveness of communication channels (Tele marketing, Emails or Direct Mails) are usually tested by comparing the results from Test & Control groups.
Using significance tests, the channel’s effect (response rate, purchase, etc) on the Test group is pronounced to be greater than the Control group by an amount unlikely to occur by chance.
The standard in most of the significance tests is 5%, and a result expected to occur less than 5% of the time is considered “statistically significant.” So if Email drives a higher response in the Test group than the Control (the non-Emailed group) by an amount that would be expected by chance only 4% of the time, it would be concluded that the Email campaign really worked.
Now suppose Direct Mail also delivered similar results – Test group having a higher response than the Control group, but by an amount that would be expected by chance 6% of the time. In that case, conventional analysis would say that such an effect lacked statistical significance and that there was insufficient evidence to conclude that Direct Mail worked.
If both channels were tested against each other, rather than separately using Control groups - one group getting Emails and another similar group receiving Direct Mails, the difference between the performance of Email and Direct Mail might very well NOT be statistically significant.
“Comparisons of the sort, ‘X is statistically significant but Y is not,’ can be misleading,” statisticians Andrew Gelman of Columbia University and Hal Stern of the University of California, Irvine, noted in an article discussing this issue in 2006 in the American Statistician. “Students and practitioners [should] be made more aware that the difference between ‘significant’ and ‘not significant’ is not itself statistically significant.”
The Control group for Email may be doing a lot worse than the Control group for Direct Mail. The difference in the response rates between Test & Control for Email thus becomes higher for Email, and a statistically significant result was observed.
Signing off, with a quote from William W. Watt:
"Do not put your faith in what statistics say until you have carefully considered what they do not say."
Friday, January 1, 2010
Why Analytics Fail To Deliver?
- The model is more accurate than the data: you try to kill a fly with a nuclear weapon.
- You spent one month designing a perfect solution when a 95% accurate solution can be designed in one day. You focus on 1% of the business revenue, lack vision / lack the big picture.
- Poor communication. Not listening to the client, not requesting the proper data. Providing too much data to decision makers rather than 4 bullets with actionable information. Failure to leverage external data sources. Not using the right metrics. Problems with gathering requirements. Poor graphics or graphics that is too complicated.
- Failure to remove or aggregate conclusions that do not have statistical significance. Or repeating many times the same statistical tests, thus negatively impacting the confidence levels.
- Sloppy modeling or design of experiments, or sampling issues. One large bucket of data has good statistical significance, but all the data in the bucket in question is from one old client with inaccurate statistics. Or you use two databases to join sales and revenue, but the join is messy, or sales and revenue data do not overlap because of a different latency.
- Lack of maintenance. The data flow is highly dynamic and patterns change over time, but the model was tested 1 year ago on a data set that has significantly evolved. Model is never revisited, or parameters / blacklists are not updated with the right frequency.
- Changes in definition (e.g. include international users in the definition of a user, or remove filtered users) resulting in metrics that lack consistency, making vertical comparisons (trending, for a same client) or horizontal comparisons (comparing multiple clients at a same time) impossible.
- Blending data from multiple sources without proper standardizations: using (non-normalized) conversion rates instead of (normalized) odds of conversion.
- Poor cross validation. Cross validation should not be about randomly splitting the training set into a number of subsets, but rather comparing before (training) with after (test). Or comparing 50 training clients with 50 different test clients, rather than 5000 training observations from100 clients with another 5000 test observations from the same 100 clients. Eliminate features with statistical significance but lack of robustness when comparing 2 time periods.
- Improper use of statistical packages. Don't feed a decision tree software with a raw metric such as IP address: it just does not make sense. Instead provide a smart binned metric such as type of IP address (corporate proxy, bot, anonymous proxy, edu proxy, static IP, IP from ISP, etc.)
- Wrong assumptions. Working with dependent independent variables and not handling the problem. Violations of the Gaussian model, multimodality ignored. External factor explains the variations in response, not your independent variables. When doing A/B testing, ignoring important changes made to the website during the A/B testing time period.
- Lack of good sense. Analytics is a science AND an art, and the best solutions require sophisticated craftsmanship (the stuff you will never learn at school), but might usually be implemented pretty fast: elegant/efficient simplicity vs. inefficient complicated solutions.
My additions:
- What is your problem?
Without a real business problem, modeling or data mining will just give you numbers. I have come across people, on both the delivery and client side, who comes up with this often repeated line “I’ve got this data, tell me what you can do?”
My answer to that – “I will give you the probability that your customer will attrite based on the last 2 digits of her transaction ID. Now, tell me what are you going to do to make her stay with your business?”
Start with a REAL business problem.
- What are you gonna do about it?
So if your customers are leaving in alarming numbers, don’t just say that you want an attrition/churn model. Think on how you would like to use the model results. Are you thinking of a retention campaign? Are you to going to reduce churn by focusing on ALL the customers most likely to leave? Or are you going to focus on a specific subset of these customers (based on their profitability, for example)? How soon can you launch a campaign? How frequently will you be targeting these customers?
Have a CLEAR idea on what you are going to do with the model results.
- What have you got?
Don’t expect a wonderful earth-shattering surprise from all modeling projects. Model results or performances are based on many factors, with data quality as the most important one, in almost all the cases. If your database is full of @#$%, remember one thing. Garbage in, Garbage out. Period.
- Modeling is not going to give you a nice easy-to-read chart on how to run businesses.
- Technique is not everything.
A complex technique like (Artificial) Neural Networks doesn’t guarantee a prize winning model. Selecting a technique depends on many factors, with the most important ones being data types, data quality and the business requirement.
- Educate yourself
It’s never too late to learn. For people on the delivery side, modeling is not about the T-test and Regression alone. For people on the client side, know what Analytics or Data Mining can do, and CANNOT do. Know when, where and how to relate the model results with your business.
Tuesday, December 1, 2009
Sunday, August 30, 2009
Means and Proportions with two populations
Call it chance or whatever, but whenever these kind of tasks came up I hear people talking about the t-tests only. No issues as long as you want to compare means or when your target variable is a continuous value. But how or why do people talk about the t-test when they want to compare ratios or proportions? Whatever happened to the Chi-Square tests or the Z-test for difference in proportions?
I did a bit of research on the net, a bit of calculation using pen and paper [very good exercise for the brain in this age of calculators and spreadsheets :-) ], read a very good article by Gerard E. Dallal, and I found the answers.
Going back to our introductory class in statistics, let’s check out the formulae for the t-tests.
1. Assuming that the population variances are equal,
T = (X1 – X2)/sqrt (Sp2(1/n1 + 1/n2) ..........Equation 1
where
X1, X2 = means of sample 1 and 2
n1, n2 = size of sample 1 and 2
Sp2 = pooled variance = [((n1-1)S12+(n2-1)S22)/(n1+n2-2)]
2. Assuming that the population variances are not equal,
T = (X1 – X2)/sqrt(S12/n1 + S22/n2) ..........Equation 2
We have also been taught that the test statistic Z is used to determine the difference between two population proportions based on the difference between the two sample proportions.
And the formula for the Z statistic is given by
Z = (P1 – P2)/ sqrt(P(1-P)(1/n1 + 1/n2)) ..........Equation 3
where
P1, P2 = proportions of success (or target category) in samples 1 and 2
S1, S2 = variances for samples 1 and 2
n1, n2 = size of samples 1 and 2
P = pooled estimate of the sample proportion of successes =(X1 + X2)/(n1 + n2)
X1, X2 = number of successes (or target category) in samples 1 and 2
The test statistic Z (equation 3) is equivalent to the chi- square goodness-of-fit test, also called a test of homogeneity of proportions.
But how different is the proportions from means? The proportion having the desired outcome is the number of individuals/observations with the outcome divided by total number of individuals/observations. Suppose we create a variable that equals 1 if the subject has the outcome and 0 if not. The proportion of individuals/observations with the outcome is the mean of this variable because the sum of these 0s and 1s is the number of individuals/observations with the outcome.
Let's suppose there are m 1s and (n-m) 0s among the n observations. Then, XMean (=P) = m/n and Xi - XMean is equal to (1-m/n) for m observations and 0-m/n for (n-m) observations. When these results are combined, the final result is
∑(Xi – XMean)2 = m(1-m/n)2 + (n – m) (0 – m/n)2
= m(1 – 2m/n + m2/n2) + (n – m) m2/n2
= m – 2(m2/n2) + (m3/n2) + (m2/n) – (m3/n2)
= m – (m2/n)
= m(1-m/n)
= nP(1-P)
So, variance = ∑(Xi – XMean)2/n = P(1-P)
Substituting this in the equation 3 (for Z statistic), we get
(P1 – P2)/ sqrt(Variance/n1 + Variance/n2)), which is not so different from equation 2 (the formula for the "equal variances not assumed" version of t test).
As long as the sample size is relatively large, the distributional assumptions are met, and the response is binomial – the t test and the z test will give p-values that are very close to one another.
And in the case where we have only two categories, the z test and the chi-square test turn out to be exactly equivalent, though the chi-square is by nature a two-tailed test. The chi-square distribution for 1 df is just the square of the z distribution.
The various tests and their assumptions as listed in Wikipedia are given below:
1. Two-sample pooled t-test, equal variances
(Normal populations or n1 + n2 > 40) and independent observations and σ1 = σ2 and (σ1 and σ2 unknown)
2. Two-sample unpooled t-test, unequal variances
(Normal populations or n1 + n2 > 40) and independent observations and σ1 ≠ σ2 and (σ1 and σ2 unknown)
3. Two-proportion z-test, equal variances
n1 p1 > 5 and n1(1 − p1) > 5 and n2 p2 > 5 and n2(1 − p2) > 5 and independent observations
4. Two-proportion z-test, unequal variances
n1 p1 > 5 and n1(1 − p1) > 5 and n2 p2 > 5 and n2(1 − p2) > 5 and independent observations
Sunday, May 31, 2009
Analytics: Reality and the Growing Interest
InRev Systems is a Bangalore based Decision Management Company, which works on Data Based Information Systems. Their interest areas are Marketing Services, Web Information, MIS Reporting, Social Media Services and Economic Research. Bhupendra also maintains a personal blog at Business Analytics.
Introduction
Huge amount of data is collected by any business houses today. There are also certain data collection agencies who have information like economic variables, demographic variables, police fraud list, loan default list, telephone and electricity bill payment history etc. All these data if analyzed, tend to separate people into various similar groups. These groups can be fraudulent group, defaulter group, risk averse and risk takers group, high income and low income groups, etc.
Based on this information, many business decisions can be made in a better and rational way. Analytics leverages almost the same concept but with assumptions:
· The behavior of people do not change with time
· People with similar profile behave similarly
Predictive Modeling and Segmentation are the major components of Analytics. The profile and the behavior of a set of people are taken, and a relation is found out. This same relation is used for building future profiles and predicting the behavior of people having the same profiles. This is commonly done using advanced statistical techniques like Regression Modeling (Linear, Logistic, and Poisson etc) and Neural Networks.
The Business Analytics Services Market comprises solutions for storing, analyzing, modeling, and delivering information in support of decision-making and reporting processes.
Analytics, regardless of its complexity, serve the same purpose – to assist in improving or standardizing decisions at all levels of an organization.
Size and Type of Market
The size of the Analytics market globally is estimated around $25 billion today. It is increasing very fast, doubling almost every five years for the last few decades, with $19 billion dollars in 2006. It is expected to grow to $31 billion by the end of 2011 (source - IDC 2007).
There are many areas for the implementation of Analytics. The most common Analytics practices are Risk Management, Marketing Analytics, Web Analytics, and Fraud Prediction etc. These functions are handled by different organizations in different ways, with most companies maintaining a fine balance with the in-house team, outsourcing partner and consulting project vendors. Such variation makes calculation of market size very difficult.
Risk Management is one of the largest component of the Analytics Industry today, and it is the pioneer component too. The market is huge in the US and Europe; Asia Pacific is coming up fast and it is yet to get full swing in China and South Asia, including Nepal.
Web Analytics (analytics of Web Data) has picked up very fast and it is increasing by more than 20% for the last few years, thanks to the revolution led by Amazon, Google and Yahoo. Web Analytics market is a late entrant but have already passed the one billion mark.
The biggest of all is the Marketing Analytics (MA) and Strategy Science component. This is really huge owing to the effort put by companies like Dunhumby, Axiom and others. MA is critical as it tends to compete directly with Marketing Research Firms and Strategy Consulting firms in the type of work it does. This makes it difficult to calculate the market size but it is worth billions of dollars.
Big Players in the Area
The size of the Analytics market is huge today but the industry is fragmented. None of the core Analytics and Decision Management Company has ever touch a billion dollar revenue mark.
Fair Isaac is the pioneer and the largest Decision Management Company with revenue of around 800 million dollars. The other core DM and Analytics companies are quite smaller. This has happened due to the aggressive moves by IT Services Companies and Information Bureaus to acquire Analytics companies.
Experian, Transunion and Equifax are three major bureaus in the US while there are others too - Innovis, Axiom, Teletrack, Lexis Nexis etc. Each of these bureaus has analytics services as their offerings. Apart from these the BI majors like SAS, SPSS, Salford Systems etc. offer Analytics Services.
The Indian Analytics Market is small but growing. This can be justified by the more than double salary hike rates in Analytics compared to the Software domain. The outsourcing shops and India-focused companies both have mushroomed in the last five years in India while the problem remains in getting proper talent and retaining them.
Major Banks like ICICI and HDFC have their strong in-house analytics unit while SBI has partnered with GE Money for Analytics support for their Cards portfolio. The smaller banks are yet to start.
Other majors in non banking industries are also using Analytics through Outsourcing or Consulting. Airtel and Reliance are leading the way for the use of Analytics in India in Telecom.
Scope on how it can grow: Indian Context
The Analytics Industry is growing fast. It has the scope of forming a separate process across industries like HR, Operations, and IT System. IT is now in the early stage of process metamorphosis, where each process start through consulting, grow through in-house establishments and finally settle down as outsourcing to third parties.
The future growth depends on the approach of major players and consolidation of the industry. All this will make it a high value and high growth industry, where players can provide high quality products and services while maintaining their profitability.
Another big challenge is the supply of quality man-power and training. Today India neither has good institutes training people on Analytics nor has it got an Infosys for Analytics (who can employ and train huge number of freshers). Even the number of good Statistics and Mathematics Institutes in India is less.
Amidst all these challenges, India is positioned fairly well in the world today, and it will be interesting to see if it can become a Knowledge Process and Analytics hub in the days ahead.
Monday, May 18, 2009
A Tale Of Two Banks and One Telecom Service Provider
My account getting credited above
My account getting debited above
Salary credited to my account #
Cheque deposited in my account bounced
Account Balance above
Account Balance below
Debit Card Purchases above
Messages for the above can be received through SMS, Email or both.
Citibank calls it Alerts, and they offer the following options:
Withdrawal balance by account
Withdrawal balance by account
Time deposit maturity advice
Cheque Status
Cheque Bounce Alert
Time Deposit Redemption Notice
Cheque dishonor
These messages can be received through SMS, Email or Both depending on the alert type. Also, for some of the alerts, message frequencies can be chosen as Daily, Weekly, or Monthly.
At first glance, it seems that Citibank offers more options but a closer look will reveal ICICI has done more research and come up with a better offering. The Citibank alerts are based on how frequently you want them while ICICI’s alerts are based on particular (defined by the customer) credited and debited amounts. ICICI’s options make more sense as I don’t want alerts daily or weekly, but only when I have made a transaction.
And because of the lack of options, I continue to receive daily alerts on my Citibank account balance irrespective of the fact that the balance remains the same or I haven’t done any transaction for a month. Also, the system at Citibank looks like a typical CRM system while the one at ICICI looks more like a BI system.
Now, let’s discuss their Credit Cards and their Customer Analytics.
I have been using an ICICI Gold credit card for almost 3 years now. I used it to pay all my bills – electricity bill, mobile bill, internet bill, shopping bills….and I paid all dues in time. I applied for and got the Citibank Gold card about a year after I got my ICICI card. I used the Citibank Gold card for 2-3 purchases, paid all the dues in time, and Citibank increased my credit limit every time.
Encouraged by their response (I got very nice emails from their customer service) and actions (increase of credit limit), I started using the Citibank Gold card more frequently. In a few months’ time, I got a free-for-life Citibank Platinum card with all these attractive features and benefits. I even got an invitation to join a wine club though I’m more of a rum and whiskey guy. Don’t blame them though; getting your hands on such kind of consumer lifestyle/preferences data will be next to impossible in India.
I have now almost forgotten the ICICI Gold card; I use it very rarely these days. The credit limit given to me 3 years back still remains the same. And I have never received a single email or communication from ICICI. I also know that ICICI bank outsources its Customer Analytics to an Analytics Service Provider in Mumbai. So where is the up sell analytics? Doesn’t their data show the fact that I am “almost” leaving now? That again reveals the fact that customer spending information is not at all analyzed and they are not doing much about customer churn either.
On a different note, I am a post-paid mobile customer of India’s largest telecom service provider, Airtel. Every month, whenever my bill is generated I receive about 6 SMSs from Airtel within the next 5-6 days. The messages can be summarized as:
1. Your bill has been generated…
2. You can view your bill at your online account…
3. Your bill has been emailed to abc@gmail.com and the password to open it is… and if you haven’t received it… (this is actually 2 SMSs because of the length of the message)
4. Your bill amount is XXX…
5. Your bill amount is XXX and the last date of payment is…
For the last 3 years or so, I have been paying all my bills before the due date. Once I got so irritated that I emailed their customer service, and their reply? These are server generated messages and we can’t do anything about it. I got more irritated and asked my email to be forwarded to Airtel’s CRM, Business Intelligence, Analytics or whatever the team there like to call themselves!
I got just one SMS alert when my next month’s bill was generated. But it was back to square one from the 2nd month onwards.
My question is which “smart” manager came up with the idea that 6 SMSs should be sent to all their post-paid customers every month? Why doesn’t one SMS saying “Your bill amount of XXX has been generated and the due date is ABC” suffice? Has anyone at Airtel calculated the cost of sending 5-6 SMSs to all their post-paid customers, every month? Shouldn’t the last SMS be sent only to those customers who have a habit of making late payments? Do they send the second SMS to those customers who don’t have online accounts too? And why can’t these alerts be customized based on a customer’s usage and payment behavior?
I can give more examples of Indian companies in the retail, entertainment, and services sector that are not doing nothing or very little about all the customer data they have, inspite of mentioning or advertising that they use BI & Analytics. So how mature is the Business Intelligence, and CRM Analytics setup at Indian companies? And how skilled or knowledgeable are the senior people associated with it?




