Comments on Datalligence: Fraud Prediction - Decision Trees & Support Vector Machines (Classification)

have you considered random forrests?

2008-12-22T19:42:00.000-06:00

have you considered random forrests?

Bhupendra wrote: But issues with SVM and NN is tha...

2008-12-17T15:12:00.000-06:00

Bhupendra wrote: But issues with SVM and NN is that they too much overfit with the data.

Could you expand on this, or provide references? Is there some reason that early stopping (in the case of neural networks) or constraining the number of hidden nodes do not address the over-fitting issue?

-Will Dwinnell
Data Mining in MATLAB

Oracle Data Mining, release 10.2 introduced a new ...

2008-12-11T15:14:00.000-06:00

Oracle Data Mining, release 10.2 introduced a new variation of the Support Vector Machine algorithm, called 1-class SVM. 1-class SVMs are specially designed for fraud and anomaly detection when you lack examples of the "rare events". In those cases, you can do various tricks w/ stratified samples, ROC, and/or cost matrix for false pos/neg costs, but all struggle.

1-class SVMs work on the principle of learning what is considered "normal" e.g. expenses, phone calls, employees, etc. If your training data contains examples of the rare events, remove them for 1-class SVM model building. Applying the 1-class SVM model then scores each record on the likelihood that it is "abnormal". Oracle Data Mining's 1-class SVM can also mine transactional (nested data), unstructured data (i.e. text), and star schema data. You can read more about them in the documentation available online at: http://www.oracle.com/technology/products/bi/odm/index.htmlhttp://www.oracle.com/technology/products/bi/odm/index.html and in the OTN web site tech info posted at:

Hope this helps!

cb

yup bhupen, SVM & ANN would be a more appropri...

2008-12-01T23:23:00.000-06:00

yup bhupen, SVM & ANN would be a more appropriate comparison but ODM (10g) doesn't have ANN :-)
i tested the models on a few more datasets; performances dropped in both cases, with a lot more in the case of DT.

have heard a lot about MS SQL Server DM, would love to check it out. yeah, you can be sure that the 2 biggies are going to be major players in the DM market soon, 'coz the world's data reside in their systems!

sandro: SPSS has 4 algos for Decision Tree. I buil...

2008-11-27T06:23:00.000-06:00

sandro: SPSS has 4 algos for Decision Tree. I built and tested a DT model on the same dataset with the same inputs (same transformation) using CHAID as the tree-growing criteria. The accuracy was in the 50's range. I will check on the other 3 algos in SPSS and let you know if i have time:-)

Great article. Thanks for the details.I have two f...

2008-11-26T03:37:00.000-06:00

Great article. Thanks for the details.

I have two feedback on the article.

1. I am not surprised that SVM over-performed DT, as it almost always does. Neural Networks would have been a better comparison. But issues with SVM and NN is that they too much overfit with the data. It will be interested to see if you have similar results in intime and outtime validation data sets. I have always seen a significant drop in performance for SVM models.

2. It is really nice to see that Oracle's tools provide so much of facilities. I worked on Microsoft SQL Server Analytics Services for a week, and was impressed with their tool too. With biggies joining this market, it is going to be interesting.

-- Bhupendra

you are absolutely right jonathan. this post was m...

2008-11-25T23:03:00.000-06:00

you are absolutely right jonathan. this post was meant to be a comparison of the 2 techniques on a very very basic level (i mentioned that too).

any model's performance has to be decided on a whole lot of parameters - true +ves/-ves, false -ves/-ves, misclassification costs (if info is available), gains/lift charts....

also, in a majority of fraud prediction problems, the emphasis is on true positives as the cost of fraud typically tends to be higher (i'm generalizing here!) than the time/cost/inconvenience arising from the false positives.

I don't necessarily agree that SVM outperforms dec...

2008-11-25T21:42:00.000-06:00

I don't necessarily agree that SVM outperforms decision trees. While SVM correctly classified a larger number of the fraudulent cases, it also had a greater number of false positives (non-fraud classified as fraud). The better model depends on the relative costs of misclassifying fraud as non-fraud and misclassifying non-fraud as fraud. Alternatively, you can adjust some settings (like the prior probabilities or misclassification matrix on the trees) until both methods correctly predict the same number of fraudulent cases - the better model will then be the one with fewer false positives.

Thanks for the details Romakanta. Did you obtain t...

2008-11-25T11:12:00.000-06:00

Thanks for the details Romakanta. Did you obtain the same kind of results with Decision Trees (using SPSS Answer Tree) than your current results (72%)?