credit risk analysis

已取消 已发布的 6 年前 货到付款
已取消 货到付款

1. Use the data in the contents tab labeled Prosper.csv. Create a good ‘risk’ model.

a. For the model building use SAS eMiner.

b. Greater thoroughness in the descriptive statistics, the writeup/documentation, the formatting of the report, and in how you came to your conclusions.

2. Descriptive stats. For each variable in the dataset summarize it (max,min,mean,etc) and make a histogram. Label and format appropriately. Do the same for any variables you create. Comment on anomalies in your data and what you did to address them. Defend your reasoning. It might be easiest to just make one page per variable. You should be very thorough in this area. There is code and some of the midterm papers and presentations posted in d2l for you to learn from. Whether you use R or SAS for the descriptive statistics is unimportant however it would seem to me that R would be best given you have worked in that environment more at this point.

3. Create a logistic regression model. In your attempt you should justify why you used the variables you did and for each rejected variable explain why you rejected it. Many groups on the midterm did a little hand waving here which is fine but now that you have gone through it you need to be more thorough and specify each variable. The variables you use in the final model should be binned and transformed to WOE. This is a change from the midterm. On the midterm there were several variables that had ‘goofy’ values. There were 999’s. There were odd values of outliers. There were categorical variables that were too many categories. A WOE transformation can ‘fix’ all of this. You can do the WOE transformation in R if you like or in eMiner the ‘interactive grouping’ node is specially made just for this task. If you use the interactive grouping node DO NOT simply accept the base grouping SAS throws at you. Go in and adjust the bins/groups in a logical fashion. Use the WOE_variables as the inputs to the model rather than the original variables.

4. For modeling you should sample the data into two groups. Typically I use 60/40 but you can use 70/30 or 50/50 if you like.

5. In addition to the above logistic regression you should attempt a more advanced model in eMiner and report its results in comparision. This might be a form of Neural Net, LARS, Dmine regression, etc. There are a number of them in eMiner. This is to show how different models can be applied in a straightforward manner once the problem has been set up correctly. For the ‘other’ model it is preferable you not use the WOE variables since those are special purpose things for logistic regression. It isn’t invalid to use WOE but it is more interesting to compare without the WOE.

6. Submit your report and any code you made. Neatness, organization, style, etc will be part of the consideration. The two nodes that can help you here are the ‘model comparison’ node which makes nice output and ROC curves and tables and such. The second node is the ‘reporter’ node which outputs a complete log of the process. In the reporter node when you are selecting the parameters you can make your life easier by telling it to create ‘rtf’ which is a ‘rich text file’ which opens in Word. Additionally you can alter the quality/type of output as shown in class.

7. We will ‘present’ during the final exam period. I will randomly draw names to present. You should prepare and submit a presentation that is not longer than 15 minutes. In other words we’ll have enough time to do several.

Final deliverables then are:

a. Code file as appropriate (or just put it in an appendix of the report).

b. Report.

c. PowerPoint presentation.

R 编程语言 SAS

项目ID: #13830753

关于项目

3个方案 远程项目 活跃的6 年前

有3名威客正在参与此工作的竞标,均价$97/小时

ashitjha7

20+ years industry work experience in the area of IT, Finance & Banking. Also, I have 4+ years of experience in Economics, Business Analytics and Advanced statistics projects using software such as Python, R, SPSS, Min 更多

$160 USD 在4天内
(1条评论)
2.6
ExpertzWorld

“Time. It’s what we writers fight for. Without it, we have no hope of bringing our written creations to life. We need time to study, time to read, time to ponder, time to dream, and of course – time to write.” I am 更多

$30USD 在1天里
(0条评论)
0.0
techfinity3

DDear Prospect Hiring Manager. Thank you for giving me a chance to bid on your project. i am a serious bidder here and i have already worked on a similar project before and can deliver as u have mentioned I have 更多

$208 USD 在6天内
(0条评论)
0.0