Be yourself; Everyone else is already taken.
— Oscar Wilde.
This is the first post on my new blog. I’m just getting this new blog going, so stay tuned for more. Subscribe below to get notified when I post new updates.
Be yourself; Everyone else is already taken.
— Oscar Wilde.
This is the first post on my new blog. I’m just getting this new blog going, so stay tuned for more. Subscribe below to get notified when I post new updates.
This week in BAN 7002, we return from “Spring Break” to conclude our learning modules in the SAS programming language and move on to learn R programming. As we submitted our Home Price Prediction Midterm projects that helped us tie all our SAS skills together into one, cohesive and relevant environment we transition into the world of R.
This week we were introduced to R, in its similarities and differences of the language as compared to SAS and other programming languages. R is known to be more complex than SAS, although it provides the user more flexibility. Another major difference between the two languages is R’s ability to generate complex graphical models. R is more commonly used in Financial markets and for data imports and cleansing, whereas SAS is better for predictive analytics, business analytics and the like.
I am cautiously approaching R for the reasons stated above. I know it will be a beast of a language to learn but with the right resources, will be a valuable tool to have in my skillset.
In week 7 of BAN 7002, we tied all of the SAS skills we have been learning thus far into one conhesive project. This involved using all the DATA and PROC steps we have utilized so far to organize, cleanse, re-structure and model data on House prices and the various factors that make up a home’s value in order to build a model that can adequately predict house prices using regression modeling techniques. This project incorporates all the skills into a real-life, relatable way that this coding language can be used to efficiently and effectively manipulate data to produce meaningful information.
The challenge of this project will be to rely on the skills we have learned but identify which data step and proc step to use to most efficiently reach the solutions we are seeking. All in all, this will be a very cohesive way to sum up what we have learned so far in this course in SAS before we move on to the R programming language.
This week in BAN7002 we explored using linear regressions to perform estimations on datasets. We learned two new commands: PROC REG and T-TEST.
I found this to be very helpful in testing hypotheses when dealing with data and using the results to compare against a baseline. We used this concept to perform comparative regressions across multiple datasets that contained information about red and white wines, specifically using a TEST data set to compare its accuracy against a SCORE (baseline) data set.
This concept will be extraordinarily helpful in the data analysis world in the future when performing benchmark analyses. This provides much more insight than the typical 5 number summer (mean, median, mode, max, min) that only provides very high level comparison metrics in analyzing a dataset against a baseline.
This week, in BAN 7002 we reviewed our mastery of the basic DATA and PROC functions in SAS, and applied a new concept of clustering using PROC fastclus.
I really enjoyed the assignment this week because it allowed me to gain confidence in my ability to use the SAS programming language that we have been learning so far. The assignments thus far have seemed rather detail-oriented and complex, and I found myself focusing more on the specific data we were working with rather than the code language we were using to manipulate and work with the data. In this assignment, I was able to focus more on the programming language and found myself breezing through the PROC steps and configuring the data to show relationships based on two or more characteristics for comparison and analysis purposes. Now that I have a better feel for the language, I am confident that I can use this program for real-life data analyses in my work environment and future career endeavors.
We also learned a new PROC function, known as PROC fastclus. This performs disjoint cluster analysis on the basis of distances computed from one or more quantitative variables. K-means clustering aims to partition n observations into k clusters in which each observation belongs to the cluster with the nearest mean. This allows us to get another look at data spreads by visualizing how data can be grouped based on similar means.
PROC fastclus was a little more difficult for me to understand conceptually than the other PROC commands we have learned. I am not sure how this will play a role in real-life data analyses, but look forward to exploring the topic and the context it provides in live datasets.
This week we covered the topics of SAS Macros language. SAS Macro language can be a very versatile and useful way to process data in SAS efficiently. SAS Macros allow for programs that are more dynamic and can manipulate data faster. SAS macros are made up of macro variables and macro programs which together form commands that can supplement the SAS program. Some of the more common Macro commands are %PUT, %RETURN, and %END.
Although macro language can make programming easier by solving the problem of having to re-enter the same code multiple times, and can allow for more efficient data manipulation, it is notoriously a more difficult topic in learning SAS. Macros are complex and difficult to master. I definitely found myself getting frustrated with the macro language as I went through the course material this week. It requires extreme attention to detail and a thorough understanding of the commands built into the language as well as the output being generated. I found myself Nevertheless, once macro language is mastered, I can certainly see it being a useful resource when completing projects that involve repeated use of a single command, like those that we did in the last project involving opioid prescriptions in North Carolina.
This week the focus of our studies involved exploring SQL (structured query language) and its compatibilities and functions within SAS software. Essentially, SQL is a database management language that is among the database languages built into SAS. SQL has the ability to manage and organize structured data, and SAS can use the simple SQL commands to product additional data management using the PROC SQL command.
PROC SQL is yet another function that SAS uses to effectively and efficiently manage data to generate relevant information. It also further proves the broad capabilities and complexity of the SAS software in that it had the ability to comprehend different query languages. SQL is often considered a more straightforward and simplistic language, which makes using SAS all the more user-friendly. Nevertheless, it is an entirely separate language with different syntax rules and commands, which can stand as a point of difficulty for someone learning SAS. I, myself, found it tricky to keep the two languages and the rules associated with each clear.
This week, our project focused on combining and organizing several data sets concerned with Opioid use rates to reach more precise and conclusive analyses around the investigation into the opioid abuse crisis. The procedures we utilized would have been very useful in several projects I have worked on in the past, including projects where the goal was to collect, organize and analyze patient data for hundreds of physicians that were participating in a specific clinical trial. Previously, I would use simple Excel commands to search and select desired data, which proved to be a much more limiting and time consuming method of data withdrawal.
This week, we focused on using the SAS programming techniques we have learned so far to conduct a real-world analysis of financial portfolio performances. This project included merging several groups with large data quantities, filtering and organizing them by date, standardizing the data to be on the same scale, and using portfolio weights to accurately depict the daily returns for each portfolio against a benchmark. To summarize our findings, we created charts and graphical visualizations to compare the actual returns versus the expected returns against the benchmark. This project served as a prime example of a real-life situation where SAS programming skills and the ability to effectively execute data analysis would be highly valued. While working through this scenario, I found myself wanting to tackle several steps in the programming language at once, which would often result in an error that I would have to dig through my work to decipher. For example, when executing the merge by date step, it required inputting each asset and the associated date function one by one with proper syntax to avoid any simple mistakes. This was a primary example of why it is often worthwhile to do each programming step one at a time to avoid simple mistakes and ensure the language is properly formatted.
The experience we gained through this case scenario will be useful in future endeavors where comparing large groups of data, which may or may not be standardized in its original form, and using the information extracted from the analysis to make informed and validated business decisions. Looking back on my own career endeavors, I could see this program being very helpful in previous medical case studies I worked on that involved filtering through thousands of patient records with infinite data tags for a more precise and specific patient cohort. Prior to working with SAS, I primarily relied on Excel to execute data filtering, which does not have nearly the same data capacity or analytical capability as SAS does.
During week 1 of my Wake Forest University Masters in Science of Business Analytics program, I began my studies of Data Analytic Software. This week, we were introduced to the SAS program. SAS is a data analysis software that allows users to process, sort and organize data, compute values, generate reports and graphs, and present data results. As an introduction to the programming software, we were presented with an overview on how to write a basic SAS program, including the two primary steps: DATA steps and PROC steps. Data steps create or modify tables, whereas PROC steps enable you to process data and present descriptive statistics and informative graphs. Furthermore, I learned the basics about program statements, including the essential elements of a functioning program statement. As we further explore the SAS data program over the course of this semester I will ultimately be able to write SAS code to seamlessly extract, organize, and configure datasets for any statistical or informative purpose.
After a rough introduction to the program, I understand the complexity and the intricacy of the software and anticipate that the programming language will be initially difficult to learn but will prove easier to master. Along with such intricacies that the software maintains comes a precise attention to detail with the code and SAS program language. As our instructor and the modules emphasized, there will be initial frustration involved with the minute details involved in writing the programs, but will take time and practice. This stood out to me as something I ought to be wary of as I dive into the software.
I am highly anticipatory in mastering the SAS program software, as I know it will prove to be remarkably helpful in my data analysis endeavors going forward. The software is far more capable of organizing and configuring large data sets to present analyses and generate analytical results than any other software I have previously utilized, such as Excel. I look forward to learning the program in and out so that I have the capability to run SAS programs in my career.
This is an example post, originally published as part of Blogging University. Enroll in one of our ten programs, and start your blog right.
You’re going to publish a post today. Don’t worry about how your blog looks. Don’t worry if you haven’t given it a name yet, or you’re feeling overwhelmed. Just click the “New Post” button, and tell us why you’re here.
Why do this?
The post can be short or long, a personal intro to your life or a bloggy mission statement, a manifesto for the future or a simple outline of your the types of things you hope to publish.
To help you get started, here are a few questions:
You’re not locked into any of this; one of the wonderful things about blogs is how they constantly evolve as we learn, grow, and interact with one another — but it’s good to know where and why you started, and articulating your goals may just give you a few other post ideas.
Can’t think how to get started? Just write the first thing that pops into your head. Anne Lamott, author of a book on writing we love, says that you need to give yourself permission to write a “crappy first draft”. Anne makes a great point — just start writing, and worry about editing it later.
When you’re ready to publish, give your post three to five tags that describe your blog’s focus — writing, photography, fiction, parenting, food, cars, movies, sports, whatever. These tags will help others who care about your topics find you in the Reader. Make sure one of the tags is “zerotohero,” so other new bloggers can find you, too.
Have a WordPress.com account? Log in now.