data science · machine learning

Time Series Analysis

TIME SERIES BASICS Difference between regression and time series: time series are not necessarily independent and not necessarily identically distributed.  They are lists of observations where the ordering matters.  Ordering is very important because there is dependency and changing the order could change the meaning of the data. Characteristics: Is there a trend,  on average, the… Continue reading Time Series Analysis

data science · machine learning

Stats and Probability Theory

How to choose a statistical model? Are My data Normally Distributed? Problems: Excess kurtosis (forth moment, very big tails, due to extreme values away from the mean) Excess skewness (third moment, lopsided) Others: lognormal (a RV whose logarithm is normally-distributed), uniform, weibull, exponential… Routine: Histogram (largely depends on the bin size) Stem and leaf plots… Continue reading Stats and Probability Theory

data science · machine learning

When we talk about data science, what we talk about

One and a half year ago, I did not know what is logistic regression. Now, I love machine learning and data science, and decide to delve into it for my future career. I can still remember the first time I audited a machine learning class in Harvard. Without any basic knowledge in algorithms, I still found… Continue reading When we talk about data science, what we talk about