Apache Hadoop (projects)

QUESTIONS setInputFormat comparator top k frequent words HADOOP SYSTEM Apache Hadoop is an open source software framework for storage and large scale processing of data-sets on clusters of commodity hardware. HDFS(Hadoop distributed file system): data storage (data split and data replication) Map Reduce(data processing): how to leverage job; how do nodes communicate; how to deal with node…

Time Series Analysis

Time Series Analysis

TIME SERIES BASICS Difference between regression and time series: time series are not necessarily independent and not necessarily identically distributed.  They are lists of observations where the ordering matters.  Ordering is very important because there is dependency and changing the order could change the meaning of the data. Characteristics: Is there a trend,  on average, the…

Stats and Probability Theory

Stats and Probability Theory

How to choose a statistical model? Are My data Normally Distributed? Problems: Excess kurtosis (forth moment, very big tails, due to extreme values away from the mean) Excess skewness (third moment, lopsided) Others: lognormal (a RV whose logarithm is normally-distributed), uniform, weibull, exponential… Routine: Histogram (largely depends on the bin size) Stem and leaf plots…

Representation Learning

Representation Learning

From the perspective statistics, many of the methods discussed below can be considered as multivariate analysis methods. I would like to refer it as dimension reduction/unsupervised learning methods from the understanding of machine learning, even not all of them are typicall used as dimension reduction techniques. PRINCINPLE COMPONENT ANALYSIS (PCA) project the original data…

Academic Activities and Meetups

Academic Activities and Meetups

By attending academic activities,  we can get access to the most cutting-edge research topics and techniques. Even though we cannot accomplish the cool works presented by those big names, it is still worth our time to understand the basic ideas of the research community, which may help us to keep the curiosity about the area.…

When we talk about data science, what we talk about

When we talk about data science, what we talk about

One and a half year ago, I did not know what is logistic regression. Now, I love machine learning and data science, and decide to delve into it for my future career. I can still remember the first time I audited a machine learning class in Harvard. Without any basic knowledge in algorithms, I still found…