Skip to content

Latest commit

 

History

62 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Data Science Portfolio


2018 - present

UPDATE: I have made some of the repositories private that i can share when requested.

In this repository, I have put some of the Data science projects, I have worked on or currently working. All the projects will mostly focus on utilizing Machine Learning and Deep Learning Techniques to design data science or statistical models, that either solves a problem or discovers important information about the data.

NOTE:

  • I am in the process of Documenting the projects like how to reproduce results and more.
  • Most of it can also be found in the respective Jupyter notebooks even if the README files are not yet updated

Click on the projects to see the documentation and code. (Project Repository README are being built,meanwhile check the ipython notebooks, they are sufficiently documented)

Projects:

  • Predicted Activity based on Accelerometer and Gyroscope readings from smart phone | 1. Walking | 2. WalkingUpstairs | 3. WalkingDownstairs | 4. Standing | 5. Sitting | 6. Lying
  • Fitted Logistic Regression|Linear SVC |rbf SVM classifier|DecisionTree |Random Forest |GradientBoosting DT on featured data
  • LSTM model trained on Raw data
  • t-sne Visualization


  • Predict similar apparel and recommend those apparell based on which apparel(query product) the user is watching.
  • Data acquired in policy compliant manner using Amazon API.
  • Text based product recommendation, i.e using product title, brand, description, color and price.
  • Results with Text Featurized using BOW,TF-IDF, IDF,average word2vec,IDF weighted word2vec are subjectively compared.
  • Product Image based model using VGG-16 CNN also trained.
  • Final model build as a weighted Nearest neighbor model using Image,Title,Brand and Color
  • Final model uses title:Idf-Word2vec, brand:one hot encoding, colour:one hot encoding, image: VGG-16 CNN
  • Query product

  • Suggested product


  • Predicted number of pickups, given location cordinates(latitude and longitude) and time of the day.
  • Time-series forecasting and Regression.
  • New York citymap is divided into several regions based on region radius and density of passenger.
  • K-means clustering used for geometric division of the city map
  • Time series data is divided into 10 minutes interval.
  • Frequency domain features are also used along with time domain data to improve the model.
  • Linear Regression, Random Forest and Gradient Boosted Decision trees(GBDT) models were trained and tested.
  • GBDT model performed better than other models
  • Plots shows how the map of the city is divided into different regions, each color is a region


  • Identify which questions asked on Quora are duplicates of questions that have already been asked.
  • Modelled as classification problem with class as duplicate and not duplicate.
  • 15 features are extracted, some of them are:
  • word_common: Number of common unique words in question1 and question2.
  • word_total: toal number of words in qs1 + total number of words in qs2.
  • word_share: (word_common) / (word_total)
  • common_word_count / min( len(q1 word), len(q2 word))
  • Along with the above extracted features question text is featured using tf-idf.
  • Fitted Logistic Regression, Linear SVM and XGboost gradient boosted decision trees.
  • Question appearance count.

  • Word Cloud for Duplicate Question pairs.

  • Word Cloud for non-Duplicate Question pairs.


  • Predicted the rating that a user would give to a movie that they have not yet rated.
  • Minimized the difference between predicted and actual rating (RMSE and MAPE).
  • Generated user-user and movie-movie similarity matrix.
  • Featurized the data further into 13 features.
  • Models trained: Surprise-baseline, Surprise-knn-baseline, Surprise-SVD++, XGboost
  • Importance in Gradient Boosted Decision Trees


  • Classified the given genetic variations/mutations based on evidence from text-based clinical literature into 9 classes.
  • Dataset has 3 important attributes 'Gene' 'Variation' and 'Clinical text'.
  • Dataset is imbalanced, modeled with and with out balancing to compare results.
  • Clinical text featurized using tf-idf. Gene and variation are featured using both One-hot encoding and response coding.
  • Dimensionality reduction of text vector using truncated SVD.
  • Fitted and compared various models including Logisting Regression, Linear SVC , Naive Bayes and Random forest
  • Experimented with stacking the models and also to use max-voting on the results of the models.
  • Interpretability of results is important here so a calibrated classifier is stacked with the models
  • Ploted Confusion,precision and recall matrices to have more Interpretability.
  • Random-forest-precision matrix

About

Collection of some of my datascience projects

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages