Hey there! Iβm Neha, a Data & ML Specialist who finds joy in uncovering stories hidden within data. β¨π·
Iβve spent 3+ years designing data pipelines, building ML models, and transforming raw data into meaningful insights. I recently earned my Masterβs in Applied Data Science from San JosΓ© State University, where I also served as a Graduate Teaching Assistant for Machine Learning and Distributed Systems.
When Iβm not coding or analyzing trends, Iβm probably dancing to my favorite tunes or training for my next run - because balance is everything. πβπ»
π I'm currently exploring Explainable AI and Foundation Models
π― I'm looking to collaborate on ML/AI projects that solve real-world problems
π¬ Ask me about Data Engineering, ETL Pipelines, Machine Learning, Deep Learning
π MS Applied Data Science @ SJSU
πΌ Ex-DXC Technology Data Engineer
β‘ Fun fact: I've analyzed everything from fashion trends to Higgs Boson particles!
|
Developed explainable reasoning approaches for educational problem-solving using LLMs. Built end-to-end language model pipeline. |
Explainability-driven evaluation framework for code-generating LLMs. Analyzed 10 models with Python, uncovering actionable insights through execution testing. |
|
Built end-to-end ETL pipeline ingesting data from NOAA GHCN-D AWS S3. Processed 10GB+ data with scheduled jobs, queries, and visualization. |
Developed automated ETL pipelines analyzing global food insecurity patterns. Applied K-means clustering and regression across 217 countries. |
|
Developed and compared multiple object detection models including YOLO, Faster R-CNN, and Detectron2. Handled real-world images and videos using PyTorch and OpenCV. |
Designed and trained GAN and LSGAN using Fashion-MNIST. Explored Wasserstein GAN variant to generate realistic fashion images with improved visual quality. |
|
Engineered features and optimized training on 10GB+ data with PyTorch for personalized recommendations. Implemented hybrid model architecture. |
Developed NLTK and TensorFlow-based detection system for semantically similar questions. Improved accuracy by 5% through novel graph-based features. |
|
Utilized Apache Spark to analyze 8GB ROM data from high-energy collision experiments. Achieved 55% predictive accuracy using XGBoost hyperparameter optimization. |
Integrated VTA ridership and NOAA weather data using Pandas and Airflow. Trained Random Forest models achieving 55% accuracy and identified weather impact. |
π Master of Science in Applied Data Science (Data Analytics)
San Jose State University | Aug 2023 - May 2025
π Bachelor of Engineering in Information Technology
Jawaharlal Nehru Technological University | 2014 - 2018
π Certifications:
- Women in Software Engineering (WISE) Program
- Programming for Everybody (Getting Started with Python)
- Performance and interpretability analysis of code generation large language models - Research on explainability of LLMs in code conversion tasks
- Explainable Use of Foundation Models for Job Hiring - Framework for transparent AI in recruitment
π¬ Graduate Teaching Assistant | San Jose State University
Machine Learning Technologies & Distributed Systems | Aug 2024 - Present
π Data Engineer | DXC Technology
Built scalable ETL pipelines and ML solutions | Aug 2018 - Jan 2022
I'm always excited to collaborate on innovative ML/AI projects, discuss data science trends, or explore opportunities that create real-world impact. Feel free to reach out!
