AI Research Engineer and Computer Engineering undergraduate at SSUET, working as a Software Engineer at a UK-based firm. My research focuses on LLM Optimization & Efficient Inference, Distributed ML Systems, Post-Training Alignment, and Mechanistic Interpretability.
I study and implement intelligent systems from first principles — spanning custom CUDA/ROCm kernels, distributed training workflows, post-training alignment (GRPO, DPO), and empirical hidden-state probing.
2,400+ HF Model Downloads • 4+ Research Papers • 5+ Systems & ML Projects • AMD ROCm Certified
- LLM Systems & Inference: KV-cache optimization, Speculative Decoding, Continuous Batching, vLLM acceleration.
- Kernel & GPU Engineering: Custom CUDA C++ kernels, AMD ROCm/HIP pipelines, shared memory tiling,
cp.asyncstaging. - Post-Training & Alignment: Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), RLHF reward engineering.
- Empirical Interpretability & Probing: Internal hidden-state linear probing, knowledge distillation, activation-normalization interaction dynamics.
Languages : Python, C++, C, CUDA C++, SQL, Bash
ML & Systems : PyTorch, CUDA, AMD ROCm / HIP, Hugging Face (TRL, Transformers, PEFT), vLLM
Agentic AI : LangGraph, LangChain, Tool Orchestration
Backend & DB : FastAPI, Node.js, PostgreSQL, Redis, MongoDB
Infrastructure: Docker, Linux / Unix, Git, GitHub Actions, CI/CD
- Nova-MoE: Sparse Mixture-of-Experts architecture with top-$k$ load-balanced routing and auxiliary entropy loss.
- FastTransformer: GPU-optimized Transformer implementation featuring fused attention kernels, FlashAttention/SDPA, and
torch.compileprofiling. - Qwen2.5-3B GRPO Alignment: LoRA-finetuned reasoning model trained with Group Relative Policy Optimization on GSM8K with verifiable XML reasoning traces.
- Cross-VM HPC Intrusion Detection: Empirical study and benchmark dataset analyzing Hardware Performance Counter (HPC) domain shift under cross-virtual-machine deployments.


