Skip to content
View sarimahsan's full-sized avatar
💻
Focusing
💻
Focusing

Block or report sarimahsan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sarimahsan/README.md

Portfolio  Google Scholar  HuggingFace  LinkedIn  ORCID  ResearchGate  Medium


AI Research Engineer and Computer Engineering undergraduate at SSUET, working as a Software Engineer at a UK-based firm. My research focuses on LLM Optimization & Efficient Inference, Distributed ML Systems, Post-Training Alignment, and Mechanistic Interpretability.

I study and implement intelligent systems from first principles — spanning custom CUDA/ROCm kernels, distributed training workflows, post-training alignment (GRPO, DPO), and empirical hidden-state probing.


2,400+ HF Model Downloads  •  4+ Research Papers  •  5+ Systems & ML Projects  •  AMD ROCm Certified



🔬 Research & Engineering Focus

  • LLM Systems & Inference: KV-cache optimization, Speculative Decoding, Continuous Batching, vLLM acceleration.
  • Kernel & GPU Engineering: Custom CUDA C++ kernels, AMD ROCm/HIP pipelines, shared memory tiling, cp.async staging.
  • Post-Training & Alignment: Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), RLHF reward engineering.
  • Empirical Interpretability & Probing: Internal hidden-state linear probing, knowledge distillation, activation-normalization interaction dynamics.


🛠️ Technical Stack

Languages     : Python, C++, C, CUDA C++, SQL, Bash
ML & Systems  : PyTorch, CUDA, AMD ROCm / HIP, Hugging Face (TRL, Transformers, PEFT), vLLM
Agentic AI    : LangGraph, LangChain, Tool Orchestration
Backend & DB  : FastAPI, Node.js, PostgreSQL, Redis, MongoDB
Infrastructure: Docker, Linux / Unix, Git, GitHub Actions, CI/CD


📚 Selected Open-Source & Research Highlights

  • Nova-MoE: Sparse Mixture-of-Experts architecture with top-$k$ load-balanced routing and auxiliary entropy loss.
  • FastTransformer: GPU-optimized Transformer implementation featuring fused attention kernels, FlashAttention/SDPA, and torch.compile profiling.
  • Qwen2.5-3B GRPO Alignment: LoRA-finetuned reasoning model trained with Group Relative Policy Optimization on GSM8K with verifiable XML reasoning traces.
  • Cross-VM HPC Intrusion Detection: Empirical study and benchmark dataset analyzing Hardware Performance Counter (HPC) domain shift under cross-virtual-machine deployments.


📊 GitHub Activity

  


© 2026 Syed Sarim Ahsan • syedsarimahsan.netlify.app

Pinned Loading

  1. arxiv-lens arxiv-lens Public

    AI-powered research tool that summarizes ArXiv papers, builds citation knowledge graphs, and lets you chat with any paper using RAG — built with FastAPI, HuggingFace, and React.

    Python 1

  2. gpt2-inference-engine-scratch gpt2-inference-engine-scratch Public

    A pure NumPy implementation of GPT-2 inference. No PyTorch in the forward pass — every operation (attention, layer norm, GELU, sampling) is written by hand using only numpy. Weights are loaded once…

    Jupyter Notebook

  3. Mini-Qwen_from_scratch Mini-Qwen_from_scratch Public

    A compact, educational implementation of a Qwen-inspired autoregressive language model built with PyTorch. The project includes transformer components, a custom AdamW optimizer, a Hugging Face data…

    Python

  4. interplens interplens Public

    InterpLens is an open-source mechanistic interpretability toolkit that accelerates neural network reverse-engineering with interactive circuit debugging and real-time telemetry. It bridges research…

    Python 2

  5. nova-moe nova-moe Public

    Nova-MoE is a clean, modular, from-scratch PyTorch implementation of a Mixture-of-Experts (MoE) Causal Language Model designed for controlled sparsity ablations against an active-parameter-matched …

    Python

  6. cuda-ml-from-scratch cuda-ml-from-scratch Public

    Building deep learning and ML architectures from bare-metal CUDA C++ to Python. Zero frameworks, raw GPU kernels, custom autograd, and hardware-level optimizations.

    Python