AI/ML Engineer with over 4 years of cumulative experience (full-time and freelance/contract) across the full ML lifecycle: LLM fine-tuning (SFT, DPO, RLHF, LoRA/QLoRA), inference optimization (vLLM, TensorRT-LLM, ONNX, speculative decoding, quantization), medical imaging and computer vision, Graph-RAG retrieval, and multi-agent orchestration. Independent researcher in runtime-adaptive speculative decoding for CPU-constrained LLM inference, with a first-author arXiv preprint. Track record of measurable production impact — 60% P99 latency reduction, 4x inference throughput gain, 42% hallucination reduction, 35% ASR accuracy improvement. Open to AI/ML Engineer, Deep Learning Engineer, and Data Science roles internationally, with full flexibility for relocation.
No employment history.
No education history.