AI/ML Engineer LLM & vision model optimization on NVIDIA GPUs (PyTorch/CUDA)
I build and optimize AI models for real deployment. My focus is making large language and vision models fast and efficient on NVIDIA GPUs low-precision quantization (NVFP4/FP8), CUDA/Triton kernels, and inference serving with PyTorch, transformers, and vLLM. I also wrap research models into usable products (FastAPI backends with React/Vue front-ends) and build developer tooling and automation in Python and Go.
Ph.D. researcher in CS/AI (SKKU), previously at Samsung Research (AI Service Lab). ~12 years programming, ~8 years in AI/ML, with recent work submitted to NeurIPS 2026. Public work: github.com/Dev-Jahn (vit-nvfp4, repvis, developer tooling).
Happy to start with a small paid diagnostic or scoped sample so you can judge the work before committing.