LLM & Generative AI Engineer | RAG, AI Agents, Fine-Tuning & FastAPI
I am an AI Engineer with 5+ years of experience building production-ready solutions with Large Language Models, Generative AI, and Computer Vision.
I help businesses turn complex requirements into reliable AI products—from LLM fine-tuning and multimodal RAG to agent workflows, scalable APIs, and private deployment. My work covers public-sector, healthcare, e-commerce, and industrial use cases.
My core expertise includes LoRA/QLoRA fine-tuning, RAG and knowledge-graph systems, AI agent and workflow design, vLLM inference optimization, multimodal document understanding, and high-concurrency Python backends with FastAPI, aiohttp, Docker, MongoDB, and SSE streaming.
I have led complete AI delivery pipelines, including million-scale data preparation, model training, inference optimization, secure on-premise deployment, and production integration.
I work remotely from China and communicate fluently in English and Chinese. I am experienced with async collaboration across time zones and focus on clear communication, dependable delivery, and maintainable engineering.
Work Terms
Open to remote, long-term engagements and focused AI engineering projects. Available for hourly or milestone-based work.
Working hours: Monday–Saturday, 09:00–23:00 China Standard Time (GMT+8), with flexible overlap for Europe and North America. Async collaboration is welcome; I typically respond within 12 hours.
Preferred communication: Whatsapp, E-mail, written project specifications, and regular progress updates. I can work independently or integrate with an existing engineering team. I am comfortable with English and Chinese communication.