I evaluate and stress-test AI models for correctness, safety, and reliability. My work includes reviewing AI-generated responses for accuracy, safety, contextual understanding, and instruction-following quality, designing prompts that surface edge cases and reasoning failures, and systematically stress-testing AI systems to identify failure points.
This includes hands-on experience fine-tuning and evaluating language models, building evaluation pipelines, and analyzing model outputs for subtle correctness and reasoning errors that automated metrics miss.
What I offer:
- AI model evaluation and benchmarking
- Prompt engineering and prompt library development
- Systematic stress-testing and edge-case testing of AI systems
- Reasoning and edge-case analysis for LLM outputs
- Training data curation and quality review
- Clear, well-written technical explanations of findings
I bring genuine technical depth here, not just prompt tweaking, real experience across model training, evaluation, and systematic adversarial testing.