I build AI products that work outside the demo.
For close to seven years I've worked across software engineering and applied machine learning, most recently on LLM powered features and the evaluation loops behind them.
That means retrieval systems grounded in real client data, the Python and Go services that serve them at production traffic and the benchmarks that tell you honestly whether a change made the model better or just different.
Before AI took over most of my work I was a founding engineer, building a company's first platform from an empty repository through to paying customers.
That shaped how I work now. I like owning a system from the data layer through to deployment, I care about clean architecture and measurable quality and I'd rather ship something narrow that holds up than something broad that impresses once.
A year evaluating frontier models at DataAnnotation gave me an unusually close view of where these systems fail, which turns out to be the most useful thing I know as an engineer building on top of them.
Currently open to senior AI engineering roles.