April 2026 - Present
LLM Engineer and Manager
I am on a fellowship at Handshake AI, running human-in-the-loop evaluation of a leading large language model as it learns how professional design software actually works.
The work has two halves. One is producing expert-level reference deliverables in industry-standard creative tools, 2D graphics, CAD, and game development, so the model has something correct to learn from. The other is writing and evaluating the prompts that test whether it learned anything, across real workflows rather than toy examples.
I also write the rubrics that define what quality means for a professional design task, which turns out to be the hard part. Saying an illustration is good is easy. Saying precisely what would make it wrong, in terms a model and a reviewer can both apply, is not.
Alongside my own output I review other designers’ deliverables, giving the feedback that keeps the reference set consistent. If you’re trying to pin down what quality means in an expert domain, I’d like to hear about it.
- Role
- AI Trainer, 2D Graphics and Game Development Expert
- Program
- Handshake AI Fellowship, contract, remote
- Dates
- April 2026 to present
- Work
- Prompt development, LLM evaluation, reference deliverables, rubric authoring, review
- Tools
- Adobe Creative Suite, Figma, GIMP, Inkscape, Blender, game development tooling