Projects

Practice · Product design & research

How I work

The method behind everything else on this site: watch the work first, prototype in code, and measure whether AI actually helps the person using it.

What I build

Interfaces for AI systems that people actually trust: approval and review flows, evidence and citation UX, and the copy conventions that make agentic tools legible. I spent three and a half years doing this at Ubik Studio before “human-in-the-loop” was an industry phrase, and I do it today at Circleheads — shipping production agents behind approval gates — and in the open at akaOSS, where the HITL Kit packages the patterns as fifteen installable primitives.

Watch the work first

Before designing anything, I watch people do the job the software is supposed to help with. Mixed-methods research — structured interviews, behavioral observation, session replays — synthesized into prioritized UX decisions rather than slide-deck summaries. Every cycle is anchored in three questions: does the person trust what they're seeing, can they trace where it came from, and do they stay in control?

Findings feed interaction specs, flow changes, system prompts, and microcopy in the same shipping rhythm as the product. One window into that loop is the public team test log: real feedback and observation turned into concrete improvements, documented so you can read the arc — not only the conclusions.

Prototype in code, at scale

I don't hand off static mockups. Claude Code and Cursor turn intent into working surfaces — web and desktop — fast enough that the build stays aligned with what research is finding the same week. The demos on this site are that practice in public: Research OS is a deliberately small slice of an agentic research workspace built to stress-test legibility, density, and control; Music Analysis Chat rehearses rich agent output in a domain I know from the inside; the HITL-AI showcase and component sheet are the earlier in-repo iterations that became the Kit.

Measure what matters

95% of enterprise AI initiatives deliver zero measurable return — not because the models are bad, but because we measure the wrong thing. Benchmarks ask whether a model can complete a task alone; deployment asks whether it respects the user's authority, preserves their agency, and makes them better over time. I design and evaluate for the second question. The full argument is my paper, An AI Measurement Problem, and the working version is the eval-kit — an evaluation framework where humans score, not LLMs. Every primitive in the HITL Kit is tied to a claim the paper defends; if it can't be, it doesn't ship.

The demos aren't decoration

Everything on this site rehearses the same discipline that ships in product: Research OS for end-to-end flows, the HITL-AI pages for primitive comparison, the Kit and the Kraa write-ups for the argument and the field notes. If you want the full index, it's all on the projects page.