agent evaluation ·
grounded generation ·
local-first safety ·
developer experience
I build AI agents, evaluation systems, and developer tools that turn LLM demos into measurable delivery decisions. At Thoughtworks, I work across TypeScript, React, Node.js, Python, and Java/Spring Boot—from prototypes and eval harnesses to production workflows.
agent-eval-harness
•
pr-review-agent
•
beforeshare
| ≈90% less time multi-market rollout setup |
1,018 KB → 92 KB evaluation worker bundle |
31 HTTP tests integration coverage |
Top 3 · APAC Thoughtworks AI/works challenge |
- Built a project-wide AI code review platform that runs on every PR across multiple TypeScript repositories; it caught a runtime crash missed by ESLint and AI-generated tests.
- Built and piloted an agent-configuration evaluation engine with two delivery teams, including a human-in-the-loop improvement cycle.
- Encoded multi-market rollout knowledge into a reusable coding-agent workflow, reducing setup from roughly one week to half a day.
Fresh notes on AI delivery, design engineering, and the parts of software that only become visible after the demo works.
2026-08-27— We Can Build an Agent in Ten Minutes — Why Does Shipping It Still Take Weeks?2026-05-31— Reading baoyu-skills Source Code Through a Design Systems Lens — Three Familiar Patterns2026-05-28— I Thought All 9 Token Styles Were Valid — Until I Changed the IA
AI engineering: Agent evaluation · RAG / embeddings · MCP · LLM APIs · Prompt / Skill Engineering
Product engineering: TypeScript · React · Next.js · Node.js · Hono · Python · Java · Spring Boot
Delivery: Azure DevOps · CI/CD · Google Cloud · Playwright · Honeycomb
Why architecture still matters to my engineering
I studied architecture before I wrote code. It trained me to see software as a system people move through—not just a set of screens—and still shapes how I design agent interactions, failure paths, and developer tools.
A little tree, grown from my GitHub history. Tended daily.



