Skip to content
View yuki-uix's full-sized avatar
🤗
keep learning
🤗
keep learning

Block or report yuki-uix

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yuki-uix/README.md

Kunyu Xu (Yuki)

A terminal session. yuki whoami prints an ASCII fish beside: hi, i'm yuki. I build AI agents, eval systems, and developer tools. yuki background prints: I studied architecture before I wrote code. yuki what-keeps-me-curious prints: the jump from it works to it's useful, building it, testing it, finding out.

Portfolio LinkedIn Email

agent evaluation · grounded generation · local-first safety · developer experience

I build AI agents, evaluation systems, and developer tools that turn LLM demos into measurable delivery decisions. At Thoughtworks, I work across TypeScript, React, Node.js, Python, and Java/Spring Boot—from prototypes and eval harnesses to production workflows.

01 / systems, not demos

Insurance enquiry triage and reply drafting with a frozen golden set, human review queue, and 2×2 evaluation matrix.

The useful result: every tested configuration missed the safety bars, producing a defensible no-ship decision.

Decision: no ship 2 by 2 evaluation matrix

A source-learning agent with constructive grounding, complete exit-path safety gates, and measured agent-loop economics.

Validated: agent-loop economics and exit-path safety across 638 tests and 22 merged PRs, with the findings turned into reusable engineering guidance.

638 tests 22 merged pull requests

Grounded evidence becomes schema-constrained, interactive knowledge cards—not arbitrary model-generated UI code.

Measured: dense retrieval reached 63.7% Recall@10 versus 47.6% for BM25; fusion did not beat the best single retriever.

Recall at 10: 63.7 percent 60 evaluation questions

A measurement lab for the uncomfortable fact that fewer tokens can still mean a larger bill when cache economics change.

First result: one compaction paid back after 18–19 turns, not the 2–4 turns predicted before measurement.

Compaction payback: 18 to 19 turns Predictions locked before runs

agent-eval-harness  •  pr-review-agent  •  beforeshare

02 / measured impact

≈90% less time
multi-market rollout setup
1,018 KB → 92 KB
evaluation worker bundle
31 HTTP tests
integration coverage
Top 3 · APAC
Thoughtworks AI/works challenge
  • Built a project-wide AI code review platform that runs on every PR across multiple TypeScript repositories; it caught a runtime crash missed by ESLint and AI-generated tests.
  • Built and piloted an agent-configuration evaluation engine with two delivery teams, including a human-in-the-loop improvement cycle.
  • Encoded multi-market rollout knowledge into a reusable coding-agent workflow, reducing setup from roughly one week to half a day.

03 / latest writing

Fresh notes on AI delivery, design engineering, and the parts of software that only become visible after the demo works.

Read the full archive →

04 / contribution graph

Animated snake eating Yuki's GitHub contributions

05 / toolbox

TypeScript, React, Next.js, Node.js, Python, Java, Spring, PostgreSQL, Google Cloud, Docker, Git, and GitHub Actions

AI engineering: Agent evaluation · RAG / embeddings · MCP · LLM APIs · Prompt / Skill Engineering
Product engineering: TypeScript · React · Next.js · Node.js · Hono · Python · Java · Spring Boot
Delivery: Azure DevOps · CI/CD · Google Cloud · Playwright · Honeycomb

Why architecture still matters to my engineering
I studied architecture before I wrote code. It trained me to see software as a system people move through—not just a set of screens—and still shapes how I design agent interactions, failure paths, and developer tools.

Yuki's pixel bonsai, grown from GitHub activity and updated daily
A little tree, grown from my GitHub history. Tended daily.

Portfolio · Juejin · LinkedIn · Dev.to

Profile views

Pinned Loading

  1. pace-triage-agent pace-triage-agent Public

    Evaluated insurance triage agent with a frozen golden set, human review queue, and evidence-backed no-ship decision

    Python

  2. rag-generative-ui-explorer rag-generative-ui-explorer Public

    RAG-powered generative UI explorer for grounded, interactive knowledge cards

    TypeScript 2 2

  3. RepoCoach RepoCoach Public

    Source-learning agent for measuring agent-loop cost, grounded evidence, safety gates, and cross-process state

    TypeScript 2

  4. agent-cost-lab agent-cost-lab Public

    Measuring what coding-agent cost optimisations actually cost — tokens are not money

    Python 2 2

  5. pr-review-agent pr-review-agent Public

    Read-only, evidence-based PR review agent built on mini-swe-agent with replayable optimization cases

    Python

  6. agent-eval-harness agent-eval-harness Public

    Zero-dependency Python harness for paired agent regression tests, exact McNemar analysis, and sample-size planning

    Python