Evaluating Agentic AI Systems
Work on approaches for assessing multi-agent AI applications, including evaluation workflows, centralized evidence and governance-oriented reporting.
I work on the questions behind AI systems — how agents collaborate, how applications are evaluated, how retrieval adapts, and how emerging ideas become reusable engineering.
Selected work across agentic AI, developer infrastructure and applied AI — described at the level I can responsibly share.
Work on approaches for assessing multi-agent AI applications, including evaluation workflows, centralized evidence and governance-oriented reporting.
Work on a unified SDK abstraction that gives developers a consistent way to create, host, orchestrate and consume tools and agents.
A personal exploration into using document characteristics to inform retrieval configuration and reduce manual RAG tuning.
It should solve the problem in front of you — and leave behind a better way for someone else to solve the next one.