Evaluation & Oversight
How AI behavior and AI-assisted work are tested, verified, escalated, and audited.
Developing Research Direction
I'm exploring how increasingly capable AI systems should be evaluated, reviewed, and governed as direct human oversight becomes harder to scale. My current focus is the design and comparison of human and model review structures: when they improve reliability, where they fail, and how empirical evidence can inform deployment and governance decisions.
This work sits at the intersection of AI safety, evaluation, human-AI systems, and governance.
How AI behavior and AI-assisted work are tested, verified, escalated, and audited.
When parallel review, recursive critique, human escalation, or hybrid human-model systems improve reliability, and when they do not.
How empirical evaluations and system behavior should inform deployment thresholds, organizational controls, and governance decisions.
When do different human and model review structures meaningfully improve reliability rather than add redundant oversight?
How should human oversight change as AI systems become more autonomous, capable, or high-throughput?
What evidence should organizations use to decide when AI systems are reliable enough for consequential use?
This research direction is developing. I'm currently reviewing related work, narrowing the research questions, and evaluating potential empirical approaches. The scope may evolve as I test ideas and learn from existing AI safety and evaluation research.