Developing Research Direction

AI Evaluation, Oversight, & Governance

I'm exploring how increasingly capable AI systems should be evaluated, reviewed, and governed as direct human oversight becomes harder to scale. My current focus is the design and comparison of human and model review structures: when they improve reliability, where they fail, and how empirical evidence can inform deployment and governance decisions.

This work sits at the intersection of AI safety, evaluation, human-AI systems, and governance.

Current Areas of Inquiry

Where the research is focused now

A

Evaluation & Oversight

How AI behavior and AI-assisted work are tested, verified, escalated, and audited.

B

Review Architectures

When parallel review, recursive critique, human escalation, or hybrid human-model systems improve reliability, and when they do not.

C

Governance from Evidence

How empirical evaluations and system behavior should inform deployment thresholds, organizational controls, and governance decisions.

Questions I'm Exploring

Open questions guiding the work

  1. Q01

    When do different human and model review structures meaningfully improve reliability rather than add redundant oversight?

  2. Q02

    How should human oversight change as AI systems become more autonomous, capable, or high-throughput?

  3. Q03

    What evidence should organizations use to decide when AI systems are reliable enough for consequential use?

Current Status

This research direction is developing. I'm currently reviewing related work, narrowing the research questions, and evaluating potential empirical approaches. The scope may evolve as I test ideas and learn from existing AI safety and evaluation research.