Research

Five threads, one question: what does it take to know whether an AI system is working?

The threads below are not separate projects so much as different angles on the same difficulty. Measurement is the through-line — what gets counted as evidence, who or what produces it, and what that leaves unaccounted for. Each thread accumulates writing, and where the history warrants it, a spine.

This is active work. Descriptions here state what each thread is about, not what has been concluded.


01
Evaluation and measurement
Benchmarks report numbers, but a number is the output of an instrument, and every instrument encodes assumptions about what it is measuring. Ground truth is not found; it is produced — by annotation processes, adjudication procedures, and choices about what counts as correct. This thread treats evaluation claims as measurement claims, and asks what follows when a measure becomes a target.
ground truth · instrumentation · validity · construct · benchmarks
02
Information and meaning
Shannon made communication measurable by deliberately setting meaning aside. Seventy years later, systems trained on next-token prediction produce fluent language, and the question of what they represent is live again. This thread follows the long argument between formal semantics, pragmatics, distributional linguistics, and learned representation — and what each says about the difference between prediction and understanding.
semantics · pragmatics · embeddings · context · grounding
03
Affective computing
Systems increasingly detect, model, and respond to human emotional states. The measurement problem is acute here: emotion categories are contested in psychology itself, self-report is unreliable, and observable signals underdetermine internal states. This thread examines what affective systems are actually measuring, and what follows for the people being measured.
emotion recognition · construct validity · machine behavior · judgment
04
Governance
Standards, taxonomies, and regulatory frameworks translate research findings into operational requirements — and in doing so, fix particular definitions of harm, risk, and adequacy. This thread reads governance instruments as measurement instruments, examining what they make legible and what falls outside their scope.
standards · taxonomies · policy · risk frameworks · accountability
05
Human-AI collaboration
When an AI system completes a task, the work does not end — someone still has to verify, correct, and take responsibility for the result. That effort is real, unevenly distributed, and largely unmeasured. This thread concerns oversight as a first-class object of study: what supervision costs, how it is distributed, and what makes a system supervisable rather than merely capable.
oversight · supervision · teaming · verification · accountability

A note on what is published here. Writing on this site is intended to be readable and useful on its own terms rather than to substitute for peer-reviewed work. Where something is preliminary, it says so.