Research note 10 min read

Metrics that still discriminate

What funders and partners can still measure when agents saturate hiring, review, and drafting — and polish stops signaling depth.

  • Institutions & evaluation
  • Human–AI interaction
  • Learning & cognition

Institutions are adopting agents faster than they are revising how they know what they know. Hiring panels read agent-polished statements. Reviewers encounter agent-shaped manuscripts. Operations teams inherit agent-generated summaries they cannot fully audit. Funders receive proposals that share structure, tone, and risk language — and must decide whom to back anyway.

This note is about evaluation under saturation: what signals still discriminate when expressive form converges, verification capacity shrinks, and compute-intensive checking is budget-limited. We write it in September because the threads we tracked earlier in the year — semantic homogenization, delegation and lock-in, and the power wall — are no longer separate concerns. They meet in the committee room.

The failure mode is false legibility

Most institutional metrics were built for a world where voice, process, and outcome were loosely coupled. A well-written grant did not guarantee good science, but distinctive reasoning left traces: surprising definitions, contested citations, visible trade-offs, errors that revealed where thinking actually lived.

Semantic homogenization compresses those traces. Materials can differ in content yet feel interchangeable in register — the same hedges, section arcs, and confidence calibration. Reviewers experience false legibility: everything reads competently; less reads specifically.

Meanwhile delegation shifts where competence is exercised. Work that once displayed junior friction — rough drafts, wrong abstractions, corrected dead ends — may never appear. What remains is finish. Institutions that reward polished outputs without measuring verification process systematically overpay for style and underfund depth.

And the power wall caps how much verification institutions can afford. Deep review of long agentic chains — multi-step retrieval, drafting, revision — has an energy and labour bill. Shallow review is not always negligence. Sometimes it is rationing.

The composite risk is an evaluation stack that looks rigorous while discriminating weakly: checklists satisfied, tone appropriate, citations present — and intellectual lineage invisible.

What still discriminates (and what does not)

We do not believe evaluation is hopeless. We believe default metrics are decaying and must be replaced with ones that stress-test lineage, disagreement, and constraint — not surface form.

Weak discriminators — useful for hygiene, poor alone for high-stakes choice:

  • Readability and formatting quality
  • Generic safety and ethics boilerplate
  • Keyword overlap with funding priorities
  • Self-reported tool-use disclosures without corroboration
  • Benchmark scores on tasks disconnected from the proposed work

Stronger discriminators — harder to fake at scale, especially when combined:

  • Pre-registered predictions with dated priors and explicit falsifiers
  • Intermediate artifacts: lab notebooks, failed models, abandoned hypotheses, raw data paths — not only final narratives
  • Adversarial review where critics receive time and mandate to find concrete errors, not stylistic nits
  • Replication or reproduction commitments with resource lines itemized (including compute and verification labour)
  • Process interviews that probe how a claim was reached, not only what was concluded
  • Disagreement maps: where the team expects expert pushback, and what would change their mind
  • Operational metrics with failure visibility — especially in civilian systems where lives are involved (warning latency distributions by community, not demo-city medians)

None of these is novel in isolation. The point is composition. Institutions need bundles that remain informative when agents participate in every stage of production.

Three measurement commitments for partners and funders

We suggest three commitments that travel across domains — research, hiring, disaster risk, public-sector modelling.

1. Measure lineage, not only output

Ask: Where did this claim live before it was polished? Require artifact trails proportionate to stakes. A partnership deck is not a paper; a seismic operations plan is not a blog post. Scale demands accordingly — but the principle holds: polish is not evidence of thought.

2. Budget verification as a first-class line item

Verification is skilled work. It competes for the same junior pipeline that automation erodes. Funders who want trustworthy outputs should fund time to disagree well — including compute for independent re-runs where models are part of the method. Treating review as volunteer labour atop agentic production is how moral crumple zones form: humans nominally responsible for outcomes they lacked capacity to check.

3. Report distributions, not hero numbers

Agent-assisted systems fail unevenly — across languages, regions, experience levels, and edge cases. Institutions should publish distributional performance: who received how much warning time; which cohorts saw employment effects; which communities show semantic flattening in institutional records. Hero metrics are compatible with systemic harm. Funders should reward transparency about tails.

Where this connects to our applied work

In earthquake early warning, the metric that matters is not whether an alert was sent, but how many seconds each population segment had — and whether trust survived false alarms. In energy-constrained reasoning systems, the metric is not peak benchmark score, but cost per verified conclusion at declared depth. In language and learning, the metric is not fluency, but whether distinctive expertise remains detectable across agent-mediated materials.

Different domains. Same institutional question: what still discriminates when production gets cheap and conformity gets rewarded?

Open questions we are pursuing

  • Can we build homogenization-aware review tools that flag stylistic convergence without punishing non-native or non-dominant registers?
  • What minimum artifact bundle predicts reproducible research in agent-assisted workflows — and what is waste?
  • How should funders score proposals that disclose heavy agent use versus those that omit it — without creating perverse incentives to hide tooling?
  • In high-stakes civilian systems, what audit cadence keeps metrics honest after deployment, not only at launch?

We do not treat this note as a framework to bolt onto existing programmes unchanged. It is an argument that evaluation itself is infrastructure — as load-bearing as models, sensors, and grids — and that it is currently underbuilt for the year we are actually in.

If your mandate includes building institutions that can still tell depth from polish, start a conversation. We are looking for partners and funders who want to measure what matters — not only what is easy to score.