The Knowledge LedgerThe Knowledge LedgerFollow on WhatsApp
← Back to feed
July 10, 2026·gilharilabs.com

What If Evaluations Had Three Dimensions?

Most evaluations today reduce complex outcomes to a single scalar score or leaderboard number. This compression makes it impossible to distinguish between a genuine, reliable improvement and a fragile result that looks good only on the measured metric.

Nagaraju Gangaraju proposes shifting from one-dimensional scalars to a three-dimensional evaluation space called Vec3(T, R, V), where:

•⁠ ⁠T (Truthness): Did the desired outcome actually occur?

•⁠ ⁠R (Reliability): Would it occur again under similar conditions?

•⁠ ⁠V (Validity): Does the measurement align with what we actually care about?

In this bounded cube, outcomes become points with independent coordinates rather than positions on a line. This makes it possible to visualize trajectories and detect issues like "Goodhart Drift" (where the metric improves while validity declines) or "Pyrrhic Passes" (high truthness but low reliability and validity).

The framework is especially powerful for AI development, where rapid optimization can exploit blind spots in evaluation, leading to specification gaming, reward hacking, or brittle gains that collapse under real-world conditions.

This Ledger Entry expands how readers think about evaluation and optimization in AI by showing that compressing outcomes into a single scalar score obscures important distinctions — and that representing evaluation in a three-dimensional space (Truthness, Reliability, Validity) enables clearer detection of drift, fragility, and misalignment during optimization.

Read the full article

Share this post