Geometric Evaluation Theory concept
| Definition | An evaluator-relative theory of preference, value, and choice. An evaluation object is a six-tuple of states with a belief, actions, an evaluator mapping actions and states into a consequence space, a metric of evaluation with an ideal point, a resolution budget, and an admissible set. The distinctions the evaluator can make are derived from the object, and preference and choice are derived from the distinctions. The world supplies the states, the actions, and the map; the evaluator owns the metric, the ideal, the budget, and the admissible set. Written GET, always spelled out, and distinct from Geometric Decision Theory. Also evaluation object. |
|---|---|
| Example | Two actions with consequences (1, 0) and (0.9, 2), an ideal at the origin, and an evaluator that resolves only the first coordinate rank the second action first at distance 0.9 against 1, while resolving both coordinates ranks it last at distance 2.19, so one evaluator with one metric and one ideal reverses its preference when its budget changes. |
| Book | Data Mining as Observation, draft 0.2, commit f3914f0; entry id geometric-evaluation-theory, kind concept. |
| Status | no ledger row names this entry. Corrections: none recorded. |
| Defining equation | none |
| Assumptions and scope |
|
| Prior art | The ideal-point models of Coombs and of spatial voting theory (weighted distance from an ideal) are its unbudgeted single-evaluator case. Luce's semiorders supply the form of finite-resolution preference, with the threshold here fixed by the budget. Rational inattention and limited attention are the nearest theories with a budget, and theirs acts on information about states while this one acts on the resolution of consequences. Dawid and Lauritzen's decision geometry is the identity-evaluator, unbudgeted case. Sen's menu-dependence argument supplies the admissibility theorem. Observation Theory supplies the pullback, the regret, and the budget cliff, and an evaluator is an observer whose output is scored from an ideal point. |
| Evidence | none |
| Reviewed | semantic review 2026-09-07; generated 2026-09-10 from records at the commits on the provenance page. |
Equation
none
Conditions
- The evaluation object is E = (X, A, C, G, B, K),
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:108-130. The world supplies (X, A, C) and the evaluator supplies (G, B, K),geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:131-140. Without that split every weak order on a finite set is trivially an evaluation object, so the split is where the theory’s content lives. - Unbudgeted distinctions are an equivalence relation whose local
tangents are the kernel of the evaluator’s Jacobian. A rank budget
coarsens it into a coarser equivalence. A length budget gives a
tolerance, reflexive and symmetric, that is transitive only when no
three distances form a chain of steps within the budget whose ends lie
outside it,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:174-208. - Preference is derived, not assumed. At zero threshold it is a weak
order represented by minus the distance to the ideal. At a positive
threshold it is a Luce semiorder whose threshold is the budget itself,
so intransitive indifference is the signature of finite resolution and
its size is a property of the evaluation object,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:209-241. Choice on a finite admissible menu is nonempty, and satisficing is the length budget with the aspiration as ideal,geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:242-268. - For a fixed world, a preference is representable exactly when a
semidefinite feasibility problem has a solution. Every representable
preference obeys the hull law, that no action is strictly worse than
every action whose consequences surround it, which holds for every
convex metric of evaluation. Affinely independent consequences represent
every order on at most m+1 actions, and consequences on a line represent
exactly the single-peaked orders,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:269-291. The hull law is machine checked ashull_lawingeometric-evaluation-theory/lean/GET/HullLaw.lean:1-43against Mathlib, with the semidefinite characterization and the line case not checked. A closed-form combinatorial characterization is open. - Uniqueness. From the order on an open connected set of consequences,
the metric is identified up to a positive scale and the ideal up to the
metric’s null directions,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:292-336. This is what the learned-geometry protocol can recover from choices and no more. - Special cases with their conditions,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:337-374. Expected utility, mean-variance evaluation, additive multi-attribute value and the ideal-point models, the geometric decision cost of the TCSS paper, satisficing, the Dawid-Lauritzen decision geometry, and Nash equilibrium through the behavioral game. Loss aversion and lexicographic priority as a metric are not in the theory as written. - Four theorems a scalar utility on the action set does not provide. A
change of rank budget reverses preference between two fixed actions with
the evaluator unchanged,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:375-396. Two evaluators share a standard of correctness exactly when their induced distances are ordinally equivalent, and agreement on every pair of one menu does not transfer to another,geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:397-419. A representation shared by several evaluators at rank k is optimal for all of them exactly when their geometries share a top-k eigenspace, and otherwise the least weighted regret is attained by the top-k eigenspace of the weighted sum and is strictly positive,geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:420-447. Admissibility that depends on the menu produces choice violating the weak axiom of revealed preference, so an obligation of that kind is not a preference at any strength,geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:448-481. - Five predictions are stated in a form a measurement can fail. The
threshold tracks the budget, rank changes reverse preference, shared
representations pay the regret bound, the hull law holds in a population
with a known consequence map, and menu-dependent obligations produce the
weak-axiom pattern,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:482-502. The campaign that grades them isgeometric-evaluation-theory/CAMPAIGN.md:1-40, gates G0 to G8. Done at commit 91095c9: G0, the frozen foundations; G1, the Lean core; G2, the hull law, which passed; G3, the threshold tracks the budget on a language-model scorer, which passed; G4 with its follow-up G4b, the shared-code regret, both indeterminate, together establishing what the prediction can and cannot test; and G5, identification from choices in its synthetic stage, which passed in all twelve cells on a fresh seed, recovering the metric and the ideal at 64 points above the dimension and recovering nothing at or below the design bound, off a subspace battery’s span, or in a singular metric’s kernel,geometric-evaluation-theory/CAMPAIGN.md:108-125at b54d55f. - The hull law has been measured once, under a registration sealed
before any rating was read, on the 1972 American National Election
Study, where respondents placed Nixon, McGovern, Wallace, and the two
parties on seven-point issue scales and rated them on feeling
thermometers,
geometric-evaluation-theory/CAMPAIGN.md:42-53at dc14015. On the registered pair of scales, liberal-conservative and guaranteed jobs, 489 respondents had a menu in which the law could bind, 246 of them, 50.3 percent, violated it at least once, against 66.4 percent with a standard deviation of 1.6 when each respondent’s ratings are shuffled among that respondent’s own objects, and no shuffle in 200 reached the observed rate,geometric-evaluation-theory/experiments/G2/results.json:1-87. Five other pairs of the four scales gave the same ordering, observed rates of 54 to 63 percent against shuffled rates of 67 to 74 percent, the two tax-rate pairs at margins of 0.108 and 0.109 against the sealed bar of 0.10. The verdict is a pass by the sealed bars, and the number that matters more than the verdict is that half of the testable respondents violate the law at least once, so on this world the hull law holds as a population tendency, 10 to 16 points below the geometric base rate of violation, and fails as the deterministic statement of the representation theorem for about half the respondents. Placements are integers on seven-point scales, ratings are integers on a thermometer capped at 97, ties are counted as not strictly better, and no error model was registered, so no part of the half is attributed to rounding. - The shared-representation regret prediction was measured once, under
a registration sealed after one recorded revision and before any loss
under any code was read, on Llama-3.2-3B, with the three query heads of
each grouped-query group as the evaluators of one key cache and shared
codes of rank 2, 4 and 8 in 128 dimensions,
geometric-evaluation-theory/CAMPAIGN.md:73-81at 8eda439. The verdict by the sealed rule is indeterminate at every rank,geometric-evaluation-theory/experiments/G4/grade.json:1-239: at those ranks the measured attention loss is 20 to 330 times the second-order prediction, median 75, with a mean KL divergence of about 4 nats per query, so the regret formula and the deficiency bound were not tested. The registration’s second-order bar was also defective as written, comparing the probe’s one-key prediction to an all-keys measurement, and that is recorded as a registration defect. The ordering claim, measured on the actual losses, held in 32 of 32 non-vacuous cell-ranks: the leading eigenspace of the summed read operators beat all 32 random codes and the key-covariance code for the group’s total loss. The own-code claim failed: a head’s own leading eigenspace was the best code for that head in 1, 3 and 6 of 16 cells at ranks 2, 4 and 8, and the compromise was often better for a head than the head’s own code. Outside the quadratic regime the theorem orders the shared codes and does not identify the private optimum. The probe that preceded the run, at the drafted ladder 8 to 64, had found the operators of effective rank 2.6 to 8.3 and the three heads largely coincident,geometric-evaluation-theory/CAMPAIGN.md:71-71at 8eda439. - A second sealed run, G4b, measured the object the formula predicts:
one key at each probed operating point perturbed by a small isotropic
error in the discarded subspace, with an antithetic pair and common
random numbers, on the same operators,
geometric-evaluation-theory/CAMPAIGN.md:89-98at 8eda439. The second-order regime is reached, the ratio of measured to predicted loss is flat in the perturbation size up to a tenth of a whitened unit and its median over cells is 1.02 to 1.05, so the recovered operators predict the loss without bias; but 64 draws per cell leave each cell’s ratio offset by a factor with standard deviation 0.16 to 0.24, the registered per-triple bar of 80 percent within 25 percent is missed at 64 to 70 percent, and the sealed verdict is indeterminate,geometric-evaluation-theory/experiments/G4b/grade.json:1-814. An exploratory seed check with fresh draws re-centered the five most extreme cells on one, so the offsets are draw noise. The registration’s Monte Carlo error estimate and its diagnosis rule were both wrong and are recorded as registration errors. What the two gates establish together: inside the quadratic regime the regret formula, the bound and the compromise’s optimality follow from the recovered operators by the theorem, so the prediction’s only empirical content there is whether the loss is quadratic with the recovered operator, which holds at the median; outside it, the ordering claim survives and the own-code claim does not. - The threshold prediction was measured once, under a registration
sealed before any pair of the run seed was scored, on a language model,
Qwen2.5-7B-Instruct, used as a scorer that reports how far a value is
from a target, with the induced order comparing reported distances and
equal reports counted as indifference,
geometric-evaluation-theory/CAMPAIGN.md:56-64at 91095c9. Three budgets were set independently, the decimals at which values are rendered, the number of tokens the report may use, and the weight precision. On fresh pairs the threshold, the gap at which 90 percent of pairs are ordered correctly, was 0.869, 0.0868, 0.00856 and 0.00087 at 0 to 3 decimals against rendering steps of 1, 0.1, 0.01 and 0.001, and 0.872, 0.872, 0.0873, 0.00877 and 0.00087 at 2 to 6 report tokens against the resolutions 1, 1, 0.1, 0.01 and 0.001 those tokens afford, every graded threshold at 0.87 to 0.88 of its step, which is where the 90 percent level of a rounding scorer falls,geometric-evaluation-theory/experiments/G3/grade.json:1-120. Pairs separated by 20 were ordered alike in every cell at accuracy at least 0.985. Reducing the weights to 8 or 4 bits left every threshold where bfloat16 put it, so on this evaluator the two output budgets are the resolution budget of the semiorder and the weight precision at 8 or 4 bits is not one. Sealed verdict PASS on all four bars at the registered tolerance 1.5. Two instruments tried before the scorer are kept in the record, a pairwise chooser that was vacuous with options on both sides of the target and had a floor of 7.5 from position bias with options on one side. Not established: anything about a human evaluator, which is gate G7, or a weight ladder coarse enough to move the threshold. - The three branches under the theory are descriptive choice,
normative judgment, and multi-agent interaction, each with what is
measured and what is posited stated,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:503-526. The correspondence of the ethics stack’s deontic gate to the admissible set and of the invariance principle to factoring through the quotient is posited.
Conditions are curated in entries.toml rather than read
from a record.
Ledger
none
First stated
The foundational paper, draft 0.2,
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:1-100,
in the repository ahb-sjsu/geometric-evaluation-theory, with the
campaign, prior-art record, and claim ledger beside it.
Measurements
none
Failures and corrections
none
Invariance envelope
Proved invariant.
- GET:change-of-basis, invertible change of basis of the consequence
space, z’ = A z, with the metric transported as G’ = A^-T G A^-1. Claim:
The evaluative distance (y - t)^T G (y - t) is unchanged.
geometric-evaluation-theory/paper/geometric-evaluation-theory.tex:108-130.
Survived.
- GET:battery-below-design-bound, restriction of the battery of
consequences to fewer than m + 1 points, or to a proper affine subspace,
in the identification gate. Claim: The metric’s action off the affine
span of the battery and the ideal’s component in the metric’s kernel are
not revealed by choices, while the in-span part and the range component
are.
geometric-evaluation-theory/CAMPAIGN.md:108-125. - GET:output-resolution-coarsening, coarsening of the resolution at
which a computational evaluator receives and reports consequences:
values rendered at fewer decimals, or the report limited to fewer
generated tokens. Claim: The indifference threshold of the induced
semiorder equals the resolution budget within a registered factor 1.5
wherever the budget exceeds the evaluator’s floor, is monotone along
each ladder, and the ordering of pairs separated by more than twice the
largest threshold is unchanged.
geometric-evaluation-theory/CAMPAIGN.md:56-64. Witness: on fresh pairs every graded threshold sits at 0.87 to 0.88 of its step (rendering 0.869, 0.0868, 0.00856, 0.00087 against 1, 0.1, 0.01, 0.001; report 0.872, 0.872, 0.0873, 0.00877, 0.00087 at 2 to 6 tokens against 1, 1, 0.1, 0.01, 0.001), the constant a rounding scorer imposes; accuracy at the one well-separated gap, 20, at least 0.985 in every cell. Absorbed by: . - GET:weight-precision-reduction, reduction of the evaluator’s own
weight precision at fixed rendering and report length, bfloat16 to 8-bit
to 4-bit NF4. Claim: The threshold is non-decreasing as weight bits are
removed, within the factor 1.5, and well-separated pairs stay ordered.
geometric-evaluation-theory/CAMPAIGN.md:56-64. Witness: held as an equality: 8-bit and 4-bit weights left every threshold where bfloat16 put it (4-bit at 3 decimals 0.00093 against 0.00087, accuracy at gap 20 0.985 against 1.00), so on this scorer and task weight precision at 8 or 4 bits is not a resolution budget; a ladder coarse enough to move the threshold was not registered. Absorbed by: . - OD:budget, change of the observation budget B at fixed consumer and
output metric. Claim: A single-direction probe along coordinate e_i at
size rho is reported identifiable at budget B exactly when P_ii > B^2
/ rho^2, and a ladder of sizes brackets P_ii between B^2 / r_hi^2 and
B^2 / r_lo^2 (article Theorem 2(b), Corollary 3).
experiments/DISCOVERY-TRACK.md, D1 record; experiments/OD/D1/grade.json axis. Witness: exact in every evaluator of all fifteen World A cells (three worlds, five budgets, 20 evaluators each) on the run seed, as in every pilot. Absorbed by: . - OD:observer-family, change of observer within a declared family:
read operators of declared spectrum, coordinate subsets, coarsenings,
and learned consumers’ recovered read operators. Claim: In a positive
definite world the analytic centre returns the pencil’s end P* = s* P +
(1 - s) c I, s = c / (c - lambda_min), so the estimate is
nearer P* than P (article Corollary 4).
experiments/OD/D1/grade.json, B4; experiments/OD/D1/pencil_probe.json. Witness: nearer P* in 20 of 20 evaluators at B = 1 and at B = 1.5 on the run seed (Frobenius 0.223 against P, 0.181 against P*; 0.117 against 0.102), 20 of 20 in the third pilot, 6 of 6 cells in the pre-seal probe; in the kernel worlds the pencil’s end is P itself and the two errors coincide. Absorbed by: . - OD:observer-family, change of observer within a declared family:
read operators of declared spectrum, coordinate subsets, coarsenings,
and learned consumers’ recovered read operators. Claim: Second version,
the rehabilitation of D1: mixed probes at one radius, where the sphere
crosses the ellipsoid substantially, recover the read operator up to the
pencil s P + (1 - s)(B^2 / rho^2) I, below-threshold eigenvalues
included, with Frobenius, above-threshold and below-threshold medians at
most REC = 0.39 of chance at 960 queries, REC fixed before the pilot as
1.5 times the pilot’s largest graded ratio.
experiments/DISCOVERY-TRACK.md, D1v2 record; experiments/OD/D1v2/grade.json. Witness: all six well-crossed cells on a second fresh seed: Frobenius 0.045 to 0.212 of chance, above-threshold 0.049 to 0.303, below-threshold 0.022 to 0.146; single-direction verdicts and brackets exact in all fifteen World A cells; the pencil’s end nearer than the truth in both positive definite cells; the cell that decided D1 at 0.303 again, 0.087 under the bar. Absorbed by: . Revision: none; the 0.30 of chance for the above-threshold eigenvalues at 960 queries in the positive definite world is a measured property of the analytic-centre estimator, stable across seeds.
Boundary measured.
- GET:one-key-isotropic-perturbation, one key at a probed operating
point perturbed by eps (I - Q) u with u standard normal in whitened
coordinates, antithetic pair, common random numbers across codes. Claim:
The measured loss equals eps^2 tr(Pt (I - Q)) within 25 percent per
triple.
geometric-evaluation-theory/CAMPAIGN.md:89-98. Boundary: the ratio of measured to predicted loss is flat in eps up to 0.1 whitened units (median over cells 1.02 to 1.05) and rises by 4 percent at 0.2; the per-triple bar is missed at every eps (64 to 70 percent within tolerance against 80) by a per-cell scatter of standard deviation 0.16 to 0.24. Witness: a seed check with fresh draws re-centred the five most extreme cells on one, so the scatter is the fluctuation of 64 common draws through operators of low local rank; the registration’s Monte Carlo error estimate used the averaged operator and was wrong, recorded as a registration error. Absorbed by: declaration. Revision: none proposed: inside the second-order regime the regret formula and bound follow from the recovered operators by the theorem, so the only empirical content there is the quadratic model itself, which holds at the median. - GET:population-transfer, transfer of the hull law from a
representable evaluator to a human population with a known consequence
map (self-placed candidates), scored against a within-respondent shuffle
null. Claim: No action is ranked strictly below every action whose
consequences surround it (the hull law), deterministically.
geometric-evaluation-theory/CAMPAIGN.md:43-54. Boundary: the law holds as a population tendency, 50 percent of 489 testable respondents violate it at least once against 66 percent under the shuffle null (permutation p below 0.005, five replication pairs agree), and fails as a deterministic statement for about half the respondents. Witness: the population itself; no error model for integer placements and integer ratings was registered, so no part of the half is attributed to rounding. Absorbed by: declaration. Revision: none yet: an error model would have to be registered before any of the half is attributed. - OD:observer-family, change of observer within a declared family:
read operators of declared spectrum, coordinate subsets, coarsenings,
and learned consumers’ recovered read operators. Claim: Mixed probes at
one radius, where the sphere crosses the ellipsoid substantially,
recover the read operator up to the pencil s P + (1 - s)(B^2 / rho^2) I,
below-threshold eigenvalues included, with Frobenius, above-threshold
and below-threshold medians at most 0.30 of chance at 960 queries.
experiments/DISCOVERY-TRACK.md, D1 record; experiments/OD/D1/grade.json mixed. Boundary: held in five of six well-crossed cells (Frobenius 0.045 to 0.111 of chance, above 0.055 to 0.111, below 0.037 to 0.172); in the positive definite world at B = 1 the Frobenius ratio 0.198 and the below ratio 0.068 held and the above ratio was 0.303 against 0.30. Witness: the tolerance was fixed as the pilot’s own maximum in that same cell (0.299) rounded up to two decimals, so it carried no margin for a fresh seed; the estimator, the crossing and the recovery are unchanged between pilot and run. Absorbed by: declaration. Revision: none for the claim; for later gates, fix a tolerance as a declared multiple of the pilot’s maximum.
Failed, with witness.
- GET:all-keys-rank-truncation, a shared rank-k orthogonal projector
in whitened coordinates applied to every key of a cell at once. Claim:
Each head’s loss under a shared code is the second-order prediction
tr(Pt (I - Q)) from its recovered read operator, and the regret formula
and bound follow.
geometric-evaluation-theory/CAMPAIGN.md:73-81. Witness: measured attention loss 20 to 330 times the prediction (mean KL about 4 nats per query) at ranks 2 to 8; reduced to two causes, the registered prediction was a one-key operator while the measurement coded all 1,024 keys at once, and the codes sat far outside the second-order regime; the ordering claim survived on measured losses in 32 of 32 non-vacuous cell-ranks while the own-code claim failed in 10, 13 and 15 of 16 cells; ledger row GET-9w, geometric-evaluation-theory/claims/LEDGER.md:22-22. Absorbed by: measurement. Revision: the second registration PREREG-G4B.md (blob d6d1c44fcd8454f749bd204e914b02059b8e658f), the object the formula predicts measured directly; ledger row GET-9r, geometric-evaluation-theory/claims/LEDGER.md:23-23.
Machine checked
none
Used in
none
Related
geometric decision cost; observer; consumer; output metric; budget; read operator; quotient; nuisance; budget cliff; Mahalanobis distance.
See also
none
Status
Generated 2026-09-10 by encyclopedia/generate.py; book
at observation-data-mining f3914f0; the commit of every record is listed
in the encyclopedia’s provenance.