# Exploratory comparison of ten theorem explanations

Prepared 2026-10-08. One automated review; no expert human adjudication.

The review found **59 retained checklist entries, 15 with something not stated explicitly, one uncertain interpretation, and no demonstrated scope error**. These are descriptions of 75 frozen checklist entries across ten responses, not an accuracy rate. An unstated technical detail is not automatically a mathematical error. No source theorem was proof-checked.

The original answers were left unchanged. The evidence file preserves the complete supplied collection packet, including prompts, answers, settings and completion records, and adds judgments separately. Its serialized original packet was checked for equality with the input.

## Per-theorem summary

R = retained meaning; N = not explicit; E = demonstrated scope error; U = uncertain. Counts refer to whole frozen checklist bullets, some of which contain multiple components.

| Theorem | R | N | E | U | Retained meaning and remaining limits |
|---|---:|---:|---:|---:|---|
| T01 — Annular variation | 6 | 2 | 0 | 0 | Preserves r>2, the norm estimate, point-dependent partitions and constant dependency. Complex-valued inputs and the common-null-set representative convention are not explicit. |
| T02 — Perfect matchings | 4 | 3 | 0 | 0 | Preserves the graph class, multiplicative probability guarantee, certain zero case and worst-case runtime. Rational parameter ranges, nonnegative rational output and the explicit bit-operation model are not fully stated. |
| T03 — Contingency tables | 8 | 1 | 0 | 0 | Preserves arbitrary margins, uniform laws, feasibility and expected versus bounded runtime. Unbiased random bits/bit operations are not explicit. |
| T04 — Average sensitivity | 6 | 1 | 0 | 0 | Preserves real multilinearity, uniform input, sign at zero and 8d√n. The full declared integer range n≥1, 1≤d≤n is not explicit. |
| T05 — Fractal measure dimension | 7 | 1 | 0 | 1 | Preserves map entropy, contraction rate, dimension cap, overlaps and reflections. Strictly positive weights are not explicit; the affine/nondegenerate map class is uncertain in the verbal abstraction. |
| T06 — Ergodic averages | 6 | 2 | 0 | 0 | Preserves mixing, bounded fixed measurements, the product limit, full integer sequence and dependent exceptional sets. Invertibility/bimeasurability and explicit n≥2 are not stated. |
| T07 — Sums and products | 6 | 0 | 0 | 0 | Preserves finite-set quantifiers, distinct positive integers, all nonempty subsets and optional rapid growth. Correctly avoids an infinite-set conclusion. |
| T08 — Symmetric powers | 5 | 1 | 0 | 0 | Preserves b≥6, finite dimension, injection direction, equivariance and Schur positivity. The complex base field is not explicit. |
| T09 — Relative generation | 7 | 0 | 0 | 0 | Preserves irreducibility, separable predual, generation relative to P and trace-2-norm dense G_delta typicality. Explicitly distinguishes topology from probability. |
| T10 — Elasticity uniqueness | 4 | 4 | 0 | 0 | Preserves smoothness, physical positivity, complete exact static data and both coefficient equalities. Boundedness/connectedness, explicit real coefficients, closure-wide regularity/positivity and fixed coordinates are not all stated. |

![Counts of automated checklist judgments by theorem](checklist-summary.png)

The figure counts mixed kinds of checklist content. It does not rank fields or responses by mathematical quality. A row marked N can retain most of a compound criterion.

## Concrete evidence and distinctions

**A retained conclusion: runtime versus expected runtime (T02 and T03).** T02 says the bound holds for “**every execution**” and “not merely on average”; its source states “The worst-case bit running time is polynomial”. T03 instead says “Individual runs can take much longer”, matching the source's “expected running time polynomial” guarantee for exact sampling. This distinction was retained. T02's omission of an explicit bit-operation model is recorded separately within its compound runtime criterion, not treated as confusing worst-case and expected time.

**An unstated assumption, not a demonstrated false extension (T06.01).** The response describes T as “a probability-preserving evolution”. Source context states that T is “invertible, bimeasurable, and measure preserving”. Invertibility and bimeasurability are not verbalized. But the response does not expressly claim the result for noninvertible transformations, and refers to the supplied T. The status is N, not E. The same conservative reasoning applies to T08's unstated complex field.

**A retained quantifier restriction (T06.06).** The response says: “The exceptional set may depend on the chosen measurements; the theorem does not assert a single set of good starting points that works for every bounded function at once.” The source scope note says: “The exceptional null set may depend on the system and the fixed tuple.” This preserves the order of fixing functions and obtaining almost-everywhere convergence.

**An uncertain abstraction (T05.01).** The response begins the construction with “finitely many maps that shrink the real line, possibly also reflecting and translating it”. The source specifies affine maps, `varphi_i(x)=r_i x+t_i`, with `0<|r_i|<1`. In the supplied context, the response can reasonably refer to exactly that class. In isolation, “maps that shrink” could include nonlinear contractions or constant maps. Later references to r_i and log|r_i| constrain the interpretation but do not resolve the verbal class completely. This is U, not a declared scope error. Expert review could reasonably move it to retained or not explicit; no alternative count is presented as equally established.

**Equivalent meaning is accepted (T03.05, T05.07, T10.03).** T03 describes total variation using the largest discrepancy in probabilities of collections of tables; on a finite set this is equivalent to the source's half-sum convention. T05 uses typical local mass scaling rather than the positive-mass-set definition of measure dimension; the source context explicitly states exact dimensionality and equivalence of these descriptions. T10 says “both resistances are positive”, after identifying shear modulus mu and bulk modulus lambda+2mu/3. This preserves the source inequalities mu>0 and 3lambda+2mu>0. T10.03 is nevertheless N because positivity on the entire closure is not explicit—not because the physical reformulation is wrong.

**A representative convention is not an extra hypothesis (T01.07).** The source sets quantities “to zero on the exceptional set”. The answer says variation is finite “at almost every point” but does not repeat the zero assignment. This is recorded as an unstated technical convention, not loss of an assumption on F or G and not an error in the norm estimate.

Source quotations in this report may join source line wraps; mathematical expressions above are readable renderings. The evidence ledger contains verbatim passages with exact source paths and line/character locators. Response quotations are excerpts, not edited replacement answers.

## Adjudication method

The four status definitions are those recorded in the collection packet. Retained allows equivalent language and clear notation. Not explicit means absent without a demonstrated contradiction or enlarged claim. Scope error requires affirmative source/response evidence of changed scope. Uncertain reserves an unresolved reading.

Every frozen bullet has one row, identified T01.01 through the relevant theorem's final entry. The original criterion text is preserved. Columns separate semantic kinds: assumptions/domains, quantifiers, definitions, conclusions, exceptions, computational models and technical conventions; mixed criteria retain mixed labels. Full source evidence, a verbatim response paragraph, retained components, absent components and rationale are available in the ledger. For an absence judgment, the whole response was reviewed; its quoted paragraph is nearby evidence, not proof of absence by itself.

For descriptive counting, a compound bullet is N when at least one clearly missing component remains; retained components are still named in the rationale. U takes precedence when an implicit technical reading cannot reliably be resolved. These operational choices were made during this review, after reading the answers. They were not preregistered, and the frozen checklist was not rewritten or split to change its denominator. A different reasonable treatment of compound bullets could change counts.

No response passage met the threshold for a demonstrated scope error. That finding does not certify mathematical correctness or complete retention, and does not mean errors would be absent under expert review.

## Integrity, timing and study limits

- **Local integrity checks passed.** All three frozen file hashes, ten prompt hashes and ten final-response hashes match. All ten records have distinct project and session IDs, identical recorded configurations, empty visible-tool arrays and true final-output-verification flags. The original attachment's hash remains unchanged.
- **Recorded settings:** builtin Codex subscription, model `gpt-6-astra`, medium reasoning, memory disabled, delegation denied, no configured compute hosts. Tool calls were not observed in the packet. This is not an independent inspection of unseen session state.
- **Freeze and scoring history:** materials were frozen at 11:41:13.412488 UTC. Rules were recorded at 11:43:57.736070 UTC after the operator had seen T01/T02. The supplied timestamps show T01–T07 had started and T01–T05 had completed by then; completion does not establish which additional answers the operator read. This is not a fully preregistered or blinded scoring exercise.
- **Timing deviation:** an idle flag appeared while output was still streaming. T03–T10 started in order but adjacent runs overlapped; the seven overlaps range from 5.469 to 12.027 seconds. The realized collection was ordered initiation, not strictly sequential completion.
- **Repair and its limits:** the packet reports checking official completed CLI outputs and preserving initial partial captures privately, without rerunning inference. This review scores only supplied final answers. Their hashes agree with the recorded terminal-output hashes; the raw per-trial status field still reads `idle`. The live CLI and private partial captures were not independently rechecked.
- **Exploratory inference only:** ten purposively selected theorems, one response each, no human expert review and no independent second adjudicator. Entries within a theorem are dependent; no confidence interval, pooled accuracy percentage, model-wide estimate or cross-field ranking is justified. The chart is a descriptive inventory.
- **Context and source limits:** the prompts included definitions and explicit scope notes, which can aid retention. Source length and technical vocabulary vary. Fidelity to these supplied papers does not verify their theorems, proofs or publication status.

## Files

- `item-adjudications.csv`: all 75 criterion judgments, semantic roles, response evidence, source evidence and rationale.
- `adjudication-evidence.json`: complete audit, all judgments, unchanged collection packet and frozen materials.
- `checklist-summary.png`: compact descriptive chart; geometry and visual legibility checked.
