# Exploratory theorem-explanation study: researcher packet

Prepared 2026-10-08. Preparation only: ten source statements and checklists; no explanations collected, generated, or scored.

## Question and scope

Research question: which stated hypotheses, domains, quantifier dependencies, and exceptions survive a plain-English explanation? There is no empirical answer yet.

“Necessary” below means necessary to preserve the scope of the supplied statement, not proved logically minimal hypotheses. A hypothesis might be removable by a different theorem. These are provisional source-fidelity checklists, not proof verification or a mathematical gold standard. Conclusions and definitions are marked separately from assumptions.

The input identifies repository https://github.com/openai/math at commit fd4aeeb2ee4fc729c18d98444fed42fd0529eeeb, retrieved 2026-10-08T10:23:39.060769+00:00, sample seed 20261008. This is attachment-supplied provenance, not an independent repository audit. Its source files claim substantial mathematical results; their correctness, originality, and publication status have not been verified. Do not report an explanation's agreement with these files as evidence that the mathematics is valid.

## Selection and reproducibility

The attachment contains 20 sampled paper families and 221 source records. A bounded scan found 43 literal LaTeX theorem environments. Ten statements were purposively selected, one per paper, for explicit conditions, manageable contextual dependencies, and varied mathematical language. This is not a random selection of ten theorems, and the attachment's seed does not make this second-stage selection random.

Preference went to main statements with nearby definitions; the relative-generation statement was chosen instead of its shorter corollary-like main theorem because it exposes topology and irreducibility. The contingency-table introduction theorem was selected instead of the internal exact-sampling theorem because it states input and complexity conditions directly. Selection used source text only, before any explanations.

| ID | Family | Statement | Selection reason |
|---|---|---|---|
| T01 | 082 | Full annular variation | Strict exponent bound, function spaces, constant dependency |
| T02 | 113 | Perfect-matching FPRAS | Relative error, failure probability, certain zero case |
| T03 | 115 | Contingency-table sampling | Exact versus approximate laws; expected versus bounded time |
| T04 | 127 | Average sensitivity | Uniform input law, degree bounds, convention at zero |
| T05 | 148 | Self-similar measure dimension | Entropy definition, contraction domain, allowed overlaps |
| T06 | 154 | Multiple ergodic averages | Inherited system assumptions and dependent null sets |
| T07 | 164 | Finite sums and products | For-every/there-exists order and nonempty subsets |
| T08 | 210 | Sixth symmetric power | Field, parameter range, direction of injection |
| T09 | 296 | Relative generation | Irreducibility, separable predual, specific topology |
| T10 | 372 | Isotropic elasticity uniqueness | Smoothness, positivity, equality of complete operators |

Other supplied families were not selected: 015, 024, 046, 065, 192, 215, 230, 245, 351, 369. Their exclusion is a scope choice, not a quality or validity judgment. No balance across all fields is claimed.

The machine-readable source packet records exact source paths, supplied URLs, theorem labels, 1-based source lines, 0-based Unicode character spans with exclusive end, and SHA-256 hashes of UTF-8 text. Line ranges can share a boundary line because some supplemental spans start after the theorem on that line; character spans are authoritative. Text is decoded JSON content preserved without mathematical editing. All ten full-source hashes matched the supplied hashes. These are LaTeX source passages, not OCR or rendered-PDF transcriptions; PDF page numbers are not asserted.

The original upload Version identifier could not be resolved by the notebook provenance service. The exact notebook input path was readable; input-file hashing and source locators provide the fallback audit trail. No independent download was made.

## Keeping collection separate

Use collection-prompts.md in ten separate, fresh conversations. Copy only one complete block, from BEGIN PROMPT to END PROMPT, into each conversation. The target theorem, context, and any source scope note are included; no researcher checklist or rationale is included. Do not upload this researcher packet, the source audit packet, or other responses into those conversations. Source definitions inevitably contain conditions: they are part of the task material, not hidden checklist annotations.

The same short instruction is used for all ten. No hard word limit is imposed because forced brevity could itself cause omission. This design measures explanation with supplied context; it does not measure theorem recall from memory or the effect of withholding definitions. LaTeX accessibility and passage-length differences remain potential confounders.

Record model and displayed version, provider, date/time, prompt ID, exact prompt, model settings if available, tool access, and the entire unedited response (including refusals or uncertainty). Use one initial response per item; record any retry separately. Keep model/settings consistent for an initial run where possible. Avoid follow-up questions that reveal checklist items. A neutral shuffled order is optional; record the order actually used.

Freeze these files before collection. If mathematical review later changes a checklist, retain both versions and document the reason rather than silently changing criteria after seeing responses. Expert review of T05, T09, and T10 is particularly useful for technical definitions. No scoring categories, thresholds, or hypothetical response evaluations are supplied here.

## How to read the checklists

S = target statement, C = supplied pre-statement context, N = supplied scope note. These roles map directly to passages in source-packet.json and collection-prompts.md. Short quoted anchors identify evidence within those passages. All boxes are unassessed; they are not observations about a model.

H/D = hypothesis or domain; Q = quantifier/dependency; R = result or defining relationship; E = exception, permitted edge case, or scope limit. Interpretive cautions are provisional deductions, not extra source hypotheses. Equivalent wording can preserve a condition without repeating its terminology.

### T01 — Full annular variation

- [ ] H/D: Real r strictly greater than 2; the endpoint r=2 is outside this statement. Anchor S: “For every real $r>2$”.
- [ ] H/D: Both inputs are complex functions in L³(R²), not arbitrary functions on any space. Anchor S: “for all complex”.
- [ ] R: The output is the annular variation V_r defined from B with shifts (x+t,y) and (x,y+t) and kernel dt/t. Anchor C: eq:bilinear-truncation.
- [ ] R: Preserve the L^(3/2) norm bound by C_r times the product of the two L³ norms. Anchor S: displayed inequality.
- [ ] Q: For each r there exists a finite C_r that works for all F,G; the constant is not selected separately for each pair. Anchor S: “there is a finite constant” before “for all”.
- [ ] D/Q: Variation takes a pointwise supremum over all finite increasing positive rational endpoint lists; partitions may depend on location. Anchor C: eq:variation-definition and “The supremum is pointwise”.
- [ ] E: Finite integrals use a common null-set convention; quantities are set to zero there. Anchor C: “one common null set”.
- [ ] E: There is no restriction on the number of endpoints within a dyadic scale interval. Anchor C: “no restriction”.

Uncertainty/scope: No uniform control of C_r as r approaches 2 is stated. This norm estimate should not be rewritten as a uniform bound at every spatial point. Those are scope deductions; the statement does not claim endpoint failure. Rational versus real endpoints need not become a substantive restriction: C explicitly says continuity leaves the supremum unchanged. The proof of that continuity was not audited.

### T02 — Perfect-matching approximation

- [ ] H/D: G is finite, simple, and undirected. Anchor C: first sentence.
- [ ] R: Z(G) counts perfect matchings, meaning each vertex is incident to exactly one selected edge. Anchor C: definition of Z(G).
- [ ] H/D: Epsilon and delta are rational with 0<epsilon<1 and 0<delta<1/2. Anchor S: “rational parameters”.
- [ ] Q: A single uniform classical randomized algorithm accepts these inputs. Anchor S: “There is a uniform classical randomized algorithm”.
- [ ] R: Output is a nonnegative rational estimate with multiplicative relative error at most epsilon, with probability at least 1−delta. Anchor S: displayed event.
- [ ] E: If Z(G)=0, output is zero with certainty, not merely high probability. Anchor S: “zero with certainty”.
- [ ] R/Q: Worst-case bit runtime is polynomial in input encoding length, epsilon inverse, and log(delta inverse). Anchor S: final sentence.

Uncertainty/scope: The probability is interpreted over algorithm randomness, not a random graph distribution. “Uniform” describes the algorithm across inputs, not uniform sampling of estimates. The source theorem asserts an existence/complexity result; it supplies neither a small practical exponent here nor a guaranteed exact count for positive Z. Graphs with loops, parallel edges, or directed edges are outside the stated input domain, without implying that no extensions exist.

### T03 — Contingency-table sampling

- [ ] H/D: Tables have nonnegative integer entries and prescribed nonnegative integer row/column margins with equal total N. Anchor C: definition of Omega; S: opening sentence.
- [ ] D: All cells are allowed; there are no additional cell bounds or forbidden positions. Anchor C: “There are no cell bounds or forbidden positions”.
- [ ] D/Q: Both dimensions m,n vary and margins are binary encoded. Anchor C: “both dimensions vary” and “encoded in binary”.
- [ ] R: Uniform means equal mass for every feasible table. Anchor C: “every table equal mass”.
- [ ] H/R: For each integer k≥1, part (i) returns a feasible table within 2^(−k) total variation of uniform. Anchor S: part (i); C fixes TV as half the sum of absolute probability differences.
- [ ] R: Part (i) has bounded runtime polynomial in m,n,log(N+1),k. Anchor S: part (i).
- [ ] R/E: Part (ii) is exactly uniform, terminates almost surely, and has expected runtime polynomial in m,n,log(N+1). Anchor S: part (ii).
- [ ] D: Computation uses unbiased random bits and bit operations. Anchor S: “model of unbiased random bits and bit operations”.
- [ ] Q/E: Polynomials are uniform over margins; no positivity, sparsity, balance, or fixed-dimension assumption is imposed. Anchor S: final two sentences.

Uncertainty/scope: “No positivity” means strict positivity is unnecessary, while nonnegativity remains required. Zero totals/margins are included by the wording. Expected polynomial time for exact sampling must not be promoted to a bound on every execution: N explicitly mentions a rare exhaustive branch. Dimensions are matrix dimensions under ordinary positive-integer conventions; zero-dimensional matrices are not expressly discussed. No claim about constrained-cell variants is supplied.

### T04 — Polynomial-threshold average sensitivity

- [ ] H/D: Integers n≥1 and 1≤d≤n. Anchor S: opening sentence.
- [ ] H/D: p is a real multilinear polynomial of degree at most d. Anchor S: second sentence.
- [ ] D/R: f maps the ±1 Boolean cube to ±1 by the sign of p. Anchor S: definition of f.
- [ ] E: Zero is assigned sign +1; p is allowed to vanish on the cube. Anchor S: “sgn(0)=1”; N: “may vanish”.
- [ ] D: Average sensitivity is under the uniform distribution on the cube. Anchor S: “under the uniform law”; C: definition of I(f).
- [ ] R: I(f) is the sum of coordinate-flip disagreement probabilities and is at most 8d√n. Anchor C: eq:influence-definition; S: inequality.
- [ ] Q/E: The constant does not depend on d or n; d may grow with n; no regularity condition is imposed on p. Anchor N.

Uncertainty/scope: Degree at most d does not demand exact degree d or exclude constant polynomials. This is an average of coordinate effects, not a bound of the same size on sensitivity at every vertex. C explains multilinearization on the cube, but a broad extension to arbitrary input distributions is not provided. Numeric factor 8 is part of the stated bound; plain-language paraphrase may retain it as notation.

### T05 — Self-similar measure dimension

- [ ] H/D: A finite, nonempty indexed family of affine maps on the real line, with real translations and 0<|r_i|<1. Anchor S and C: definition of Phi.
- [ ] H/D: p is a strictly positive probability vector: each weight is positive and the weights sum to one. Anchor S: “strictly positive probability vector”.
- [ ] R: The measure is the Borel probability measure satisfying the weighted pushforward fixed-point equation. Anchor C: definition of mu.
- [ ] R: h_RW is the limit H(G_n)/n for random composed affine maps; equal maps aggregate their word probabilities. Anchor C: law of G_n and eq:rate-and-lyapunov.
- [ ] R: chi = −sum p_i log|r_i|, with a common logarithm base; it is positive. Anchor C: eq:rate-and-lyapunov.
- [ ] R: Hausdorff dimension of the measure equals min{1,h_RW/chi}, not uncapped entropy divided by contraction. Anchor S: displayed equality.
- [ ] D: C specifies the lower Hausdorff dimension of a measure using positive-mass Borel sets and discusses exact dimensionality. Anchor C: “We use the lower Hausdorff dimension”.
- [ ] E: No separation hypothesis; exact overlaps and repeated indexed generators are allowed. Anchor S: final sentence.
- [ ] E: Negative contraction ratios are allowed by the absolute-value condition. Anchor C: “negative ratios”.

Uncertainty/scope: This is a highly technical claimed formula. Mathematical validity and the cited exact-dimensionality input have not been checked. h_RW must not automatically be replaced by the one-step symbol entropy H(p); the text distinguishes them. Zero weights and zero contraction ratios are outside the literal theorem, even if some reductions might extend it. Dimension of the measure should not silently become dimension of its support. Expert adjudication is recommended for a paraphrase that uses “information rate” without specifying what is counted.

### T06 — Multiple ergodic averages

- [ ] H/D: A probability space; T is invertible, bimeasurable, and measure preserving. Anchor C: opening sentence.
- [ ] H: T is mixing in the stated sense: for every measurable A,B, mu(A intersect T^(−r)B) tends to mu(A)mu(B) as |r|→infinity. Anchor C: eq:mixing; S invokes it.
- [ ] Q/D: For every integer n≥2 and every fixed tuple f_1,...,f_n in L-infinity(mu). Anchor S.
- [ ] R: Averages are N inverse times the sum over k=1,...,N of the product f_j(T^(jk)x). Anchor C: definition of A_N.
- [ ] R: Limit is the product of the separate integrals, not the integral of the pointwise product. Anchor S: displayed limit.
- [ ] E/Q: Convergence is mu-almost everywhere for each fixed tuple; the exceptional null set may depend on the system and tuple. Anchor S and N.
- [ ] R/Q: N tends through all positive integers. Anchor S: final sentence.
- [ ] E: Space need not be standard; no mixing rate is required; no single common exceptional set for all bounded functions is required. Anchor S and N.

Uncertainty/scope: The inherited invertibility and bimeasurability conditions must remain in scope despite not being repeated in S. Mixing cannot be replaced by mere ergodicity. “Almost every point” is not all points, and quantifiers over tuples cannot be interchanged with choosing one full-measure set. L-infinity means essentially bounded; the source uses bounded measurable representatives. No quantitative convergence rate is asserted.

### T07 — Finite sums and products

- [ ] H/D: Integers r,m≥1 and fixed reals R≥2,D≥1; coloring of positive integers by r colors. Anchor S; C fixes N={1,2,...}.
- [ ] Q: For every such coloring there exist a_1<...<a_m and one color c. Choices may depend on the coloring and the fixed parameters. Anchor S: “For every coloring” then “there are”.
- [ ] D: Chosen numbers are distinct positive integers, exactly m in number. Anchor S.
- [ ] Q/R: For every nonempty subset J of [m], both its sum and its product have that same color c. Anchor S: displayed equality.
- [ ] D/E: Each selected element occurs at most once in an individual expression; singleton subsets are included, empty subsets excluded. Anchor C: “at most once” and “singleton subsets”.
- [ ] R: The same selected sequence can additionally satisfy a_1>R and a_d>R(sum of previous a_k + product of previous a_k)^D for every 2≤d≤m. Anchor S: eq:separated-elements.

Uncertainty/scope: Arbitrarily large finite m does not by itself assert a single infinite sequence with the property. Distinct generators do not by themselves mean every sum/product expression is distinct. The separation inequality is an additional promised property, not an assumption on a set supplied as input. Its d is an index, distinct from exponent D. At m=1 the later-index inequalities have no instances.

### T08 — Sixth symmetric power

- [ ] H/D: Integer b≥6 and every finite-dimensional complex vector space V. Anchor S.
- [ ] Q: For every such b,V there exists an injection. Anchor S: quantifier order.
- [ ] R: Direction is Sym^6(Sym^b V) into Sym^b(Sym^6 V). Anchor S: displayed arrow.
- [ ] R: Injection is GL(V)-equivariant, preserving the group action, not merely an injection of underlying sets. Anchor S.
- [ ] R: Equivalent formulation is Schur positivity of h_b[h_6]−h_6[h_b]. Anchor S; C defines Schur positivity as nonnegative integer Schur coefficients.
- [ ] E: No restriction on finite dim V beyond the stated domain; b=6 is included. Anchor N: “no restriction” and “diagonal case”.

Uncertainty/scope: The existence statement does not identify the canonical Foulkes–Howe map as injective. No isomorphism or equality of the two representations is asserted for general b. Complex field and finite dimension should not disappear; other fields or infinite-dimensional V are not covered here. A sixth-power case should not become the assertion for all outer powers a. The equivalence to Schur positivity is taken from the source, not independently proved.

### T09 — Relative generation of finite factors

- [ ] H/D: P subset M is an inclusion of type II_1 factors. Anchor S.
- [ ] H: Inclusion is irreducible, defined by P-prime intersect M = complex scalar multiples of the identity. Anchor C: “irreducible”.
- [ ] H/D: M has separable predual. Anchor S; N distinguishes this from operator-norm separability.
- [ ] D/R: u ranges over unitary elements of M. Anchor S: U(M); C defines that notation.
- [ ] R: Generation means W*(P,u)=M; adjoints and weak operator closure are included, and the existing subfactor P is retained. Anchor S and C: definition of W*.
- [ ] R: Such unitaries form a dense G_delta subset, not a claim about every unitary. Anchor S.
- [ ] D: Density and G_delta refer specifically to the trace 2-norm topology, with normalized trace tau and norm tau(x*x)^(1/2). Anchor S and C.

Uncertainty/scope: “Generic” is a topological description here, not a probability-one claim for an unspecified random unitary. G_delta means countable intersection of open sets (standard terminology, not separately defined in the excerpt). Do not replace the 2-norm topology by operator norm or drop P from generation. The usual unital inclusion convention is implicit in this subject; it is not added as a new explicit criterion. Expert review is recommended if an explanation reduces irreducibility to an informal slogan.

### T10 — Isotropic elasticity uniqueness

- [ ] H/D: Omega is a bounded connected domain in R³ with C-infinity boundary. Anchor S.
- [ ] H/D: Both coefficient pairs lambda_j,mu_j are real and C-infinity on the closure, j=1,2. Anchor S.
- [ ] H: mu_j>0 and 3lambda_j+2mu_j>0 on the entire closure. Anchor S. Strict inequalities apply to both pairs.
- [ ] D: Lambda is the static isotropic displacement-to-traction operator defined using the stated strain/stress equation. Anchor C: eq:physical-operator and eq:physical-dn.
- [ ] H/R: Equality is of the complete maps H^(1/2)(boundary;R³)→H^(−1/2)(boundary;R³), for the same domain. Anchor S: operator equality.
- [ ] R: It implies both lambda_1=lambda_2 and mu_1=mu_2 throughout Omega. Anchor S: conclusion.
- [ ] E: Data are the full zero-frequency boundary operator in fixed Euclidean coordinates. Anchor N.
- [ ] E: Analyticity, closeness to constants, and prior coefficient agreement near the boundary are not additional hypotheses. Anchor N.

Uncertainty/scope: Lambda equality denotes exact complete operator data, not approximate, noisy, finite, or partial-boundary measurements. The claim is uniqueness, not a stability bound or practical reconstruction algorithm. The source's C-infinity-on-closure convention is retained; alternative extension conventions were not resolved. The positivity condition does not separately require lambda>0. The theorem's mathematical validity and operator-analytic details require expert assessment beyond this preparation.

## Limits on interpretation

Ten purposively chosen statements cannot establish how often assumptions are lost across mathematics, papers, or models. Conditions within one theorem are dependent; treating all checklist boxes as independent observations would overstate the sample size. Differences between fields, technical vocabulary, source length, and model familiarity may affect responses. Explicit scope notes may help a model preserve exceptions, so results apply to this context-supplied protocol.

No checklist asserts that its items must all be recited verbatim. A faithful explanation can preserve scope through equivalent language or clearly delimited notation. Whether a particular omission changes meaning remains an interpretive judgment to be made after collection, with disagreements and unresolved cases recorded rather than concealed.
