# Eight-model download audit — corrected edition

**Audit date: 8 October 2026.** This report uses user-supplied official-source snapshots and four replacement text records preserving original line endings. It is not live verification of pages fetched in this session. No weights or training-data shards were downloaded and no models were run.

**Seven releases list weights, configuration and tokenizer materials and report `gated:false`. Mistral Large 4 promises weights by October-end.** Nemotron and OLMo provide substantial model-specific training materials, but exact end-to-end training reproduction is not established by these snapshots.

The [corrected comparison](download-audit-corrected.html) contains all 80 component judgments, links, excerpts and revision IDs. The [CSV](download-audit-comparison-corrected.csv), [evidence ledger](download-audit-evidence-corrected.json) and [visual summary](download-audit-summary-corrected.png) use the corrected findings.

## Final integrity results

| Result | Records |
|---|---:|
| Exact recorded SHA-256 **and** UTF-8 byte-count match | **76** |
| Unresolved mismatch | **0** |
| Original oversized retrieval failures, no content | **3** |
| Total source records | **79** |

All four replacements match their **original** hashes, sizes, retrieval timestamps and other recorded metadata. Comparison with the first attachment confirms that only line endings changed.

| Source ID | Corrected record | Exact bytes | Preserved CRLF endings |
|---|---|---:|---:|
| S002 | DeepSeek LICENSE | 1,084 | 20 |
| S007 | Qwen LICENSE | 11,544 | 201 |
| S008 | Qwen model card | 65,012 | 582 |
| S038 | NVIDIA license page | 300,192 | 184, among 6,863 total LF characters |

The earlier three CRLF reconstruction cases now match directly; NVIDIA's mixed-line-ending mismatch is resolved. All **14 comparable GitHub files** still match their tree-recorded Git blob SHA-1.

**S034, S035 and S039 remain failed oversized dataset metadata retrievals.** Their separately successful reduced-field replies (S075–S077) are retained as distinct records. They do not retroactively make the failed requests successful, and they do not establish complete shard inventories.

Original input SHA-256: `edceb401799bf19058d7b7365f585cc061d4b2d180aa1a4ca3373ec588f1d9b7`  
Corrections input SHA-256: `119c558b76af56aa4d27ca2e8d5cd6f6fadc0b1c4dceb28bc0e4f9b7903e4f88`

Hash agreement establishes internal consistency of the supplied snapshots, not independent authentication of their origin. The corrected ledger preserves a correction history and marks the former integrity issues resolved.

## Availability and reproduction

“Public” means supplied contents or a file listing evidences the named material, not that its full bytes were downloaded or tested. “Partial” names a specific coverage or inspection limitation. “Not found” is limited to the checked sources; “unknown” means insufficient evidence. Promises and access gates are separate categories. No weight access gate is evidenced by the seven metadata snapshots.

| Release | Weights evidenced | What remains for reproduction |
|---|---|---|
| **Mistral Large 4** | Promise for October-end; API preview described | Obtain weights, license, configuration/tokenizer and implementation first. Original corpus, processing, trainer, pre/post-training settings, checkpoints and evaluation assets are unverified. Broad multilingual-data and hardware descriptions are insufficient. |
| **DeepSeek-V4-Pro** | 64 safetensors shards | Dedicated inference and custom encoding code are listed, but the original trainer/corpus is not found in the inspected release. Need the >32T corpus, processing, Muon/distributed settings, expert SFT/GRPO data and rewards, distillation teachers and run state, and complete evaluation harnesses. |
| **Qwen3.8-27B** | 18 shards | Serving examples and generic fine-tuning links do not provide the original recipe. Need exact multimodal data and processing, mixtures, pre/post-training settings and teacher/reward/agent materials. Evaluation also needs corrected labels and precise harness/judge versions. |
| **GLM-5.3** | 141 shards | Need GLM-5.2 base lineage plus the exact 5.3 post-training data, rewards, orchestration and settings. A family-level slime link is not the recipe. Modified evaluation harnesses/containers and judge access remain unverified. |
| **Kimi-K3** | 96 shards | Model/processor/tokenization modules are listed, not an original trainer. Need corpus, preprocessing, pretraining/post-training/QAT code, teachers/rewards and run state. Evaluation needs Kimi Code setup, in-house tasks and the H20-calibrated branch. The listed technical-report PDF was not supplied for reading. |
| **Gemma 4 31B-IT** | 2 shards | Transformers examples support inference. Broad data categories, cutoff and filtering descriptions are not downloadable corpora. Need exact source versions/mixtures, processing, trainer/settings, instruction-tuning/teacher/reward materials and fuller evaluation instructions. |
| **NVIDIA Nemotron 3 Super 120B-A12B-BF16** | 50 shards | Model-specific stage recipes and evaluation instructions exist. Some dataset files are evidenced, but original internal-only and licensed data are explicitly missing from the public corpus. Need these inputs, exact blends/settings, generators and historical run lineage. Validate containers and the exact BF16 evaluation endpoint. |
| **Olmo-3-1125-32B** | 14 shards | Actual stage scripts are supplied and official release scripts/manifests are listed. Need pinned shard inventories, exact script-to-release correspondence, cloud checkpoint/data access, optimizer state and evaluation settings. Resolve dataset identity and final checkpoint-averaging discrepancies. This is the base model; Think/Instruct training is outside its reproduction scope. |

Shard counts are filename counts, not verified tensor inventories or hardware requirements.

**Data versus descriptions.** DeepSeek and Gemma provide broad data descriptions, while exact-release data descriptions for Qwen3.8, GLM-5.3 and Kimi-K3 remain unverified in the supplied material. Statements about other family members are not transferred to these releases. NVIDIA supplies detailed public/private dataset descriptions and one audited metadata listing of three Cascade-RL-SWE JSONL files; other collection links are not complete audited inventories. OLMo supplies stage descriptions, manifest paths and reduced dataset metadata, not inspected shard contents.

**Recipes versus libraries.** Generic fine-tuning examples are excluded from evidence of original training. Nemotron's model-specific pretraining/SFT/RL scripts and OLMo's stage scripts are stronger evidence, but do not establish identical reconstruction of the historical runs.

Two substantive OLMo discrepancies remain: its requested 5.5T-1125 dataset URL returns the ID `allenai/dolma3_mix-6T`, without proven historical equivalence; its model card says four final checkpoints were averaged, while the official training README says three. Some logs are described as forthcoming and linked reports were not inspected. NVIDIA's recipe also estimates 8–10T public tokens while calling that 40–50% of 25T; the arithmetic is 32–40%, so no precise audited coverage percentage is asserted.

## Corrected license findings

| Release | License evidence |
|---|---|
| Mistral Large 4 | Terms for the promised weight release unverified. |
| DeepSeek-V4-Pro | MIT text inspected; card explicitly covers repository and weights. Corrected record matches directly. |
| Qwen3.8-27B | Apache-2.0 LICENSE text inspected; corrected record matches directly. |
| GLM-5.3 | Custom license: MaaS operators exceeding $10B aggregate affiliate revenue over any consecutive 12 months require Z.AI security review before commercial use. |
| Kimi-K3 | Custom license: MaaS operators exceeding $20M aggregate affiliate revenue over any consecutive 12 months need a separate agreement. Commercial products/services exceeding 100M monthly active users or $20M monthly revenue require prominent Kimi K3 attribution. Sections 2–3 exempt defined internal use and access through official products/certified partners. |
| Gemma 4 31B-IT | Apache-2.0 supported by the card and inspected official Gemma 4 license page. No older Gemma restrictions are inferred. |
| NVIDIA Nemotron 3 Super BF16 | **NVIDIA Nemotron Open Model License text now passes exact integrity checks.** Commercial use and derivative distribution are permitted subject to its conditions, including license/notice retention, litigation termination, indemnity and trade compliance. This is a custom license, not Apache-2.0. |
| Olmo-3-1125-32B | Apache-2.0 declared in card/metadata, but no model license body supplied. Dataset metadata labels ODC-BY without inspected terms. |

NVIDIA's license-evidence category changes from **Partial** to **Public material evidenced** because the inspected text is now hash-consistent. This resolves the integrity limitation; it does not remove the license conditions or the missing training data. The corrected HTML ledger links the exact source texts and relevant clauses.

The model availability and training-reproduction judgments otherwise remain unchanged. Unknown materials remain unknown, scoped negative findings remain scoped, and the three oversized retrievals remain failures.
