<!doctype html><html lang="en"><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>What can you download? Eight-model snapshot audit</title><style>body{font:15px/1.5 system-ui,sans-serif;color:#192734;background:#fafcfd;margin:0}main{max-width:1600px;margin:auto;padding:28px}h1{font-size:29px;line-height:1.2}h2{font-size:21px;margin-top:32px}.lead{max-width:1050px}.notice{border-left:4px solid #72549c;background:#f0ebf7;padding:14px 18px;margin:18px 0}.legend{display:flex;flex-wrap:wrap;gap:12px}.legend span{padding:6px 10px;border:1px solid #bbc7cf;border-radius:4px}.A{background:#d8edf4}.P{background:#ffe1b3}.N{background:#e4e4e4}.U{background:#f7f7f7}.R{background:#e4d9f5}.G{background:#c5e4df}.badge{display:inline-block;padding:1px 7px;border:1px solid #9aa9b3;border-radius:4px;font-weight:700;margin-right:5px}.scroll{overflow:auto;border:1px solid #c8d2da;max-height:78vh;margin-top:15px}table{border-collapse:separate;border-spacing:0;background:white;font-size:13px}th,td{padding:12px;border-right:1px solid #dce2e7;border-bottom:1px solid #dce2e7;vertical-align:top;min-width:220px;max-width:260px}thead th{position:sticky;top:0;background:#17394b;color:white;z-index:2}tbody th{position:sticky;left:0;background:#edf3f6;z-index:1;min-width:185px;width:185px}thead th:first-child{left:0;z-index:3}a{color:#075e85;overflow-wrap:anywhere}thead a{color:white}.refs{display:block;margin-top:8px}code{overflow-wrap:anywhere}.small{font-size:13px;color:#415565}details{margin:12px 0;padding:12px;background:white;border:1px solid #d3dce2;border-radius:4px}summary{cursor:pointer;font-weight:600}pre{white-space:pre-wrap;overflow-wrap:anywhere;background:#f5f7f8;padding:12px;font-size:12px;max-height:450px;overflow:auto}input{padding:9px;width:320px;max-width:90%;font:inherit;margin-top:15px}.model-detail{background:#fff;padding:12px 18px;margin:12px 0;border-left:3px solid #54819a}@media print{main{padding:8px}.scroll{max-height:none;overflow:visible}table{font-size:9px}th,td{min-width:90px;padding:4px;position:static!important}.no-print{display:none}details{break-inside:avoid}}</style><main><h1>What can you download when a model is called open?</h1><p><strong>Corrected edition:</strong> four exact-byte records incorporated; all 76 content-bearing sources pass original hash and byte-count checks. Three oversized retrieval failures are retained.</p><p class="lead"><strong>Eight-release comparison · 8 October 2026.</strong> Based on user-supplied official-source snapshots collected 2026-10-08T09:04:47.226524+00:00. Pages were not fetched in this session. No weight or dataset shards were downloaded; no models were executed.</p><div class="notice"><strong>Main finding:</strong> seven model repositories list weights and report <code>gated:false</code>; Mistral Large 4 promises weights by the end of October. Nemotron and OLMo provide substantially more training material, but the supplied evidence does not establish exact end-to-end reproduction of any model.</div><p class="small">A file listing establishes evidence of a published file, not that its bytes were acquired or tested. Categories describe evidence coverage, not model capability or a numerical openness score. Missing from an inspected repository differs from unknown elsewhere.</p><div class="legend"><span class="A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test."><b>A</b> Public material</span><span class="P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation."><b>P</b> Partial</span><span class="N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence."><b>N</b> Not found in checked sources</span><span class="U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected."><b>U</b> Unknown</span><span class="R" title="Promised: source explicitly describes a future release; not counted as currently downloadable."><b>R</b> Promised</span><span class="G" title="Gated: an access gate is explicitly evidenced. No requested model weight repository is marked gated in these snapshots."><b>G</b> Gated (none evidenced for weights)</span></div><label class="no-print"><br>Filter models: <input id="filter" placeholder="e.g. Gemma or Nemotron" aria-label="Filter model rows"></label><div class="scroll"><table><thead><tr><th>Release</th><th>Weights</th><th>Configuration</th><th>Tokenizer / processing</th><th>Inference code</th><th>Training libraries / code</th><th>Original model training recipe</th><th>Training-data descriptions</th><th>Downloadable training data</th><th>Evaluation instructions</th><th>License evidence</th></tr></thead><tbody><tr data-model="mistral large 4"><th scope="row">Mistral Large 4<br><span class="small">Announcement only</span></th><td><span class="badge R" title="Promised: source explicitly describes a future release; not counted as currently downloadable.">R</span>October 6 announcement promises weights by month-end (October 2026); preview API is available. No weight listing supplied.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>No configuration file or model repository listing in the supplied evidence; architecture details promised.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>No tokenizer files or listing supplied.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Preview API described; downloadable model-specific inference code not verified.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Mistral Forge named as the training/customization/RL environment; source release not verified.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>No original training pipeline inspected. Further post-training methodology is promised, not a promise of training code.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Broad description: substantial multilingual data covering 160+ languages; no exact corpus manifest or mixture.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>No downloadable training dataset evidenced; no dataset repository snapshot supplied.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Benchmarks and selected methodological descriptions, but no runnable evaluation package inspected; further benchmarks promised.<span class="refs"><a href="#S033">S033</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>No license for the promised weight release verified. Marketing use of open-weight does not establish terms.<span class="refs"><a href="#S033">S033</a></span></td></tr><tr data-model="deepseek-v4-pro"><th scope="row">DeepSeek-V4-Pro<br><span class="small">Revision <code>b5968e9190ef</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>64 safetensors shards and index listed; public, gated=false. Config specifies FP8 quantization, not an all-BF16 checkpoint.<span class="refs"><a href="#S001">S001</a> · <a href="#S004">S004</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>config.json read; generation_config.json listed. deepseek_v4 architecture.<span class="refs"><a href="#S001">S001</a> · <a href="#S004">S004</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json listed and tokenizer_config.json read. No Jinja template by design: custom encoding scripts/tests and documentation are supplied.<span class="refs"><a href="#S001">S001</a> · <a href="#S005">S005</a> · <a href="#S003">S003</a> · <a href="#S079">S079</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>inference/model.py, kernel.py, convert.py and generate.py listed; README provides conversion and torchrun commands. Source bodies of those Python files were not supplied.<span class="refs"><a href="#S001">S001</a> · <a href="#S078">S078</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No training implementation in the inspected model repository listing; Muon/GRPO are method descriptions, not released training code.<span class="refs"><a href="#S001">S001</a> · <a href="#S003">S003</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No original pretraining/post-training recipe in the inspected release. Card describes domain experts, SFT/RL and distillation only.<span class="refs"><a href="#S001">S001</a> · <a href="#S003">S003</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Card reports &gt;32T diverse high-quality tokens and post-training stages; exact sources, filtering and mixtures unspecified in inspected text.<span class="refs"><a href="#S003">S003</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No corpus shards or downloadable corpus release identified in the inspected card/listing; availability elsewhere remains unknown.<span class="refs"><a href="#S001">S001</a> · <a href="#S003">S003</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Benchmark tables with shot counts/metrics and encoding instructions; no full benchmark runner/config package in the listing.<span class="refs"><a href="#S001">S001</a> · <a href="#S003">S003</a> · <a href="#S079">S079</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>MIT text inspected; card explicitly applies it to repository and weights. Corrected source preserves original CRLF and matches the recorded hash and byte count directly.<span class="refs"><a href="#S002">S002</a> · <a href="#S003">S003</a></span></td></tr><tr data-model="qwen3.8-27b"><th scope="row">Qwen3.8-27B<br><span class="small">Revision <code>1d4bf0f2ff60</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>18 safetensors shards and index listed; public, gated=false. This is the post-trained vision-language model; a coming-soon hosted service does not make the weights promised.<span class="refs"><a href="#S006">S006</a> · <a href="#S008">S008</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>config.json read; image/video preprocessing and generation configurations listed. Model reuses qwen3_5 architecture identifiers.<span class="refs"><a href="#S006">S006</a> · <a href="#S009">S009</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json, vocab.json, merges.txt and chat_template.jinja listed; tokenizer_config.json read.<span class="refs"><a href="#S006">S006</a> · <a href="#S010">S010</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Public usage/serving examples and links to Transformers, SGLang, vLLM and TokenSpeed; no self-contained inference implementation in either inspected release tree. Linked runtime source not inspected.<span class="refs"><a href="#S008">S008</a> · <a href="#S064">S064</a> · <a href="#S067">S067</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Generic Unsloth, Swift and Llama-Factory fine-tuning recommendations; linked libraries not inspected. They are not the original pipeline.<span class="refs"><a href="#S067">S067</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No original model training recipe in the inspected HF/GitHub trees. Generic fine-tuning guidance is excluded.<span class="refs"><a href="#S006">S006</a> · <a href="#S064">S064</a> · <a href="#S067">S067</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Exact 3.8-27B training corpus/mixture not described in inspected material. The family README&#x27;s trillions-of-tokens discussion is under Qwen3.5 and is not evidence of this release&#x27;s exact data.<span class="refs"><a href="#S008">S008</a> · <a href="#S067">S067</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No training corpus release in inspected model/project listings or cards; other sources unverified.<span class="refs"><a href="#S006">S006</a> · <a href="#S064">S064</a> · <a href="#S067">S067</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Benchmark footnotes identify prompts, harnesses, judges and some corrected labels; no complete runnable evaluation package or corrected-label files inspected.<span class="refs"><a href="#S008">S008</a> · <a href="#S064">S064</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Apache-2.0 LICENSE text inspected. Corrected license and model-card records preserve original CRLF and match their recorded hashes and byte counts directly.<span class="refs"><a href="#S007">S007</a> · <a href="#S008">S008</a></span></td></tr><tr data-model="glm-5.3"><th scope="row">GLM-5.3<br><span class="small">Revision <code>aca966e4e027</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>141 safetensors shards and index listed; public, gated=false. Config includes FP8 quantization.<span class="refs"><a href="#S011">S011</a> · <a href="#S014">S014</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>config.json read; glm_moe_dsa architecture and generation config listed.<span class="refs"><a href="#S011">S011</a> · <a href="#S014">S014</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json and chat_template.jinja listed; tokenizer_config.json read.<span class="refs"><a href="#S011">S011</a> · <a href="#S015">S015</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Serving instructions link SGLang, vLLM and TokenSpeed; no standalone implementation in the inspected GLM-5 project tree. Linked engines not inspected.<span class="refs"><a href="#S013">S013</a> · <a href="#S062">S062</a> · <a href="#S065">S065</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Family README links slime asynchronous RL infrastructure for GLM-5; slime source is not in this snapshot and does not establish the GLM-5.3 recipe.<span class="refs"><a href="#S065">S065</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No exact GLM-5.3 original pipeline in inspected trees. Card says it shares GLM-5.2&#x27;s base and gains come from post-training.<span class="refs"><a href="#S011">S011</a> · <a href="#S013">S013</a> · <a href="#S062">S062</a> · <a href="#S065">S065</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Exact GLM-5.3 data not verified. GLM-5 and GLM-5.3-Flash corpus figures in the family README are not attributed to this exact release.<span class="refs"><a href="#S013">S013</a> · <a href="#S065">S065</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No training corpus found in inspected release/project listings; no independent corpus inventory supplied.<span class="refs"><a href="#S011">S011</a> · <a href="#S062">S062</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Detailed benchmark footnotes specify sampling, harness versions, timeouts and modifications; result YAMLs listed. Full patched harnesses/judges not inspected.<span class="refs"><a href="#S011">S011</a> · <a href="#S013">S013</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Custom GLM-5.3 license, not MIT: model-as-a-service operators with aggregate affiliate revenue &gt;$10B over any consecutive 12 months require Z.AI security review before commercial use.<span class="refs"><a href="#S012">S012</a></span></td></tr><tr data-model="kimi-k3"><th scope="row">Kimi-K3<br><span class="small">Revision <code>f831ab668142</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>96 safetensors shards and index listed; public, gated=false. Card describes native MXFP4 weights/MXFP8 activations with QAT from SFT onward.<span class="refs"><a href="#S016">S016</a> · <a href="#S018">S018</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>config.json read; custom configuration, processor and vision-processing Python files listed.<span class="refs"><a href="#S016">S016</a> · <a href="#S019">S019</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tiktoken.model, tokenization_kimi.py and encoding_k3.py listed; tokenizer_config.json read.<span class="refs"><a href="#S016">S016</a> · <a href="#S020">S020</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Model implementation and processing/tokenization Python files listed; vLLM/SGLang/TokenSpeed deployment links provided. Python file bodies not supplied.<span class="refs"><a href="#S016">S016</a> · <a href="#S018">S018</a> · <a href="#S019">S019</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No original training implementation in inspected model/project trees; architecture/inference modules do not constitute a trainer.<span class="refs"><a href="#S016">S016</a> · <a href="#S063">S063</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No original pretraining/SFT/RL/QAT pipeline in inspected trees. Technical-report PDF is listed but its contents were not supplied.<span class="refs"><a href="#S016">S016</a> · <a href="#S018">S018</a> · <a href="#S063">S063</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Exact training data descriptions not verified from supplied card. Listed k3_tech_report.pdf was not inspected.<span class="refs"><a href="#S018">S018</a> · <a href="#S063">S063</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No training corpus files in the inspected model/project listings; availability elsewhere unknown.<span class="refs"><a href="#S016">S016</a> · <a href="#S063">S063</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Sampling settings and benchmark/harness notes supplied; result YAMLs listed. In-house tasks and adjusted GPU-task branches prevent claiming a full reproducible suite.<span class="refs"><a href="#S016">S016</a> · <a href="#S018">S018</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Custom Kimi K3 license: &gt;$20M aggregate revenue over any consecutive 12 months for MaaS operators triggers separate agreement; specified large commercial products require Kimi K3 UI attribution. Sections 2–3 have internal-use and official-product/certified-partner exceptions.<span class="refs"><a href="#S017">S017</a></span></td></tr><tr data-model="gemma 4 31b-it"><th scope="row">Gemma 4 31B-IT<br><span class="small">Revision <code>842da3794eaa</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>2 safetensors shards and index listed; public, gated=false in this snapshot. Do not infer a gate from older Gemma releases.<span class="refs"><a href="#S021">S021</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>gemma4 config.json read; processor_config.json and generation config listed.<span class="refs"><a href="#S021">S021</a> · <a href="#S023">S023</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json and chat_template.jinja listed; tokenizer_config.json read.<span class="refs"><a href="#S021">S021</a> · <a href="#S024">S024</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Transformers loading/generation examples are public in the card. Runtime implementation is external and not supplied as source.<span class="refs"><a href="#S022">S022</a></span></td><td><span class="badge U" title="Unknown: evidence is insufficient or the relevant material was not supplied/inspected.">U</span>Original training implementation or relevant library source not inspected; no separate training repository snapshot supplied.<span class="refs"><a href="#S021">S021</a> · <a href="#S022">S022</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>No original pretraining/instruction-tuning pipeline in the inspected model repo/card.<span class="refs"><a href="#S021">S021</a> · <a href="#S022">S022</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Card describes web/code/math/images and multilingual data, January 2025 cutoff, and safety/quality preprocessing; no exact source manifest or mixture.<span class="refs"><a href="#S022">S022</a></span></td><td><span class="badge N" title="Not found in the explicitly inspected card/listing/tree. This is a scoped negative finding, not proof of global absence.">N</span>Data descriptions supplied, not downloadable training shards. No corpus in the inspected model listing/card.<span class="refs"><a href="#S021">S021</a> · <a href="#S022">S022</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Benchmark results and safety-evaluation approach supplied, plus one result YAML listing; full runnable evaluation recipe not inspected.<span class="refs"><a href="#S021">S021</a> · <a href="#S022">S022</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Apache-2.0 confirmed by the model card and the linked official Gemma 4 license page; that page&#x27;s stored hash matches.<span class="refs"><a href="#S022">S022</a> · <a href="#S036">S036</a></span></td></tr><tr data-model="nvidia nemotron 3 super 120b-a12b-bf16"><th scope="row">NVIDIA Nemotron 3 Super 120B-A12B-BF16<br><span class="small">Revision <code>2dc98e2afe4f</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>50 safetensors shards and index listed; public, gated=false. This exact checkpoint is BF16; NVFP4 pretraining does not change its release identity.<span class="refs"><a href="#S025">S025</a> · <a href="#S027">S027</a> · <a href="#S026">S026</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>config.json read; configuration_nemotron_h.py and generation config listed.<span class="refs"><a href="#S025">S025</a> · <a href="#S027">S027</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json, special_tokens_map.json and chat_template.jinja listed; tokenizer_config.json read.<span class="refs"><a href="#S025">S025</a> · <a href="#S028">S028</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>modeling_nemotron_h.py and reasoning parser listed; card supplies vLLM/SGLang/TRT-LLM/Transformers examples. Implementation bodies not supplied.<span class="refs"><a href="#S025">S025</a> · <a href="#S026">S026</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Nemotron developer tree has training/data-prep code, with Megatron-Bridge, NeMo-RL and related frameworks identified; recipe README contents inspected.<span class="refs"><a href="#S050">S050</a> · <a href="#S072">S072</a> · <a href="#S073">S073</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Model-specific pretrain→SFT→RL→eval recipe and phase configs listed, with run commands; more than a generic library. It is an adapted public recipe, not a demonstrated exact replay of the original run.<span class="refs"><a href="#S050">S050</a> · <a href="#S072">S072</a> · <a href="#S073">S073</a> · <a href="#S026">S026</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Card names public, crawled, synthetic and explicitly private datasets; recipe describes 20T+5T pretraining and 34B+17B long-context stages.<span class="refs"><a href="#S026">S026</a> · <a href="#S073">S073</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Some real dataset files evidenced: Cascade-RL-SWE lists three JSONL files, ungated. Full corpus is explicitly incomplete: recipe names internal-only code/crawl/academic categories and card lists private NVIDIA/third-party data. Other collections are links, not audited inventories.<span class="refs"><a href="#S026">S026</a> · <a href="#S037">S037</a> · <a href="#S073">S073</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Model-specific reproducibility tutorial and evaluation configuration paths supplied, with launch commands and benchmark settings. No runs performed; containers, credentials and exact BF16 endpoint equivalence still need validation.<span class="refs"><a href="#S026">S026</a> · <a href="#S049">S049</a> · <a href="#S074">S074</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>NVIDIA Nemotron Open Model License text inspected; corrected mixed-line-ending page matches the original recorded hash and size. It permits commercial use and derivative distribution, subject to license/notice retention and its litigation-termination, indemnity and trade-compliance terms. It is a custom license, not Apache-2.0.<span class="refs"><a href="#S026">S026</a> · <a href="#S038">S038</a></span></td></tr><tr data-model="olmo-3-1125-32b"><th scope="row">Olmo-3-1125-32B<br><span class="small">Revision <code>c2b61dae89a1</code><br>gated: false</span></th><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>14 safetensors shards and index listed; public, gated=false. This is the base model, not Think/Instruct. Intermediate checkpoints are described but their branch listing was not supplied.<span class="refs"><a href="#S029">S029</a> · <a href="#S030">S030</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>olmo3 config.json read; generation_config.json listed.<span class="refs"><a href="#S029">S029</a> · <a href="#S031">S031</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>tokenizer.json, vocab.json, merges.txt and special_tokens_map.json listed; tokenizer_config.json read.<span class="refs"><a href="#S029">S029</a> · <a href="#S032">S032</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Transformers &gt;=4.57 loading/generation code provided; OLMo-core source tree supplied. Runtime source required to execute that example not fully inspected.<span class="refs"><a href="#S030">S030</a> · <a href="#S048">S048</a> · <a href="#S052">S052</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>OLMo-core training source tree and README inspected; actual 32B stage training scripts supplied, beyond fine-tuning examples.<span class="refs"><a href="#S048">S048</a> · <a href="#S052">S052</a> · <a href="#S068">S068</a> · <a href="#S069">S069</a> · <a href="#S070">S070</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Official staged 32B scripts/manifests listed, selected train scripts read, and stage budgets/merge method described. Exact release reconstruction needs reconciliation: card says final four checkpoints merged; official training README says three.<span class="refs"><a href="#S030">S030</a> · <a href="#S048">S048</a> · <a href="#S068">S068</a> · <a href="#S069">S069</a> · <a href="#S070">S070</a> · <a href="#S071">S071</a></span></td><td><span class="badge A" title="Public material evidenced: supplied contents or a repository listing establishes the named artifact; not a weight download or execution test.">A</span>Stage-specific Dolma/Dolmino/Longmino links, budgets and composition provided, with tokenized-data access instructions and manifest paths.<span class="refs"><a href="#S030">S030</a> · <a href="#S048">S048</a> · <a href="#S071">S071</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Reduced dataset metadata says gated=false and gives data-file patterns; tokenized-data host/manifests documented. No dataset shard listing or shard content supplied. Pretrain URL returns id allenai/dolma3_mix-6T, not requested 5.5T-1125; exact run-data identity unresolved.<span class="refs"><a href="#S030">S030</a> · <a href="#S034">S034</a> · <a href="#S035">S035</a> · <a href="#S039">S039</a> · <a href="#S071">S071</a> · <a href="#S075">S075</a> · <a href="#S076">S076</a> · <a href="#S077">S077</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Public OLMo-Eval tree, task/suite/inference commands and card results. Snapshot README is a general current harness, not a verified exact original 1125 evaluation invocation/config.<span class="refs"><a href="#S030">S030</a> · <a href="#S051">S051</a> · <a href="#S055">S055</a></span></td><td><span class="badge P" title="Partial: useful material is present, but coverage, inspection, integrity or exact model/run correspondence is incomplete. Read the cell explanation.">P</span>Apache-2.0 declared in model metadata/card; no model LICENSE file listed or license body supplied for this model. This is weaker than inspected license text. Dataset labels say ODC-BY; their terms were not supplied.<span class="refs"><a href="#S029">S029</a> · <a href="#S030">S030</a> · <a href="#S075">S075</a> · <a href="#S076">S076</a> · <a href="#S077">S077</a></span></td></tr></tbody></table></div><h2>What remains for reproduction</h2><div class="model-detail"><strong>Mistral Large 4</strong><p>First obtain the promised weights and their license, config/tokenizer and compatible runtime. Training reproduction additionally needs the corpus, preprocessing, full pre/post-training code and settings, checkpoints, and evaluation assets. An API preview cannot supply these.</p></div><div class="model-detail"><strong>DeepSeek-V4-Pro</strong><p>Inference materials are substantial, but need compatible hardware/runtime, full shard acquisition, and the custom encoder. Retraining still needs the &gt;32T corpus and processing, Muon and distributed-training configuration, domain SFT/GRPO data and rewards, consolidation/distillation recipe and teachers, run state, and exact evaluation harnesses.</p></div><div class="model-detail"><strong>Qwen3.8-27B</strong><p>Acquire all weights and multimodal preprocessing/runtime dependencies first. Retraining needs exact vision-language data, processing/mixtures, pretraining and post-training run configs, teacher/reward/agent data and environments. Score reproduction needs the corrected benchmark annotations and precise harness/judge versions; generic SFT examples do not fill these gaps.</p></div><div class="model-detail"><strong>GLM-5.3</strong><p>Need the GLM-5.2 base run/data lineage plus the actual 5.3 SFT/RL data, reward functions, orchestration and run settings. The slime link is insufficient. Evaluation needs exact modified containers/harnesses and judge access; usage must account for the custom license condition.</p></div><div class="model-detail"><strong>Kimi-K3</strong><p>Need the training corpus and its exact processing, full pretraining/post-training/QAT recipe, teachers/rewards, run state and compatible kernels. Evaluation needs Kimi Code setup, in-house tasks and the H20-calibrated task branch. Check the separate-agreement and attribution conditions before affected commercial use.</p></div><div class="model-detail"><strong>Gemma 4 31B-IT</strong><p>Need exact multimodal corpus versions, preprocessing/filtering and mixture rules, original trainer and run settings, plus instruction-tuning/teacher/reward data and code. The published loading example supports inference, not retraining. Benchmark and safety results lack a fully inspected reproduction package.</p></div><div class="model-detail"><strong>NVIDIA Nemotron 3 Super 120B-A12B-BF16</strong><p>Restore the internal-only and licensed corpus categories, full SFT/RL data and generators, precise blend/iteration settings and original run lineage. Public recipe adaptation is required for missing data. Pin containers and validate the BF16 endpoint for evaluation. License text is now hash-consistent; its conditions still apply.</p></div><div class="model-detail"><strong>Olmo-3-1125-32B</strong><p>Resolve the 5.5T-1125→6T dataset identity, enumerate/pin all stage shards and official manifests, map supplied development scripts to official release scripts, replace or obtain cloud checkpoint/data paths, and reconcile final three-vs-four checkpoint averaging. Verify optimizer states, logs and exact evaluation settings. No Think/Instruct training is needed to reproduce this base model.</p></div><h2>Final integrity results, resolved corrections and remaining limitations</h2><p>Corrected final results — 79 records: 76 exact SHA-256/byte matches; 0 unresolved mismatches; 3 original oversized retrieval failures with no contents. The four replacements preserve original line endings, including mixed endings in the NVIDIA license page. All 14 comparable GitHub text files match their tree-recorded Git blob SHA-1. Hash agreement proves consistency of supplied material, not independent authentication of its official origin.</p><p class="small">Input SHA-256: <code>edceb401799bf19058d7b7365f585cc061d4b2d180aa1a4ca3373ec588f1d9b7</code></p><p><strong>I1.</strong> Resolved by corrected S002, S007 and S008 records: original CRLF line endings are preserved, so original recorded SHA-256 and byte counts now match directly. Previous reconstruction requirement is historical only. <a href="#S002">S002</a> · <a href="#S007">S007</a> · <a href="#S008">S008</a></p><p><strong>I2.</strong> Resolved by corrected S038: the NVIDIA license HTML preserves 184 CRLF endings among 6863 total LF characters; all 300192 UTF-8 bytes and the original SHA-256 match exactly. License text is now integrity-consistent. <a href="#S038">S038</a></p><p><strong>I3.</strong> Three full dataset metadata requests exceeded the supplied collector size limit. Reduced successful replies are not full file listings; this is not evidence of missing data. <a href="#S034">S034</a> · <a href="#S035">S035</a> · <a href="#S039">S039</a> · <a href="#S075">S075</a> · <a href="#S076">S076</a> · <a href="#S077">S077</a></p><p><strong>I4.</strong> The 5.5T-1125 pretraining dataset URL returns id allenai/dolma3_mix-6T. Rename/redirect or alias is possible, but the snapshot does not establish historical run-data equivalence. <a href="#S075">S075</a></p><p><strong>I5.</strong> OLMo model card says the final checkpoint averages four checkpoints; official training README says the final three. Resolve checkpoint identities and averaging recipe before exact reconstruction. <a href="#S030">S030</a> · <a href="#S071">S071</a></p><p><strong>I6.</strong> OLMo supplied script contents come from src/scripts/train/OLMo3. The separate official release scripts and manifests are listed and linked, but their bodies are not supplied. Some inspected scripts refer to gs:// checkpoint/data paths; accessibility untested. <a href="#S048">S048</a> · <a href="#S068">S068</a> · <a href="#S069">S069</a> · <a href="#S070">S070</a> · <a href="#S071">S071</a></p><p><strong>I7.</strong> Nemotron README reports 8–10T public tokens and parenthetically ~40–50% of 25T. These figures are arithmetically inconsistent: 8–10 / 25 is 32–40%. Keep the token estimate and explicit incompleteness; do not claim a precise audited coverage fraction. <a href="#S073">S073</a></p><p><strong>I8.</strong> OLMo official README says logs are coming soon while also linking monitoring reports. Link targets were not inspected, so completeness/accessibility of original logs remains unknown. <a href="#S071">S071</a> · <a href="#S030">S030</a></p><p><strong>I9.</strong> Nemotron evaluation tutorial is model-family specific and references local/remote endpoints and credentials. It is not evidence that these exact BF16 shards have reproduced published scores. <a href="#S026">S026</a> · <a href="#S074">S074</a></p><p class="small">Corrections input SHA-256: <code>119c558b76af56aa4d27ca2e8d5cd6f6fadc0b1c4dceb28bc0e4f9b7903e4f88</code></p><h2>Source ledger and selected evidence</h2><p>Click a source reference in the table, then expand its record. Links lead to the recorded original URL; pinned raw-file links retain the captured revision. HTML excerpt line numbers refer to extracted visible text, not original HTML source lines. Listings are evidence of file presence; uninspected file bodies are not silently treated as read.</p><details id="S001"><summary>S001 · deepseek-ai--DeepSeek-V4-Pro/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro</a></p><p class="small">Retrieved 2026-10-08T08:49:50.775867+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>2084701428020e4bedf7e8fa6444199cae8cc76d83b0501d82386fd6d6f9817b</code><br>Computed UTF-8 SHA-256: <code>2084701428020e4bedf7e8fa6444199cae8cc76d83b0501d82386fd6d6f9817b</code><br>Recorded bytes 9417; embedded UTF-8 bytes 9417</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;deepseek-ai/DeepSeek-V4-Pro&quot;,
  &quot;sha&quot;: &quot;b5968e9190ef611bbf34a7229255be88a0e937c1&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;mit&quot;,
  &quot;files&quot;: [
    &quot;.gitattributes&quot;,
    &quot;LICENSE&quot;,
    &quot;README.md&quot;,
    &quot;assets/dsv4_performance.png&quot;,
    &quot;config.json&quot;,
    &quot;encoding/README.md&quot;,
    &quot;encoding/encoding_dsv4.py&quot;,
    &quot;encoding/test_encoding_dsv4.py&quot;,
    &quot;encoding/tests/test_input_1.json&quot;,
    &quot;encoding/tests/test_input_2.json&quot;,
    &quot;encoding/tests/test_input_3.json&quot;,
    &quot;encoding/tests/test_input_4.json&quot;,
    &quot;encoding/tests/test_output_1.txt&quot;,
    &quot;encoding/tests/test_output_2.txt&quot;,
    &quot;encoding/tests/test_output_3.txt&quot;,
    &quot;encoding/tests/test_output_4.txt&quot;,
    &quot;generation_config.json&quot;,
    &quot;inference/README.md&quot;,
    &quot;inference/config.json&quot;,
    &quot;inference/convert.py&quot;,
    &quot;inference/generate.py&quot;,
    &quot;inference/kernel.py&quot;,
    &quot;inference/model.py&quot;,
    &quot;inference/requirements.txt&quot;,
    &quot;model-00001-of-00064.safetensors&quot;,
    &quot;model-00002-of-00064.safetensors&quot;,
    &quot;model-00003-of-00064.safetensors&quot;,
    &quot;model-00004-of-00064.safetensors&quot;,
    &quot;model-00005-of-00064.safetensors&quot;,
    &quot;model-00006-of-00064.safetensors&quot;,
    &quot;model-00007-of-00064.safetensors&quot;,
    &quot;model-00008-of-00064.safetensors&quot;,
    &quot;model-00009-of-00064.safetensors&quot;,
    &quot;model-00010-of-00064.safetensors&quot;,
    &quot;model-00011-of-00064.safetensors&quot;,
    &quot;model-00012-of-00064.safetensors&quot;,
    &quot;model-00013-of-00064.safetensors&quot;,
    &quot;model-00014-of-00064.safetensors&quot;,
    &quot;model-00015-of-00064.safetensors&quot;,
    &quot;model-00016-of-00064.safetensors&quot;,
    &quot;model-00017-of-00064.safetensors&quot;,
    &quot;model-00018-of-00064.safetensors&quot;,
    &quot;model-00019-of-00064.safetensors&quot;,
    &quot;model-00020-of-00064.safetensors&quot;,
    &quot;model-00021-of-00064.safetensors&quot;,
    &quot;model-00022-of-00064.safetensors&quot;,
    &quot;model-00023-of-00064.safetensors&quot;,
    &quot;model-00024-of-00064.safetensors&quot;,
    &quot;model-00025-of-00064.safetensors&quot;,
    &quot;model-00026-of-00064.safetensors&quot;,
    &quot;model-00027-of-00064.safetensors&quot;,
    &quot;model-00028-of-00064.safetensors&quot;,
    &quot;model-00029-of-00064.safetensors&quot;,
    &quot;model-00030-of-00064.safetensors&quot;,
    &quot;model-00031-of-00064.safetensors&quot;,
    &quot;model-00032-of-00064.safetensors&quot;,
    &quot;model-00033-of-00064.safetensors&quot;,
    &quot;model-00034-of-00064.safetensors&quot;,
    &quot;model-00035-of-00064.safetensors&quot;,
    &quot;model-00036-of-00064.safetensors&quot;,
    &quot;model-00037-of-00064.safetensors&quot;,
    &quot;model-00038-of-00064.safetensors&quot;,
    &quot;model-00039-of-00064.safetensors&quot;,
    &quot;model-00040-of-00064.safetensors&quot;,
    &quot;model-00041-of-00064.safetensors&quot;,
    &quot;model-00042-of-00064.safetensors&quot;,
    &quot;model-00043-of-00064.safetensors&quot;,
    &quot;model-00044-of-00064.safetensors&quot;,
    &quot;model-00045-of-00064.safetensors&quot;,
    &quot;model-00046-of-00064.safetensors&quot;,
    &quot;model-00047-of-00064.safetensors&quot;,
    &quot;model-00048-of-00064.safetensors&quot;,
    &quot;model-00049-of-00064.safetensors&quot;,
    &quot;model-00050-of-00064.safetensors&quot;,
    &quot;model-00051-of-00064.safetensors&quot;,
    &quot;model-00052-of-00064.safetensors&quot;,
    &quot;model-00053-of-00064.safetensors&quot;,
    &quot;model-00054-of-00064.safetensors&quot;,
    &quot;model-00055-of-00064.safetensors&quot;,
    &quot;model-00056-of-00064.safetensors&quot;,
    &quot;model-00057-of-00064.safetensors&quot;,
    &quot;model-00058-of-00064.safetensors&quot;,
    &quot;model-00059-of-00064.safetensors&quot;,
    &quot;model-00060-of-00064.safetensors&quot;,
    &quot;model-00061-of-00064.safetensors&quot;,
    &quot;model-00062-of-00064.safetensors&quot;,
    &quot;model-00063-of-00064.safetensors&quot;,
    &quot;model-00064-of-00064.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;
  ]
}</pre></details><details id="S002"><summary>S002 · deepseek-ai--DeepSeek-V4-Pro/LICENSE · exact (corrected line endings)</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/LICENSE" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/LICENSE</a></p><p class="small">Retrieved 2026-10-08T08:49:51.285131+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>f2c6c602815669d292889e5be8c802f2ed950653b77999b1584e8e6aed25d040</code><br>Computed SHA-256: <code>f2c6c602815669d292889e5be8c802f2ed950653b77999b1584e8e6aed25d040</code><br>Recorded bytes 1084; corrected UTF-8 bytes 1084<br>Source: supplied snapshot-integrity-corrections.json. Only line endings changed.</p><p><b>content lines 1-21</b></p><pre>MIT License

Copyright (c) 2023 DeepSeek

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the &quot;Software&quot;), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED &quot;AS IS&quot;, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.</pre></details><details id="S003"><summary>S003 · deepseek-ai--DeepSeek-V4-Pro/README.md · exact</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:51.782877+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>c4d714818a4d3333542edc7d38ea065825a0cf7aa8fea3605bbd1d1c18e4a610</code><br>Computed UTF-8 SHA-256: <code>c4d714818a4d3333542edc7d38ea065825a0cf7aa8fea3605bbd1d1c18e4a610</code><br>Recorded bytes 13149; embedded UTF-8 bytes 13149</p><p><b>content lines 40-53</b></p><pre>
## Introduction

We present a preview version of **DeepSeek-V4** series, including two strong Mixture-of-Experts (MoE) language models — **DeepSeek-V4-Pro** with 1.6T parameters (49B activated) and **DeepSeek-V4-Flash** with 284B parameters (13B activated) — both supporting a context length of **one million tokens**.

DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:

1. **Hybrid Attention Architecture:** We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only **27% of single-token inference FLOPs** and **10% of KV cache** compared with DeepSeek-V3.2.
2. **Manifold-Constrained Hyper-Connections (mHC):** We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
3. **Muon Optimizer:** We employ the Muon optimizer for faster convergence and greater training stability.

We pre-train both models on more than **32T** diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model.

**DeepSeek-V4-Pro-Max**, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, **DeepSeek-V4-Flash-Max** achieves comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller parameter scale naturally places it slightly behind on pure knowledge tasks and the most complex agentic workflows.</pre><p><b>content lines 74-92</b></p><pre>## Evaluation Results

### Base Model

&lt;div align=&quot;center&quot;&gt;

| Benchmark (Metric) | # Shots | DeepSeek-V3.2-Base | DeepSeek-V4-Flash-Base | DeepSeek-V4-Pro-Base |
| :--- | :---: | :---: | :---: | :---: |
| Architecture | - | MoE | MoE | MoE |
| # Activated Params | - | 37B | 13B | 49B |
| # Total Params | - | 671B | 284B | 1.6T |
| **World Knowledge** | | | | |
| AGIEval (EM) | 0-shot | 80.1 | 82.6 | **83.1** |
| MMLU (EM) | 5-shot | 87.8 | 88.7 | **90.1** |
| MMLU-Redux (EM) | 5-shot | 87.5 | 89.4 | **90.8** |
| MMLU-Pro (EM) | 5-shot | 65.5 | 68.3 | **73.5** |
| MMMLU (EM) | 5-shot | 87.9 | 88.8 | **90.3** |
| C-Eval (EM) | 5-shot | 90.4 | 92.1 | **93.1** |
| CMMLU (EM) | 5-shot | 88.9 | 90.4 | **90.8** |</pre><p><b>content lines 194-226</b></p><pre>## Chat Template

This release does not include a Jinja-format chat template. Instead, we provide a dedicated `encoding` folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model&#x27;s text output. Please refer to the [`encoding`](encoding/README.md) folder for full documentation.

A brief example:

```python
from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;hello&quot;},
    {&quot;role&quot;: &quot;assistant&quot;, &quot;content&quot;: &quot;Hello! I am DeepSeek.&quot;, &quot;reasoning_content&quot;: &quot;thinking...&quot;},
    {&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;1+1=?&quot;}
]

# messages -&gt; string
prompt = encode_messages(messages, thinking_mode=&quot;thinking&quot;)

# string -&gt; tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained(&quot;deepseek-ai/DeepSeek-V4-Pro&quot;)
tokens = tokenizer.encode(prompt)
```

## How to Run Locally

Please refer to the [inference](inference/README.md) folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.

For local deployment, we recommend setting the sampling parameters to `temperature = 1.0, top_p = 1.0`. For the Think Max reasoning mode, we recommend setting the context window to at least **384K** tokens.

## License

This repository and the model weights are licensed under the [MIT License](LICENSE).</pre></details><details id="S004"><summary>S004 · deepseek-ai--DeepSeek-V4-Pro/config.json · exact</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:52.335843+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>5fe4568daee51c208cb8a79538eaeda090ae011ade1dee2c386aa95f569c810e</code><br>Computed UTF-8 SHA-256: <code>5fe4568daee51c208cb8a79538eaeda090ae011ade1dee2c386aa95f569c810e</code><br>Recorded bytes 1828; embedded UTF-8 bytes 1828</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;deepseek_v4&quot;,
  &quot;architectures&quot;: [
    &quot;DeepseekV4ForCausalLM&quot;
  ],
  &quot;torch_dtype&quot;: &quot;bfloat16&quot;,
  &quot;expert_dtype&quot;: &quot;fp4&quot;,
  &quot;quantization_config_summary&quot;: {
    &quot;activation_scheme&quot;: &quot;dynamic&quot;,
    &quot;fmt&quot;: &quot;e4m3&quot;,
    &quot;quant_method&quot;: &quot;fp8&quot;,
    &quot;scale_fmt&quot;: &quot;ue8m0&quot;,
    &quot;weight_block_size&quot;: [
      128,
      128
    ]
  }
}</pre></details><details id="S005"><summary>S005 · deepseek-ai--DeepSeek-V4-Pro/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:52.876525+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>6ac8c8dc065ed118161d02dd532749ae3f52c243deac27872134fae2f50d8547</code><br>Computed UTF-8 SHA-256: <code>6ac8c8dc065ed118161d02dd532749ae3f52c243deac27872134fae2f50d8547</code><br>Recorded bytes 801; embedded UTF-8 bytes 801</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;PreTrainedTokenizerFast&quot;,
  &quot;model_max_length&quot;: 1048576,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S006"><summary>S006 · Qwen--Qwen3.8-27B/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/Qwen/Qwen3.8-27B" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/Qwen/Qwen3.8-27B</a></p><p class="small">Retrieved 2026-10-08T08:49:50.776509+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>4eb5d7fdfa0914039453faf55b51207f0aab626e6847db6d50358339141838e4</code><br>Computed UTF-8 SHA-256: <code>4eb5d7fdfa0914039453faf55b51207f0aab626e6847db6d50358339141838e4</code><br>Recorded bytes 24105; embedded UTF-8 bytes 24105</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;Qwen/Qwen3.8-27B&quot;,
  &quot;sha&quot;: &quot;1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;apache-2.0&quot;,
  &quot;files&quot;: [
    &quot;.gitattributes&quot;,
    &quot;LICENSE&quot;,
    &quot;README.md&quot;,
    &quot;chat_template.jinja&quot;,
    &quot;config.json&quot;,
    &quot;crc32.txt&quot;,
    &quot;generation_config.json&quot;,
    &quot;merges.txt&quot;,
    &quot;model-00001-of-00018.safetensors&quot;,
    &quot;model-00002-of-00018.safetensors&quot;,
    &quot;model-00003-of-00018.safetensors&quot;,
    &quot;model-00004-of-00018.safetensors&quot;,
    &quot;model-00005-of-00018.safetensors&quot;,
    &quot;model-00006-of-00018.safetensors&quot;,
    &quot;model-00007-of-00018.safetensors&quot;,
    &quot;model-00008-of-00018.safetensors&quot;,
    &quot;model-00009-of-00018.safetensors&quot;,
    &quot;model-00010-of-00018.safetensors&quot;,
    &quot;model-00011-of-00018.safetensors&quot;,
    &quot;model-00012-of-00018.safetensors&quot;,
    &quot;model-00013-of-00018.safetensors&quot;,
    &quot;model-00014-of-00018.safetensors&quot;,
    &quot;model-00015-of-00018.safetensors&quot;,
    &quot;model-00016-of-00018.safetensors&quot;,
    &quot;model-00017-of-00018.safetensors&quot;,
    &quot;model-00018-of-00018.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;preprocessor_config.json&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;,
    &quot;video_preprocessor_config.json&quot;,
    &quot;vocab.json&quot;
  ]
}</pre></details><details id="S007"><summary>S007 · Qwen--Qwen3.8-27B/LICENSE · exact (corrected line endings)</summary><p><a href="https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE" target="_blank" rel="noopener noreferrer">https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE</a></p><p class="small">Retrieved 2026-10-08T08:49:51.336782+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a</code><br>Computed SHA-256: <code>bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a</code><br>Recorded bytes 11544; corrected UTF-8 bytes 11544<br>Source: supplied snapshot-integrity-corrections.json. Only line endings changed.</p><p><b>content lines 1-5</b></p><pre>
                                 Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/
</pre><p><b>content lines 67-101</b></p><pre>   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.

   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.

   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and

      (c) You must retain, in the Source form of any Derivative Works</pre></details><details id="S008"><summary>S008 · Qwen--Qwen3.8-27B/README.md · exact (corrected line endings)</summary><p><a href="https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:51.842636+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>57e4bdb258ee1a7d2635c5174ebd4e56abe392505cdb5f8bbb356b0dc4293641</code><br>Computed SHA-256: <code>57e4bdb258ee1a7d2635c5174ebd4e56abe392505cdb5f8bbb356b0dc4293641</code><br>Recorded bytes 65012; corrected UTF-8 bytes 65012<br>Source: supplied snapshot-integrity-corrections.json. Only line endings changed.</p><p><b>content lines 10-53</b></p><pre>&gt; This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. 
&gt;
&gt; These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

&gt; [!Tip]
&gt; For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by [Qwen Cloud](https://www.qwencloud.com).
&gt; In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the [Qwen3.8-27B Overview](https://www.qwencloud.com/models/qwen3.8-27b). The service is coming soon. Stay tuned for updates.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

## Qwen3.8 Highlights

Qwen3.8-27B features the following enhancements:
- **Core Capabilities**: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
- **Agent Execution**: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
- **Downstream Compatibility**: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
- **Flexible Thinking Control**: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with `reasoning_effort`, and reasoning context from historical messages is retained via `preserve_thinking`.
- **Vision-Language Understanding**: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.


## Model Overview

- Type: Causal Language Model with Vision Encoder
- Training Stage: Pre-training &amp; Post-training
- Language Model
    - Number of Parameters: 27B
    - Hidden Dimension: 5120
    - Token Embedding: 248,320 (Padded)
    - Number of Layers: 64
    - Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
    - Gated DeltaNet:
        - Number of Linear Attention Heads: 48 for V and 16 for QK
        - Head Dimension: 128
    - Gated Attention:
        - Number of Attention Heads: 24 for Q and 4 for KV
        - Head Dimension: 256
        - Rotary Position Embedding Dimension: 64
    - Feed Forward Network:
        - Intermediate Dimension: 17,408
    - LM Output: 248,320 (Padded)
    - MTP (Multi-Token Prediction): trained with multiple steps
- Context Length: 262,144 natively and extensible up to 1,000,000 tokens.</pre><p><b>content lines 211-240</b></p><pre>&lt;div style=&quot;margin-top:12px;font-size:11px;line-height:1.5;color:rgba(0,0,0,0.72)&quot;&gt;
&lt;ol style=&quot;margin:0;padding-left:20px&quot;&gt;
&lt;li&gt;MathVision, BabyVision, and CharXiv (RQ): Where both settings are available, cells report “Without CI” and “With CI” separately; otherwise, only the available setting is shown. A small number of incorrect ground-truth annotations in MathVision and CharXiv (RQ) were corrected following manual verification, and all reported scores on those benchmarks were computed using the corrected annotations.&lt;/li&gt;
&lt;li&gt;MathVision: Qwen3.8-27B is evaluated using the fixed prompt: “Please reason step by step, and put your final answer within &lt;code&gt;\boxed{}&lt;/code&gt;.” For the remaining models, we report the higher score from two prompt variants—one with and one without the &lt;code&gt;\boxed{}&lt;/code&gt; formatting requirement.&lt;/li&gt;
&lt;li&gt;WebArena-Verified: Scores are computed with the official WebArena-Verified grader under the OSWorld scaffold.&lt;/li&gt;
&lt;li&gt;RecreationBench: An in-house, long-horizon application-recreation benchmark designed to evaluate hybrid-agent capabilities across five platforms: desktop (Ubuntu, macOS, and Windows), mobile (Android), and the web.&lt;/li&gt;
&lt;li&gt;ClawEval-MM: Scores are reported as “Pass@3 / average score.” Pass@3 is the percentage of tasks passed in at least one of three trials; the average score is the mean benchmark score across the three trials.&lt;/li&gt;
&lt;li&gt;Vision2Web: Scores are averaged across the frontend, webpage, and website categories. Evaluations use the Claude Code harness and are judged by &lt;code&gt;gpt-5.4-2026-03-05&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;SWE-MM: Scores are evaluated on the Claude Code harness using the public dev split of SWE-bench Multimodal, with the modifications described in Appendix 8.3 of the Claude Opus 4.7 system card.&lt;/li&gt;
&lt;li&gt;Empty cells (--) indicate that results are not yet available or not applicable.&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;/div&gt;


## Quickstart

For streamlined integration, we recommend using Qwen3.8 via APIs.

### Serving Qwen3.8

&gt; [!Important]
&gt; Inference efficiency and throughput vary significantly across frameworks. 
&gt; We recommend using the latest framework versions to ensure optimal performance and compatibility.
&gt; For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, vLLM, or TokenSpeed are recommended.

Qwen3.8 can be deployed with popular inference frameworks, e.g.:

- [SGLang](https://www.sglang.io/): [Qwen3.8 Cookbook](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B)
- [vLLM](https://vllm.ai/): [Qwen3.8 Recipe](https://recipes.vllm.ai/Qwen/Qwen3.8-27B)
- [TokenSpeed](https://lightseek.org/tokenspeed/): [Qwen3.8 Recipe](https://lightseek.org/tokenspeed/recipes/models#qwen3-8)</pre></details><details id="S009"><summary>S009 · Qwen--Qwen3.8-27B/config.json · exact</summary><p><a href="https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:52.559947+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab</code><br>Computed UTF-8 SHA-256: <code>191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab</code><br>Recorded bytes 4312; embedded UTF-8 bytes 4312</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;qwen3_5&quot;,
  &quot;architectures&quot;: [
    &quot;Qwen3_5ForConditionalGeneration&quot;
  ],
  &quot;quantization_config_summary&quot;: {}
}</pre></details><details id="S010"><summary>S010 · Qwen--Qwen3.8-27B/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:53.031738+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>b11349aafa7cdc6a320767cf7ceb29ed82f7eda5d65e8e0819e76f0ce947bf27</code><br>Computed UTF-8 SHA-256: <code>b11349aafa7cdc6a320767cf7ceb29ed82f7eda5d65e8e0819e76f0ce947bf27</code><br>Recorded bytes 17928; embedded UTF-8 bytes 17928</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;Qwen2Tokenizer&quot;,
  &quot;model_max_length&quot;: 262144,
  &quot;has_embedded_chat_template&quot;: true
}</pre></details><details id="S011"><summary>S011 · zai-org--GLM-5.3/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/zai-org/GLM-5.3" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/zai-org/GLM-5.3</a></p><p class="small">Retrieved 2026-10-08T08:49:50.776654+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>510bf9c80ad32b574a132f99cefa10bd29b7d660e432b8468b9cfaefc43cdb66</code><br>Computed UTF-8 SHA-256: <code>510bf9c80ad32b574a132f99cefa10bd29b7d660e432b8468b9cfaefc43cdb66</code><br>Recorded bytes 42446; embedded UTF-8 bytes 42446</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;zai-org/GLM-5.3&quot;,
  &quot;sha&quot;: &quot;aca966e4e02791568aa6a4ced368624b3d897f42&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;other&quot;,
  &quot;files&quot;: [
    &quot;.eval_results/deep-swe.yaml&quot;,
    &quot;.eval_results/hle.yaml&quot;,
    &quot;.eval_results/terminal-bench-2.1.yaml&quot;,
    &quot;.eval_results/terminal-bench-3.0.yaml&quot;,
    &quot;.eval_results/zai-org__GLM-5.3.yaml&quot;,
    &quot;.gitattributes&quot;,
    &quot;LICENSE&quot;,
    &quot;README.md&quot;,
    &quot;chat_template.jinja&quot;,
    &quot;config.json&quot;,
    &quot;generation_config.json&quot;,
    &quot;model-00001-of-00141.safetensors&quot;,
    &quot;model-00002-of-00141.safetensors&quot;,
    &quot;model-00003-of-00141.safetensors&quot;,
    &quot;model-00004-of-00141.safetensors&quot;,
    &quot;model-00005-of-00141.safetensors&quot;,
    &quot;model-00006-of-00141.safetensors&quot;,
    &quot;model-00007-of-00141.safetensors&quot;,
    &quot;model-00008-of-00141.safetensors&quot;,
    &quot;model-00009-of-00141.safetensors&quot;,
    &quot;model-00010-of-00141.safetensors&quot;,
    &quot;model-00011-of-00141.safetensors&quot;,
    &quot;model-00012-of-00141.safetensors&quot;,
    &quot;model-00013-of-00141.safetensors&quot;,
    &quot;model-00014-of-00141.safetensors&quot;,
    &quot;model-00015-of-00141.safetensors&quot;,
    &quot;model-00016-of-00141.safetensors&quot;,
    &quot;model-00017-of-00141.safetensors&quot;,
    &quot;model-00018-of-00141.safetensors&quot;,
    &quot;model-00019-of-00141.safetensors&quot;,
    &quot;model-00020-of-00141.safetensors&quot;,
    &quot;model-00021-of-00141.safetensors&quot;,
    &quot;model-00022-of-00141.safetensors&quot;,
    &quot;model-00023-of-00141.safetensors&quot;,
    &quot;model-00024-of-00141.safetensors&quot;,
    &quot;model-00025-of-00141.safetensors&quot;,
    &quot;model-00026-of-00141.safetensors&quot;,
    &quot;model-00027-of-00141.safetensors&quot;,
    &quot;model-00028-of-00141.safetensors&quot;,
    &quot;model-00029-of-00141.safetensors&quot;,
    &quot;model-00030-of-00141.safetensors&quot;,
    &quot;model-00031-of-00141.safetensors&quot;,
    &quot;model-00032-of-00141.safetensors&quot;,
    &quot;model-00033-of-00141.safetensors&quot;,
    &quot;model-00034-of-00141.safetensors&quot;,
    &quot;model-00035-of-00141.safetensors&quot;,
    &quot;model-00036-of-00141.safetensors&quot;,
    &quot;model-00037-of-00141.safetensors&quot;,
    &quot;model-00038-of-00141.safetensors&quot;,
    &quot;model-00039-of-00141.safetensors&quot;,
    &quot;model-00040-of-00141.safetensors&quot;,
    &quot;model-00041-of-00141.safetensors&quot;,
    &quot;model-00042-of-00141.safetensors&quot;,
    &quot;model-00043-of-00141.safetensors&quot;,
    &quot;model-00044-of-00141.safetensors&quot;,
    &quot;model-00045-of-00141.safetensors&quot;,
    &quot;model-00046-of-00141.safetensors&quot;,
    &quot;model-00047-of-00141.safetensors&quot;,
    &quot;model-00048-of-00141.safetensors&quot;,
    &quot;model-00049-of-00141.safetensors&quot;,
    &quot;model-00050-of-00141.safetensors&quot;,
    &quot;model-00051-of-00141.safetensors&quot;,
    &quot;model-00052-of-00141.safetensors&quot;,
    &quot;model-00053-of-00141.safetensors&quot;,
    &quot;model-00054-of-00141.safetensors&quot;,
    &quot;model-00055-of-00141.safetensors&quot;,
    &quot;model-00056-of-00141.safetensors&quot;,
    &quot;model-00057-of-00141.safetensors&quot;,
    &quot;model-00058-of-00141.safetensors&quot;,
    &quot;model-00059-of-00141.safetensors&quot;,
    &quot;model-00060-of-00141.safetensors&quot;,
    &quot;model-00061-of-00141.safetensors&quot;,
    &quot;model-00062-of-00141.safetensors&quot;,
    &quot;model-00063-of-00141.safetensors&quot;,
    &quot;model-00064-of-00141.safetensors&quot;,
    &quot;model-00065-of-00141.safetensors&quot;,
    &quot;model-00066-of-00141.safetensors&quot;,
    &quot;model-00067-of-00141.safetensors&quot;,
    &quot;model-00068-of-00141.safetensors&quot;,
    &quot;model-00069-of-00141.safetensors&quot;,
    &quot;model-00070-of-00141.safetensors&quot;,
    &quot;model-00071-of-00141.safetensors&quot;,
    &quot;model-00072-of-00141.safetensors&quot;,
    &quot;model-00073-of-00141.safetensors&quot;,
    &quot;model-00074-of-00141.safetensors&quot;,
    &quot;model-00075-of-00141.safetensors&quot;,
    &quot;model-00076-of-00141.safetensors&quot;,
    &quot;model-00077-of-00141.safetensors&quot;,
    &quot;model-00078-of-00141.safetensors&quot;,
    &quot;model-00079-of-00141.safetensors&quot;,
    &quot;model-00080-of-00141.safetensors&quot;,
    &quot;model-00081-of-00141.safetensors&quot;,
    &quot;model-00082-of-00141.safetensors&quot;,
    &quot;model-00083-of-00141.safetensors&quot;,
    &quot;model-00084-of-00141.safetensors&quot;,
    &quot;model-00085-of-00141.safetensors&quot;,
    &quot;model-00086-of-00141.safetensors&quot;,
    &quot;model-00087-of-00141.safetensors&quot;,
    &quot;model-00088-of-00141.safetensors&quot;,
    &quot;model-00089-of-00141.safetensors&quot;,
    &quot;model-00090-of-00141.safetensors&quot;,
    &quot;model-00091-of-00141.safetensors&quot;,
    &quot;model-00092-of-00141.safetensors&quot;,
    &quot;model-00093-of-00141.safetensors&quot;,
    &quot;model-00094-of-00141.safetensors&quot;,
    &quot;model-00095-of-00141.safetensors&quot;,
    &quot;model-00096-of-00141.safetensors&quot;,
    &quot;model-00097-of-00141.safetensors&quot;,
    &quot;model-00098-of-00141.safetensors&quot;,
    &quot;model-00099-of-00141.safetensors&quot;,
    &quot;model-00100-of-00141.safetensors&quot;,
    &quot;model-00101-of-00141.safetensors&quot;,
    &quot;model-00102-of-00141.safetensors&quot;,
    &quot;model-00103-of-00141.safetensors&quot;,
    &quot;model-00104-of-00141.safetensors&quot;,
    &quot;model-00105-of-00141.safetensors&quot;,
    &quot;model-00106-of-00141.safetensors&quot;,
    &quot;model-00107-of-00141.safetensors&quot;,
    &quot;model-00108-of-00141.safetensors&quot;,
    &quot;model-00109-of-00141.safetensors&quot;,
    &quot;model-00110-of-00141.safetensors&quot;,
    &quot;model-00111-of-00141.safetensors&quot;,
    &quot;model-00112-of-00141.safetensors&quot;,
    &quot;model-00113-of-00141.safetensors&quot;,
    &quot;model-00114-of-00141.safetensors&quot;,
    &quot;model-00115-of-00141.safetensors&quot;,
    &quot;model-00116-of-00141.safetensors&quot;,
    &quot;model-00117-of-00141.safetensors&quot;,
    &quot;model-00118-of-00141.safetensors&quot;,
    &quot;model-00119-of-00141.safetensors&quot;,
    &quot;model-00120-of-00141.safetensors&quot;,
    &quot;model-00121-of-00141.safetensors&quot;,
    &quot;model-00122-of-00141.safetensors&quot;,
    &quot;model-00123-of-00141.safetensors&quot;,
    &quot;model-00124-of-00141.safetensors&quot;,
    &quot;model-00125-of-00141.safetensors&quot;,
    &quot;model-00126-of-00141.safetensors&quot;,
    &quot;model-00127-of-00141.safetensors&quot;,
    &quot;model-00128-of-00141.safetensors&quot;,
    &quot;model-00129-of-00141.safetensors&quot;,
    &quot;model-00130-of-00141.safetensors&quot;,
    &quot;model-00131-of-00141.safetensors&quot;,
    &quot;model-00132-of-00141.safetensors&quot;,
    &quot;model-00133-of-00141.safetensors&quot;,
    &quot;model-00134-of-00141.safetensors&quot;,
    &quot;model-00135-of-00141.safetensors&quot;,
    &quot;model-00136-of-00141.safetensors&quot;,
    &quot;model-00137-of-00141.safetensors&quot;,
    &quot;model-00138-of-00141.safetensors&quot;,
    &quot;model-00139-of-00141.safetensors&quot;,
    &quot;model-00140-of-00141.safetensors&quot;,
    &quot;model-00141-of-00141.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;
  ]
}</pre></details><details id="S012"><summary>S012 · zai-org--GLM-5.3/LICENSE · exact</summary><p><a href="https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/LICENSE" target="_blank" rel="noopener noreferrer">https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/LICENSE</a></p><p class="small">Retrieved 2026-10-08T08:49:51.458686+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>96e1622099fc9d6b70c9760f007d99e66d7497eec636b63c60fe208401e9170c</code><br>Computed UTF-8 SHA-256: <code>96e1622099fc9d6b70c9760f007d99e66d7497eec636b63c60fe208401e9170c</code><br>Recorded bytes 4263; embedded UTF-8 bytes 4263</p><p><b>content lines 1-14</b></p><pre>GLM-5.3 License

Copyright (c) 2026 Z.AI

Permission is hereby granted, free of charge, to any person or entity (the &quot;Licensee&quot;) obtaining a copy of this software — including the model weights, parameters, configuration files, inference and training code, and associated documentation (collectively, the &quot;Software&quot;) — to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software; to run, deploy, fine-tune, or otherwise modify the Software and create derivative works from it; and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

1. The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. The Licensee&#x27;s use of the Software must comply with applicable laws and regulations.

2. &quot;Model as a Service&quot; means giving a third party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data. This does not include (a) end-user products with model capabilities solely embedded within specific features or harnesses, or (b) mere relaying of requests to models hosted by others.
If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must pass Z.AI&#x27;s security review before using the Software or its derivative works for any commercial purpose. The scope and method of the security review shall be reasonably determined by Z.AI.

3. THE SOFTWARE AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN &quot;AS IS&quot; BASIS, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL Z.AI OR ITS AFFILIATES OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

For any questions regarding this license, please contact glmlicense@z.ai.</pre></details><details id="S013"><summary>S013 · zai-org--GLM-5.3/README.md · exact</summary><p><a href="https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:51.850479+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>ed1c0a4563c437a32f8953d637a1db9bd831de1b24bd3062a5ace9e04340b1cf</code><br>Computed UTF-8 SHA-256: <code>ed1c0a4563c437a32f8953d637a1db9bd831de1b24bd3062a5ace9e04340b1cf</code><br>Recorded bytes 14209; embedded UTF-8 bytes 14209</p><p><b>content lines 11-17</b></p><pre># GLM-5.3

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

+ Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents&#x27; Last Exam.
+ Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
</pre><p><b>content lines 41-74</b></p><pre>### Serve GLM-5.3 Locally

GLM-5.3 supports deployment with the following frameworks. Feel free to try them out:

- [SGLang](https://github.com/sgl-project/sglang) — see [cookbook](https://cookbook.sglang.io/autoregressive/GLM/GLM-5.3)
- [vLLM](https://github.com/vllm-project/vllm) — see [recipes](https://recipes.vllm.ai/zai-org/GLM-5.3)
- [TokenSpeed](https://github.com/lightseekorg/tokenspeed) — see [here](https://lightseek.org/tokenspeed/recipes/models#glm-5-3)
- [Transformers](https://github.com/huggingface/transformers) — see [transformers docs](https://github.com/huggingface/transformers/blob/main/docs/source/en/model_doc/glm_moe_dsa.md)
- [KTransformers](https://github.com/kvcache-ai/ktransformers) — see [tutorial](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/GLM-5.2-Tutorial.md)
- [Unsloth](https://github.com/unslothai/unsloth) — see [guide](https://unsloth.ai/docs/models/GLM-5.3)
- For deployment on the `Ascend NPU` platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see [here](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md).

### Note

- GLM-5.3 supports controlling the thinking budget through the `reasoning_effort` parameter, which accepts three levels: `low`, `high`, and `max`. It defaults to `max` if not passed (or if set to any other value). To use `low` or `high`, pass them explicitly. For benchmark and leaderboard reproduction, keep the default `max`.
- In the chat template for GLM-5.3, `clear_thinking` defaults to `false` if not passed. For chat scenarios, explicitly pass `clear_thinking=true`.

## Footnotes

- **HLE w/ tools**: We use sampling parameters of `temperature=1.0` and `top_p=0.95` for evaluation, with a maximum generation length of `163,840` tokens. The evaluation is conducted with a maximum context length of `300,000` tokens, using a context management strategy. We use GPT-5.6-luna (medium) as the judge model.
- **NL2Repo**: We evaluated NL2Repo with `temperature=1.0`, `top_p=1.0`, and `max_new_tokens=64k` under 1M context. To prevent hacking, we use rule-based and a LLM-based judgement to prevent malicious behaviors (e.g., unauthorized pip or curl operations).
- **DeepSWE**: We run DeepSWE using the mini-swe-agent harness with `temperature=0.95`, `top_p=1.0`, `timeout=6h` and 400K context.
- **Terminal-Bench 2.1**: We evaluate in Claude Code 2.1.207 with `temperature=1.0`, `top_p=1`, `max_new_tokens=65536` with 6h timeout.
- **Terminal-Bench 3.0**: We evaluate Terminal-Bench-3 tasks with the Claude Code 2.1.207 harness (reasoning effort=max, 400K context, and 128K maximum output), reporting avg@3 over three rollouts per task. Each rollout runs in an isolated container built from the task&#x27;s official image, and is capped at 600 agent turns with a 10-hour timeout. Tool Search is disabled, and the artifacts each agent produces are scored by the task&#x27;s official separate verifier.
- **Agent&#x27;s Last Exam (CLI)**: We evaluate ALE using the official evaluation protocol with the Claude Code harness (reasoning effort=max, 1M context, and 64K maximum output). Each of the 105 tasks runs in an isolated Docker container using the resources declared in its Task Card. The default timeout is 4 hours, with task-specific limits taking precedence (up to 8 hours). Tool Search is disabled, and results are scored by the official ALE evaluators.
- **Toolathlon Verified**: We obtain all results via the official evaluation service and report pass@1 averaged over 3 independent runs.
- **AutomationBench**: We evaluate on AutomationBench **v1.0.6**, incorporating the fix for the `null`-type handling issue introduced in [PR #13]([#](https://github.com/zapier/AutomationBench/pull/13)).
- **GDPval-AA v2**: Models are evaluated by Artificial Analysis.
- **CyberGym**: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with `temperature=1.0`, `top_p=1.0`, `max_new_tokens=128000`). All evaluations are under unlimited timeout per task and results are single-run Pass@1 over 1,507 tasks. To simulate real-world usage scenarios, we place the agent inside the task container. We also remove all Git-related information and apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- **ExploitGym**: We evaluate GLM-5.3, Kimi-K3 and Qwen3.8 Max in Claude Code 2.1.207 (max reasoning effort, no web tools with `temperature=1.0`, `top_p=1.0`, `max_new_tokens=128000`). The reported results are single-run Pass@1 on 869 tasks under two timeout budgets: 2 hours and 6 hours, which are calculated as the API inference time rescaled by per-model tokens per second rate (per-model TPS sourced from Artificial Analysis; that is, we rescale GLM-5.3&#x27;s results by 115 TPS, Kimi K3&#x27;s results by 40 TPS and Qwen3.8 Max&#x27;s results by 47 TPS), plus the non-API overhead. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- **ExploitBench**: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with `temperature=1.0`, `top_p=1.0`, `max_new_tokens=128000`). Following the official evaluation settings, we limit the maximum number of interaction rounds between the agent and the environment to 300, and compute the average coverage score over all 41 tasks across 3 revisions. The coverage result of a task is determined by taking the union of capabilities achieved across all revisions, and the average score is obtained by averaging the results. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- **FrontierSWE**: The evaluation was conducted by [Proximal](https://www.proximal.ai/) with 1M context length, max effort level, and 128K maximum output tokens. Dominance score reported as of 2026/08/14.
- **PostTrainBench**: We evaluate GLM-5.3 using Claude Code 2.1.207 with max effort level, `temperature = 1.0`, `top_p = 1.0`, `max_new_tokens = 128000`, and a 1M-token context window. We report the weighted average over 3 runs. Runs that fail to produce a score fall back to the official zero-shot base-model baseline score. For checks intended to prevent the use of third-party APIs, we removed the original pattern-matching-based checks, as they produced false positives when a local vLLM endpoint was accessed through the OpenAI SDK. Instead, we use an LLM agent to inspect solutions for external API usage.
- **SWE-Marathon**: We evaluate GLM-5.3 using Claude Code 2.1.207 with maximum effort level, `temperature = 1.0`, `top_p = 0.95`, `max_new_tokens = 128000`, and a 1M-token context window. For `strip-clone`, the original anti-cheat checks used overly broad import detection that could reject valid implementations. We removed the affected checks and performed llm-based inspection instead to avoid false positives. For `parameter-golf` and `trimul-cuda`, changes to the NVIDIA wheels caused the Docker image builds to fail, so we added `--extra-index-url https://pypi.org/simple` to restore successful builds.</pre></details><details id="S014"><summary>S014 · zai-org--GLM-5.3/config.json · exact</summary><p><a href="https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:52.373165+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>3ac72612095574542f7fff847ada8e59d9199dd8af44bdf625d7e02615572e69</code><br>Computed UTF-8 SHA-256: <code>3ac72612095574542f7fff847ada8e59d9199dd8af44bdf625d7e02615572e69</code><br>Recorded bytes 29464; embedded UTF-8 bytes 29464</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;glm_moe_dsa&quot;,
  &quot;architectures&quot;: [
    &quot;GlmMoeDsaForCausalLM&quot;
  ],
  &quot;dtype&quot;: &quot;bfloat16&quot;,
  &quot;quantization_config_summary&quot;: {
    &quot;activation_scheme&quot;: &quot;dynamic&quot;,
    &quot;fmt&quot;: &quot;e4m3&quot;,
    &quot;quant_method&quot;: &quot;fp8&quot;,
    &quot;weight_block_size&quot;: [
      128,
      128
    ]
  }
}</pre></details><details id="S015"><summary>S015 · zai-org--GLM-5.3/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:53.090531+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>98b1271574f41abf89427ae2dda030d94dc9478f0edc5a8bd240db213c6fd5fc</code><br>Computed UTF-8 SHA-256: <code>98b1271574f41abf89427ae2dda030d94dc9478f0edc5a8bd240db213c6fd5fc</code><br>Recorded bytes 761; embedded UTF-8 bytes 761</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;TokenizersBackend&quot;,
  &quot;model_max_length&quot;: 1048576,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S016"><summary>S016 · moonshotai--Kimi-K3/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/moonshotai/Kimi-K3" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/moonshotai/Kimi-K3</a></p><p class="small">Retrieved 2026-10-08T08:49:50.776754+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>f80cc04a192ef89955c118054307e9ef5314731e3b38acfeda92b192a1258dc3</code><br>Computed UTF-8 SHA-256: <code>f80cc04a192ef89955c118054307e9ef5314731e3b38acfeda92b192a1258dc3</code><br>Recorded bytes 9840; embedded UTF-8 bytes 9840</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;moonshotai/Kimi-K3&quot;,
  &quot;sha&quot;: &quot;f831ab66814297da540d832a5235f8e904f29d06&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;other&quot;,
  &quot;files&quot;: [
    &quot;.eval_results/apex-agents.yaml&quot;,
    &quot;.eval_results/deep-swe.yaml&quot;,
    &quot;.eval_results/gpqa.yaml&quot;,
    &quot;.eval_results/hle.yaml&quot;,
    &quot;.eval_results/moonshotai__Kimi-K3.yaml&quot;,
    &quot;.gitattributes&quot;,
    &quot;LICENSE&quot;,
    &quot;README.md&quot;,
    &quot;assets/kimi-logo.png&quot;,
    &quot;config.json&quot;,
    &quot;configuration_kimi_k3.py&quot;,
    &quot;encoding_k3.py&quot;,
    &quot;generation_config.json&quot;,
    &quot;kimi_k3_processor.py&quot;,
    &quot;kimi_k3_vision_processing.py&quot;,
    &quot;media_utils.py&quot;,
    &quot;model-00001-of-000096.safetensors&quot;,
    &quot;model-00002-of-000096.safetensors&quot;,
    &quot;model-00003-of-000096.safetensors&quot;,
    &quot;model-00004-of-000096.safetensors&quot;,
    &quot;model-00005-of-000096.safetensors&quot;,
    &quot;model-00006-of-000096.safetensors&quot;,
    &quot;model-00007-of-000096.safetensors&quot;,
    &quot;model-00008-of-000096.safetensors&quot;,
    &quot;model-00009-of-000096.safetensors&quot;,
    &quot;model-00010-of-000096.safetensors&quot;,
    &quot;model-00011-of-000096.safetensors&quot;,
    &quot;model-00012-of-000096.safetensors&quot;,
    &quot;model-00013-of-000096.safetensors&quot;,
    &quot;model-00014-of-000096.safetensors&quot;,
    &quot;model-00015-of-000096.safetensors&quot;,
    &quot;model-00016-of-000096.safetensors&quot;,
    &quot;model-00017-of-000096.safetensors&quot;,
    &quot;model-00018-of-000096.safetensors&quot;,
    &quot;model-00019-of-000096.safetensors&quot;,
    &quot;model-00020-of-000096.safetensors&quot;,
    &quot;model-00021-of-000096.safetensors&quot;,
    &quot;model-00022-of-000096.safetensors&quot;,
    &quot;model-00023-of-000096.safetensors&quot;,
    &quot;model-00024-of-000096.safetensors&quot;,
    &quot;model-00025-of-000096.safetensors&quot;,
    &quot;model-00026-of-000096.safetensors&quot;,
    &quot;model-00027-of-000096.safetensors&quot;,
    &quot;model-00028-of-000096.safetensors&quot;,
    &quot;model-00029-of-000096.safetensors&quot;,
    &quot;model-00030-of-000096.safetensors&quot;,
    &quot;model-00031-of-000096.safetensors&quot;,
    &quot;model-00032-of-000096.safetensors&quot;,
    &quot;model-00033-of-000096.safetensors&quot;,
    &quot;model-00034-of-000096.safetensors&quot;,
    &quot;model-00035-of-000096.safetensors&quot;,
    &quot;model-00036-of-000096.safetensors&quot;,
    &quot;model-00037-of-000096.safetensors&quot;,
    &quot;model-00038-of-000096.safetensors&quot;,
    &quot;model-00039-of-000096.safetensors&quot;,
    &quot;model-00040-of-000096.safetensors&quot;,
    &quot;model-00041-of-000096.safetensors&quot;,
    &quot;model-00042-of-000096.safetensors&quot;,
    &quot;model-00043-of-000096.safetensors&quot;,
    &quot;model-00044-of-000096.safetensors&quot;,
    &quot;model-00045-of-000096.safetensors&quot;,
    &quot;model-00046-of-000096.safetensors&quot;,
    &quot;model-00047-of-000096.safetensors&quot;,
    &quot;model-00048-of-000096.safetensors&quot;,
    &quot;model-00049-of-000096.safetensors&quot;,
    &quot;model-00050-of-000096.safetensors&quot;,
    &quot;model-00051-of-000096.safetensors&quot;,
    &quot;model-00052-of-000096.safetensors&quot;,
    &quot;model-00053-of-000096.safetensors&quot;,
    &quot;model-00054-of-000096.safetensors&quot;,
    &quot;model-00055-of-000096.safetensors&quot;,
    &quot;model-00056-of-000096.safetensors&quot;,
    &quot;model-00057-of-000096.safetensors&quot;,
    &quot;model-00058-of-000096.safetensors&quot;,
    &quot;model-00059-of-000096.safetensors&quot;,
    &quot;model-00060-of-000096.safetensors&quot;,
    &quot;model-00061-of-000096.safetensors&quot;,
    &quot;model-00062-of-000096.safetensors&quot;,
    &quot;model-00063-of-000096.safetensors&quot;,
    &quot;model-00064-of-000096.safetensors&quot;,
    &quot;model-00065-of-000096.safetensors&quot;,
    &quot;model-00066-of-000096.safetensors&quot;,
    &quot;model-00067-of-000096.safetensors&quot;,
    &quot;model-00068-of-000096.safetensors&quot;,
    &quot;model-00069-of-000096.safetensors&quot;,
    &quot;model-00070-of-000096.safetensors&quot;,
    &quot;model-00071-of-000096.safetensors&quot;,
    &quot;model-00072-of-000096.safetensors&quot;,
    &quot;model-00073-of-000096.safetensors&quot;,
    &quot;model-00074-of-000096.safetensors&quot;,
    &quot;model-00075-of-000096.safetensors&quot;,
    &quot;model-00076-of-000096.safetensors&quot;,
    &quot;model-00077-of-000096.safetensors&quot;,
    &quot;model-00078-of-000096.safetensors&quot;,
    &quot;model-00079-of-000096.safetensors&quot;,
    &quot;model-00080-of-000096.safetensors&quot;,
    &quot;model-00081-of-000096.safetensors&quot;,
    &quot;model-00082-of-000096.safetensors&quot;,
    &quot;model-00083-of-000096.safetensors&quot;,
    &quot;model-00084-of-000096.safetensors&quot;,
    &quot;model-00085-of-000096.safetensors&quot;,
    &quot;model-00086-of-000096.safetensors&quot;,
    &quot;model-00087-of-000096.safetensors&quot;,
    &quot;model-00088-of-000096.safetensors&quot;,
    &quot;model-00089-of-000096.safetensors&quot;,
    &quot;model-00090-of-000096.safetensors&quot;,
    &quot;model-00091-of-000096.safetensors&quot;,
    &quot;model-00092-of-000096.safetensors&quot;,
    &quot;model-00093-of-000096.safetensors&quot;,
    &quot;model-00094-of-000096.safetensors&quot;,
    &quot;model-00095-of-000096.safetensors&quot;,
    &quot;model-00096-of-000096.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;modeling_kimi_k3.py&quot;,
    &quot;modeling_kimi_linear.py&quot;,
    &quot;preprocessor_config.json&quot;,
    &quot;tiktoken.model&quot;,
    &quot;tokenization_kimi.py&quot;,
    &quot;tokenizer_config.json&quot;
  ]
}</pre></details><details id="S017"><summary>S017 · moonshotai--Kimi-K3/LICENSE · exact</summary><p><a href="https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/LICENSE" target="_blank" rel="noopener noreferrer">https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/LICENSE</a></p><p class="small">Retrieved 2026-10-08T08:49:51.313628+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887</code><br>Computed UTF-8 SHA-256: <code>20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887</code><br>Recorded bytes 3065; embedded UTF-8 bytes 3065</p><p><b>content lines 1-52</b></p><pre>Kimi K3 License

Copyright (c) 2026 Moonshot AI

Permission is hereby granted, free of charge, to any person (the &quot;Licensee&quot;)
obtaining a copy of this software — including the model weights, parameters,
configuration files, inference and training code, and associated documentation
(collectively, the &quot;Software&quot;) — to deal in the Software without restriction.
This includes, without limitation, the rights to use, copy, modify, merge,
publish, distribute, sublicense, and/or sell copies of the Software; to run,
deploy, fine-tune, or otherwise modify the Software and create derivative works
from it; and to permit persons to whom the Software is furnished to do so, in
each case subject to the following conditions:

1. The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software. Licensee&#x27;s use of the
Software must comply with applicable laws and regulations.

2. &quot;Model as a Service&quot; means giving a third party access to language model
inference or fine-tuning (e.g., via API) in a manner that allows such third
party to exercise meaningful control over the inputs, parameters, or training
data. This does not include (a) end-user products with model capabilities solely
embedded within specific features or harnesses, or (b) mere relaying of requests
to models hosted by others.

If the Licensee or any of its affiliates operates a Model as a Service business,
and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
US dollars (or the equivalent in other currencies) in total over any consecutive
12 months, the Licensee must enter into a separate agreement with Moonshot AI
before using the Software or its derivative works for any commercial purpose. 

3. If the Software (or any derivative works thereof) is used for any of the
Licensee&#x27;s commercial products or services that have more than 100 million
monthly active users, or more than 20 million US dollars (or equivalent in other
currencies) in monthly revenue, &quot;Kimi K3&quot; must be prominently displayed on the
user interface of such product or service.

4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal
use of the Software, defined as any use that does not make the Software, its
outputs, or its underlying capabilities available to third parties; or (b) any
use of the Software accessed through Moonshot AI&#x27;s official products or
certified inference partners.

5. THE SOFTWARE AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS”
BASIS, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT
LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE
AND NONINFRINGEMENT. IN NO EVENT SHALL MOONSHOT AI OR ITS AFFILIATES OR
COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER
IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN
CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

For any questions regarding this license, please contact &lt;license@moonshot.ai&gt;.</pre></details><details id="S018"><summary>S018 · moonshotai--Kimi-K3/README.md · exact</summary><p><a href="https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:51.830131+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>57de265b5842dfa465c6e73b368b0e15a89b8793b5450528dad577da202cc6fe</code><br>Computed UTF-8 SHA-256: <code>57de265b5842dfa465c6e73b368b0e15a89b8793b5450528dad577da202cc6fe</code><br>Recorded bytes 45261; embedded UTF-8 bytes 45261</p><p><b>content lines 38-47</b></p><pre>## 1. Model Introduction

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world&#x27;s first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

### Key Features
- **New Architecture**: Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE sparsity with a Stable LatentMoE framework that activates 16 out of 896 experts — yielding an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.
- **Long-Horizon Coding**: Operating with minimal human oversight, Kimi K3 sustains long engineering sessions, navigates massive repositories, and orchestrates terminal tools — from GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and even chip design.
- **Agentic Knowledge Work**: Kimi K3 advances end-to-end knowledge work, producing deep research with interactive visualizations, widgets and dashboards, and motion design and video editing, powered by its native multimodal architecture.
- **Native Multimodality &amp; Long Context**: Kimi K3 understands text, images, and video within the same model, and supports a 1-million-token context window.
- **Open Frontier Weights**: We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further innovation.</pre><p><b>content lines 582-622</b></p><pre>All Kimi K3 results are obtained with reasoning effort set to &#x27;max&#x27; and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell reports the scores without and with tool augmentation (general tools for HLE-Full, Python for the vision benchmarks), in that order.

1. **Reasoning &amp; knowledge benchmarks**
   - **CritPt and AA-LCR.** Scores are cited from [Artificial Analysis](https://artificialanalysis.ai/) as of July 23, 2026.
2. **Coding benchmarks**
   - **DeepSWE.** Kimi K3 is evaluated with the Kimi Code harness. The GLM-5.2 score is taken from the [GLM-5.2 release blog](https://z.ai/blog/glm-5.2); all remaining scores are from the official [DeepSWE leaderboard](https://deepswe.datacurve.ai/), under which Kimi K3 attains 67.3 with the mini-SWE-agent harness. We report the DeepSWE v1.1 tasks.
   - **Terminal-Bench 2.1.** Kimi K3 is evaluated with the Kimi Code harness. For all other models, we report the best score across harnesses: GLM-5.2 with Claude Code ([GLM-5.2 release blog](https://z.ai/blog/glm-5.2)); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 ([Artificial Analysis](https://artificialanalysis.ai/evaluations/terminalbench-v2-1)); GPT-5.5 and GPT-5.6 Sol with Codex ([OpenAI](https://openai.com/index/previewing-gpt-5-6-sol/)).
   - **ProgramBench.** Kimi K3 is evaluated with the Kimi Code harness. The GLM-5.2 score is from the [GLM-5.2 release blog](https://z.ai/blog/glm-5.2); all other scores are from [Vals AI](https://www.vals.ai/benchmarks/programbench).
   - **SWE-Marathon.** Kimi K3, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT-5.6 Sol is evaluated with the Codex harness. The GLM-5.2 score is from the [GLM-5.2 release blog](https://z.ai/blog/glm-5.2). Our evaluation is based on an H20-calibrated branch of the [official tasks](https://www.swe-marathon.org/) as of July 9, 2026, prior to the final v1.1 release: the Docker images, performance gates, and reference oracles for the GPU tasks have been recalibrated for H20, while the correctness and anti-cheat validators remain unchanged. Additionally, Claude Fable 5 hit fallbacks on 35% of the tasks in our evaluation, which may have negatively impacted its measured performance.
   - **FrontierSWE.** Kimi K3 is evaluated with the Kimi Code harness and GPT-5.6 Sol with the Codex harness; all other results are from [FrontierSWE](https://www.frontierswe.com/). Dominance scores are recomputed from the raw scores using the official evaluation script and are current as of July 16, 2026.
   - **PostTrainBench.** Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the official [PostTrainBench](https://posttrainbench.com/) results. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are evaluated with the official Harbor implementation at maximum reasoning effort, averaged over three runs on H20 GPUs (instead of H100 in the official setting) — Kimi K3 and Claude Fable 5 with the Claude Code harness, and GPT-5.6 Sol with the Codex harness.
   - **MLS-Bench-Lite.** Kimi K3 is evaluated with the Kimi Code harness; GLM-5.2 and the Claude models with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol with the Codex harness.
   - **SciCode.** Scores are cited from [Artificial Analysis](https://artificialanalysis.ai/) as of July 23, 2026.
   - **Kimi Code Bench 2.0 (in-house).** Kimi K3 is evaluated with the Kimi Code harness (it attains 73.7 with the Claude Code harness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol with the Codex harness. All models are evaluated at maximum reasoning effort, except GPT-5.5, which uses the &quot;xhigh&quot; setting. As the benchmark includes cybersecurity and safety-related tasks, we also disclose the fraction of refused or fallback tasks: Claude Fable 5 hit 13 fallbacks and 1 refusal out of 80 tasks; 10 refusals out of 80 tasks entered GPT-5.6 Sol&#x27;s cyber guard; GPT-5.5 had 3 refusals out of 80 tasks.
3. **Agentic benchmarks**
   - **OfficeQA Pro.** Each test case provides the agent with the entire PDF corpus, with all PDFs rendered as images and no machine-readable text available.
   - **OfficeQA Pro and SpreadsheetBench 2.** Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol are evaluated with the Codex harness.
   - **MCP-Atlas.** All models are evaluated on the 500-task public subset with a 100-turn limit, using Gemini 3.1 Pro as the judge.
   - **AutomationBench.** All models are evaluated on the 600-task public subset, following the official GitHub setup in all other respects.
   - **BrowseComp.** We adopt a context-compaction strategy triggered at 300K tokens. When evaluated with the full 1M-token context window and no context management, Kimi K3 achieves a score of 90.4. The results of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited from [Anthropic](https://www.anthropic.com/news/claude-fable-5-mythos-5) and [OpenAI](https://openai.com/index/gpt-5-6/).
   - **GDPval-AA v2, AA-Briefcase, τ³-Banking, Harvey Lab-AA, and APEX-Agents.** Scores are cited from [Artificial Analysis](https://artificialanalysis.ai/) and the [APEX-Agents leaderboard](https://www.mercor.com/apex/apex-agents-leaderboard/) as of July 23, 2026. For Harvey Lab-AA, we report the criterion pass rate.
   - **CorpFin v2, Finance Agent v2, and Legal Research Bench.** Scores are cited from [Vals AI](https://www.vals.ai/).
   - **Agents&#x27; Last Exam.** Scores are cited from the [official leaderboard](https://agents-last-exam.org/leaderboard) as of July 23, 2026; we report the leaderboard&#x27;s primary pass-rate metric. On the leaderboard, each model is paired with a specific harness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code. &lt;sup&gt;†&lt;/sup&gt; The Claude Fable 5 entry runs at xhigh effort with 40% of tasks annotated as downgraded.
4. **Multimodal benchmarks**
   - Except for ZeroBench, which follows the official setting and is run five times, all multimodal scores are averaged over three runs. MMMU-Pro is evaluated following the official protocol, preserving the original input order and prepending images to the text input.
   - **PerceptionBench** is an in-house benchmark that focuses on atomic visual perception capabilities.

&lt;/details&gt;

## 4. Native MXFP4 Quantization

Kimi K3 applies quantization-aware training from the SFT stage onward, using MXFP4 weights with MXFP8 activations for broad hardware compatibility.

## 5. Deployment

&gt; [!Note]
&gt; You can access Kimi K3&#x27;s API on https://platform.kimi.ai by selecting `kimi-k3`, and we provide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is recommended to run on the following inference engines:

- [vLLM](https://github.com/vllm-project/vllm) — see [recipes](https://recipes.vllm.ai/moonshotai/Kimi-K3)
- [SGLang](https://github.com/sgl-project/sglang) — see [cookbook](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3)
- [TokenSpeed](https://github.com/lightseekorg/tokenspeed) — see [recipes](https://lightseek.org/tokenspeed/recipes/models#kimi-k3)</pre></details><details id="S019"><summary>S019 · moonshotai--Kimi-K3/config.json · exact</summary><p><a href="https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:52.559498+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>9710e121a58d03ac92c8d6da287a19541994319afbbe6d6202af001ffd379213</code><br>Computed UTF-8 SHA-256: <code>9710e121a58d03ac92c8d6da287a19541994319afbbe6d6202af001ffd379213</code><br>Recorded bytes 7006; embedded UTF-8 bytes 7006</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;kimi_k3&quot;,
  &quot;architectures&quot;: [
    &quot;KimiK3ForConditionalGeneration&quot;
  ],
  &quot;dtype&quot;: &quot;bfloat16&quot;,
  &quot;auto_map&quot;: {
    &quot;AutoConfig&quot;: &quot;configuration_kimi_k3.KimiK3Config&quot;,
    &quot;AutoModel&quot;: &quot;modeling_kimi_k3.KimiK3ForConditionalGeneration&quot;,
    &quot;AutoModelForCausalLM&quot;: &quot;modeling_kimi_k3.KimiK3ForConditionalGeneration&quot;
  },
  &quot;quantization_config_summary&quot;: {}
}</pre></details><details id="S020"><summary>S020 · moonshotai--Kimi-K3/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:53.031449+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>5d0803c94db9cd78763499e0956c95fd5a225c14a727e5a6cf5db3f96f010a6e</code><br>Computed UTF-8 SHA-256: <code>5d0803c94db9cd78763499e0956c95fd5a225c14a727e5a6cf5db3f96f010a6e</code><br>Recorded bytes 3478; embedded UTF-8 bytes 3478</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;TikTokenTokenizer&quot;,
  &quot;auto_map&quot;: {
    &quot;AutoTokenizer&quot;: [
      &quot;tokenization_kimi.TikTokenTokenizer&quot;,
      null
    ]
  },
  &quot;model_max_length&quot;: 1000000000000000019884624838656,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S021"><summary>S021 · google--gemma-4-31B-it/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/google/gemma-4-31B-it" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/google/gemma-4-31B-it</a></p><p class="small">Retrieved 2026-10-08T08:49:53.351394+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>2460d90885de08ec46899247752cc654db655edb7fe567f204c17b567b1c9f33</code><br>Computed UTF-8 SHA-256: <code>2460d90885de08ec46899247752cc654db655edb7fe567f204c17b567b1c9f33</code><br>Recorded bytes 24252; embedded UTF-8 bytes 24252</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;google/gemma-4-31B-it&quot;,
  &quot;sha&quot;: &quot;842da3794eaa0b77d5f08bae87a17459d91ff475&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;apache-2.0&quot;,
  &quot;files&quot;: [
    &quot;.eval_results/mmmu_pro.yaml&quot;,
    &quot;.gitattributes&quot;,
    &quot;README.md&quot;,
    &quot;chat_template.jinja&quot;,
    &quot;config.json&quot;,
    &quot;generation_config.json&quot;,
    &quot;model-00001-of-00002.safetensors&quot;,
    &quot;model-00002-of-00002.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;processor_config.json&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;
  ]
}</pre></details><details id="S022"><summary>S022 · google--gemma-4-31B-it/README.md · exact</summary><p><a href="https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:53.917545+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>e209f5c435b1e66055bd6d7c29185103ae8d2027a03a96b359335154f206b82b</code><br>Computed UTF-8 SHA-256: <code>e209f5c435b1e66055bd6d7c29185103ae8d2027a03a96b359335154f206b82b</code><br>Recorded bytes 27963; embedded UTF-8 bytes 27963</p><p><b>content lines 18-26</b></p><pre>    &lt;a href=&quot;https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/&quot; target=&quot;_blank&quot;&gt;Launch Blog&lt;/a&gt; |
    &lt;a href=&quot;https://ai.google.dev/gemma/docs/core&quot; target=&quot;_blank&quot;&gt;Documentation&lt;/a&gt;|
    &lt;a href=&quot;https://arxiv.org/abs/2607.02770&quot; target=&quot;_blank&quot;&gt;Technical Report&lt;/a&gt;
    &lt;br&gt;
    &lt;b&gt;License&lt;/b&gt;: &lt;a href=&quot;https://ai.google.dev/gemma/docs/gemma_4_license&quot; target=&quot;_blank&quot;&gt;Apache 2.0&lt;/a&gt; | &lt;b&gt;Authors&lt;/b&gt;: &lt;a href=&quot;https://deepmind.google/models/gemma/&quot; target=&quot;_blank&quot;&gt;Google DeepMind&lt;/a&gt;
&lt;/p&gt;

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. 
</pre><p><b>content lines 127-146</b></p><pre>## Getting Started

You can use all Gemma 4 models with the latest version of Transformers. To get started, install the necessary dependencies in your environment:

`pip install -U transformers torch accelerate`

Once you have everything installed, you can proceed to load the model with the code below:

```python
from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = &quot;google/gemma-4-31B-it&quot;

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype=&quot;auto&quot;,
    device_map=&quot;auto&quot;
)</pre><p><b>content lines 427-466</b></p><pre>## **Model Data**

Data used for model training and how the data was processed.

### **Training Dataset**

Our pre-training dataset is a large-scale, diverse collection of data encompassing a wide range of domains and modalities, which includes web documents, code, images, audio, with a cutoff date of January 2025. Here are the key components:

* **Web Documents**: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary. The training dataset includes content in over 140 languages.  
* **Code**: Exposing the model to code helps it to learn the syntax and patterns of programming languages, which improves its ability to generate code and understand code-related questions.  
* **Mathematics**: Training on mathematical text helps the model learn logical reasoning, symbolic representation, and to address mathematical queries.  
* **Images**: A wide range of images enables the model to perform image analysis and visual data extraction tasks.

The combination of these diverse data sources is crucial for training a powerful multimodal model that can handle a wide variety of different tasks and data formats.

### **Data Preprocessing**

Here are the key data cleaning and filtering methods applied to the training data:

* **CSAM Filtering**: Rigorous CSAM (Child Sexual Abuse Material) filtering was applied at multiple stages in the data preparation process to ensure the exclusion of harmful and illegal content.  
* **Sensitive Data Filtering**: As part of making Gemma pre-trained models safe and reliable, automated techniques were used to filter out certain personal information and other sensitive data from training sets.  
* **Additional methods**: Filtering based on content quality and safety in line with [our policies](https://ai.google/static/documents/ai-responsibility-update-published-february-2025.pdf).

## **Ethics and Safety**

As open models become central to enterprise infrastructure, provenance and security are paramount. Developed by Google DeepMind, Gemma 4 undergoes the same rigorous safety evaluations as our proprietary Gemini models. 

### **Evaluation Approach**

Gemma 4 models were developed in partnership with internal safety and responsible AI teams. A range of automated as well as human evaluations were conducted to help improve model safety. These evaluations align with [Google’s AI principles](https://ai.google/principles/), as well as safety policies, which aim to prevent our generative AI models from generating harmful content, including:

* Content related to child sexual abuse material and exploitation   
* Dangerous content (e.g., promoting suicide, or instructing in activities that could cause real-world harm)   
* Sexually explicit content  
* Hate speech (e.g., dehumanizing members of protected groups)   
* Harassment (e.g., encouraging violence against people)

### **Evaluation Results**

For all areas of safety testing, we saw major improvements in all categories of content safety relative to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 and 3n models in improving safety, while keeping unjustified refusals low. All testing was conducted without safety filters to evaluate the model capabilities and behaviors. For both text-to-text and image-to-text, and across all model sizes, the model produced minimal policy violations, and showed significant improvements over previous Gemma models&#x27; performance. </pre></details><details id="S023"><summary>S023 · google--gemma-4-31B-it/config.json · exact</summary><p><a href="https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:54.446733+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>e967dd38bc5cfd38bd09a995a7bf4a754075df2b46aba68f7fbb5a791e6d8dd1</code><br>Computed UTF-8 SHA-256: <code>e967dd38bc5cfd38bd09a995a7bf4a754075df2b46aba68f7fbb5a791e6d8dd1</code><br>Recorded bytes 4621; embedded UTF-8 bytes 4621</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;gemma4&quot;,
  &quot;architectures&quot;: [
    &quot;Gemma4ForConditionalGeneration&quot;
  ],
  &quot;dtype&quot;: &quot;bfloat16&quot;,
  &quot;quantization_config_summary&quot;: {}
}</pre></details><details id="S024"><summary>S024 · google--gemma-4-31B-it/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:54.923842+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>9f4fec4b1dc6ecddf8f4a92e9caea5971c0e67d81309f3f9066a2bee8c362633</code><br>Computed UTF-8 SHA-256: <code>9f4fec4b1dc6ecddf8f4a92e9caea5971c0e67d81309f3f9066a2bee8c362633</code><br>Recorded bytes 3082; embedded UTF-8 bytes 3082</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;GemmaTokenizer&quot;,
  &quot;model_max_length&quot;: 1000000000000000019884624838656,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S025"><summary>S025 · nvidia--NVIDIA-Nemotron-3-Super-120B-A12B-BF16/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16</a></p><p class="small">Retrieved 2026-10-08T08:49:53.557026+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>9aa7841acec0a8892f596504ef51fdea8c8d09a33c950f1489b03b7a8e0fb752</code><br>Computed UTF-8 SHA-256: <code>9aa7841acec0a8892f596504ef51fdea8c8d09a33c950f1489b03b7a8e0fb752</code><br>Recorded bytes 17077; embedded UTF-8 bytes 17077</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16&quot;,
  &quot;sha&quot;: &quot;2dc98e2afe4face0e4ce40972a915c45368bd34a&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;other&quot;,
  &quot;files&quot;: [
    &quot;.eval_results/gpqa.yaml&quot;,
    &quot;.eval_results/gpqa_with_tools.yaml&quot;,
    &quot;.eval_results/hle.yaml&quot;,
    &quot;.eval_results/hle_with_tools.yaml&quot;,
    &quot;.eval_results/mmlu_pro.yaml&quot;,
    &quot;.eval_results/swe_bench_verified.yaml&quot;,
    &quot;.eval_results/terminal-bench_2.0.yaml&quot;,
    &quot;.gitattributes&quot;,
    &quot;README.md&quot;,
    &quot;__init__.py&quot;,
    &quot;accuracy_chart.png&quot;,
    &quot;bias.md&quot;,
    &quot;chat_template.jinja&quot;,
    &quot;config.json&quot;,
    &quot;configuration_nemotron_h.py&quot;,
    &quot;explainability.md&quot;,
    &quot;generation_config.json&quot;,
    &quot;model-00001-of-00050.safetensors&quot;,
    &quot;model-00002-of-00050.safetensors&quot;,
    &quot;model-00003-of-00050.safetensors&quot;,
    &quot;model-00004-of-00050.safetensors&quot;,
    &quot;model-00005-of-00050.safetensors&quot;,
    &quot;model-00006-of-00050.safetensors&quot;,
    &quot;model-00007-of-00050.safetensors&quot;,
    &quot;model-00008-of-00050.safetensors&quot;,
    &quot;model-00009-of-00050.safetensors&quot;,
    &quot;model-00010-of-00050.safetensors&quot;,
    &quot;model-00011-of-00050.safetensors&quot;,
    &quot;model-00012-of-00050.safetensors&quot;,
    &quot;model-00013-of-00050.safetensors&quot;,
    &quot;model-00014-of-00050.safetensors&quot;,
    &quot;model-00015-of-00050.safetensors&quot;,
    &quot;model-00016-of-00050.safetensors&quot;,
    &quot;model-00017-of-00050.safetensors&quot;,
    &quot;model-00018-of-00050.safetensors&quot;,
    &quot;model-00019-of-00050.safetensors&quot;,
    &quot;model-00020-of-00050.safetensors&quot;,
    &quot;model-00021-of-00050.safetensors&quot;,
    &quot;model-00022-of-00050.safetensors&quot;,
    &quot;model-00023-of-00050.safetensors&quot;,
    &quot;model-00024-of-00050.safetensors&quot;,
    &quot;model-00025-of-00050.safetensors&quot;,
    &quot;model-00026-of-00050.safetensors&quot;,
    &quot;model-00027-of-00050.safetensors&quot;,
    &quot;model-00028-of-00050.safetensors&quot;,
    &quot;model-00029-of-00050.safetensors&quot;,
    &quot;model-00030-of-00050.safetensors&quot;,
    &quot;model-00031-of-00050.safetensors&quot;,
    &quot;model-00032-of-00050.safetensors&quot;,
    &quot;model-00033-of-00050.safetensors&quot;,
    &quot;model-00034-of-00050.safetensors&quot;,
    &quot;model-00035-of-00050.safetensors&quot;,
    &quot;model-00036-of-00050.safetensors&quot;,
    &quot;model-00037-of-00050.safetensors&quot;,
    &quot;model-00038-of-00050.safetensors&quot;,
    &quot;model-00039-of-00050.safetensors&quot;,
    &quot;model-00040-of-00050.safetensors&quot;,
    &quot;model-00041-of-00050.safetensors&quot;,
    &quot;model-00042-of-00050.safetensors&quot;,
    &quot;model-00043-of-00050.safetensors&quot;,
    &quot;model-00044-of-00050.safetensors&quot;,
    &quot;model-00045-of-00050.safetensors&quot;,
    &quot;model-00046-of-00050.safetensors&quot;,
    &quot;model-00047-of-00050.safetensors&quot;,
    &quot;model-00048-of-00050.safetensors&quot;,
    &quot;model-00049-of-00050.safetensors&quot;,
    &quot;model-00050-of-00050.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;modeling_nemotron_h.py&quot;,
    &quot;privacy.md&quot;,
    &quot;safety.md&quot;,
    &quot;special_tokens_map.json&quot;,
    &quot;super_v3_reasoning_parser.py&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;
  ]
}</pre></details><details id="S026"><summary>S026 · nvidia--NVIDIA-Nemotron-3-Super-120B-A12B-BF16/README.md · exact</summary><p><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:54.076854+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>cade76c064eda90d4f53e12614f45816ae6cfe43dacc8a9410abcc2fc7a17e59</code><br>Computed UTF-8 SHA-256: <code>cade76c064eda90d4f53e12614f45816ae6cfe43dacc8a9410abcc2fc7a17e59</code><br>Recorded bytes 82606; embedded UTF-8 bytes 82606</p><p><b>content lines 61-74</b></p><pre>## Model Summary

| | |
|:---|:---|
| **Total Parameters** | 120B (12B active) |
| **Architecture** | LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) |
| **Context Length** | Up to 1M tokens |
| **Minimum GPU Requirement** | 8× H100-80GB |
| **Supported Languages** | English, French, German, Italian, Japanese, Spanish, Chinese |
| **Best For** | Agentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG |
| **Reasoning Mode** | Configurable on/off via chat template (`enable_thinking=True/False`) |
| **Speculative Decoding** | Includes a built-in MTP head, with an updated MTPv2 head available as a separate [checkpoint](https://huggingface.co/nvidia/Nemotron-3-Super-120B-A12B-BF16-MTPv2) |
| **License** | [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/) |
| **Release Date** | March 11, 2026 |</pre><p><b>content lines 110-114</b></p><pre>## License/Terms of Use

**Governing Download Terms:** Use of this model is governed by the [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/).

**Governing Download Terms with NIM:** The NIM container is governed by the [NVIDIA Software License Agreement](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-license-agreement/) and [Product-Specific Terms for AI Products](https://www.nvidia.com/en-us/agreements/enterprise-software/product-specific-terms-for-ai-products/). Use of this model is governed by the [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/).</pre><p><b>content lines 159-161</b></p><pre>All evaluation results were collected via [Nemo Evaluator SDK](https://github.com/NVIDIA-NeMo/Evaluator) and for most benchmarks, the [Nemo Skills Harness](https://github.com/NVIDIA-NeMo/Skills). For reproducibility purposes, more details on the evaluation settings can be found in the [Nemo Evaluator SDK configs folder](https://github.com/NVIDIA-NeMo/Evaluator/tree/main/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super) and the [reproducibility tutorial for Nemotron 3 Super](https://github.com/NVIDIA-NeMo/Evaluator/blob/main/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md). The open source container on Nemo Skills packaged via NVIDIA&#x27;s Nemo Evaluator SDK used for evaluations can be found [here](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/eval-factory/containers/nemo_skills). In addition to Nemo Skills, the evaluations also used dedicated open-source packaged containers for Tau-2 Bench (default prompt), Terminal Bench Hard (48 tasks), ScaleAI Multi Challenge Multi-turn Instruction Following, and Ruler. 

The following benchmarks are not onboarded yet in our open source tools and for these we used either their official open source implementation or otherwise an internal scaffolding that we plan to open source in the future: SWE Bench Verified (OpenHands), SWE Bench Multilingual (OpenHands), BrowseComp with Search (internal implementation with Serp API), Terminal Bench Core 2.0 (Harbor).</pre><p><b>content lines 184-206</b></p><pre>## Model Design

The model utilizes the **LatentMoE** architecture, where tokens are projected into a smaller latent dimension for expert routing and computation, improving accuracy per byte. The Super model is pre-trained using NVFP4 quantization — the first model in the Nemotron 3 family trained at this precision. The majority of linear layers use NVFP4 for weights, activations, and gradients, while select layers (including latent projections, MTP layers, QKV/attention projections, and embeddings) are maintained in BF16 or MXFP8 for training stability. The model includes **Multi-Token Prediction (MTP)** layers using a shared-weight design across prediction heads. This improves training signal quality, enables faster inference via native speculative decoding, and supports more stable autoregressive drafting at longer draft lengths compared to independently trained offset heads.

## Training Methodology

Stage 1: Pre-Training

* [NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-Base-BF16) model was pre-trained for over 25T tokens using crawled and synthetic code, math, science, and general knowledge data. Training leveraged NVFP4 quantization for efficiency. All datasets are disclosed in the [Training and Evaluation Datasets](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16#training-and-evaluation-datasets) section of this document. Major portions of the pre-training corpus are released in the [Nemotron-Pre-Training-Datasets](https://huggingface.co/collections/nvidia/nemotron-pre-training-datasets) collection.
* Software used for pre-training: [Megatron-LM](https://github.com/NVIDIA/Megatron-LM)

Stage 2: Supervised Fine-Tuning

* The model was further fine-tuned on synthetic code, math, science, tool calling, instruction following, structured outputs, and general knowledge data. This stage incorporated data designed to support long-range retrieval and multi-document aggregation. All datasets are disclosed in the [Training and Evaluation Datasets](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16#training-and-evaluation-datasets) section of this document. Major portions of the fine-tuning corpus are released in the [Nemotron-Post-Training-v3](https://huggingface.co/collections/nvidia/nemotron-post-training-v3) collection. [Data Designer](https://github.com/NVIDIA-NeMo/DataDesigner) is one of the libraries used to prepare these corpora.

Stage 3: Reinforcement Learning

* The model underwent multi-environment reinforcement learning using asynchronous GRPO (Group Relative Policy Optimization) across math, code, science, instruction following, multi-step tool use, multi-turn conversations, and structured output environments. It utilized an asynchronous RL architecture that fully decouples training from inference across separate GPU devices, leveraging in-flight weight updates and MTP to accelerate rollout generation. Conversational quality was further refined through RLHF. All datasets are disclosed in the *Training and Evaluation Datasets* section of this document. The RL environments and datasets are released as part of [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym).
* Software used for reinforcement learning: [NeMo RL](https://github.com/NVIDIA-NeMo/RL), [NeMo Gym](https://github.com/NVIDIA-NeMo/Gym)

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 model is a result of the above work.

The end-to-end training recipe is available in the [NVIDIA Nemotron Developer Repository](https://github.com/NVIDIA-NeMo/Nemotron). Evaluation results can be replicated using the [NeMo Evaluator SDK](https://github.com/NVIDIA-NeMo/Evaluator). [Data Designer](https://github.com/NVIDIA-NeMo/DataDesigner) is one of the libraries used to prepare the pre and post training datasets. More details on the datasets and synthetic data generation methods can be found in the technical report [NVIDIA Nemotron 3 Super Technical Report](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf).</pre><p><b>content lines 547-574</b></p><pre>#### Transformers
The model has been integrated into 🤗 Transformers since v5.3.0. We recommend using the [Nemotron 3 Super](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo/tags?version=26.02.nemotron_3_super) container from the NeMo Framework to ensure all required libraries are available.

```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained(&quot;nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16&quot;)
model = AutoModelForCausalLM.from_pretrained(
    &quot;nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16&quot;,
    torch_dtype=torch.bfloat16,
    device_map=&quot;auto&quot;
)
```

If your Transformers version is lower than v5.3.0, please add `trust_remote_code=True` when loading the model:
```python
model = AutoModelForCausalLM.from_pretrained(
    &quot;nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16&quot;,
    torch_dtype=torch.bfloat16,
    device_map=&quot;auto&quot;,
    trust_remote_code=True
)
```

Please note that the model supports up to a 1M context size, although the default context size in the Hugging Face configuration is 256k due to higher VRAM requirements.

Here is an example of generating outputs with reasoning enabled (the default):</pre><p><b>content lines 650-660</b></p><pre>#### **Base Pre-Training Corpus (Nemotron 3 Foundation)**

The foundation of the model is trained on the **Nemotron-3-Nano** corpus, comprising the following collections:

| Dataset Collection | Token Counts | Description |
| :--- | :--- | :--- |
| **Nemotron-CC-v2** &amp; **v2.1** | 9.13T | A massive collection of English web data filtered from Common Crawl, including 2.5T+ tokens of new organic, translated, and synthetically rephrased content. |
| **Nemotron-CC-Code-v1** | 427.9B | High-quality code tokens extracted from Common Crawl using the Lynx + LLM pipeline to preserve structure and equations. |
| **Nemotron-Pretraining-Code-v1** &amp; **v2** | 1.09T | Curated GitHub code references with multi-stage filtering, deduplication, and large-scale synthetic code data. |
| **Nemotron-CC-Math-v1** | 133.3B | High-quality math pre-training dataset preserving LaTeX formatting and mathematical structures. |
| **Nemotron-Pretraining-Specialized-v1** | 336.4B | Synthetic datasets targeting specialized domains such as STEM reasoning and scientific coding. |</pre><p><b>content lines 720-720</b></p><pre>| Competitive Coding RL data from [Nemotron-Cascade-RL-SWE](https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-SWE) | 01/10/2026 |</pre><p><b>content lines 739-763</b></p><pre>## Private Non-publicly Accessible Datasets of Third Parties

| Dataset | Model(s) used |
|---------|---------------|
| Global Regulation | Unknown |
| TAUS Translation Memory | Unknown |
| Scale HLE | Unknown |
| HackerRank Coding | Unknown |
| RL data for Search | Gemini 3; GPT-5 * |

* Models used for prompt generation only

## Private Non-publicly Accessible Datasets by NVIDIA

| Dataset | Model(s) used |
|---------|---------------|
| Simple Minesweeper | \- |
| Simple Sudoku | \- |
| Multitool Typewriter Hard | \- |
| Machine Translation of News Commentary and TAUS Translation Memory | \- |
| Machine Translation of STEM - | [Qwen2.5-14B-Instruct](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct) |
| Competitive Coding RL data from Nemotron Cascade | \- |
| Long context RL | \- |
| Single-step SWE RL for patch generation | \- |
| OpenHands SWE | \- |</pre></details><details id="S027"><summary>S027 · nvidia--NVIDIA-Nemotron-3-Super-120B-A12B-BF16/config.json · exact</summary><p><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:54.696989+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>699f34f0fc645d29ebffa5767fb59e6ae6ec98e3a4605485eb9913256d0df7e6</code><br>Computed UTF-8 SHA-256: <code>699f34f0fc645d29ebffa5767fb59e6ae6ec98e3a4605485eb9913256d0df7e6</code><br>Recorded bytes 1924; embedded UTF-8 bytes 1924</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;nemotron_h&quot;,
  &quot;architectures&quot;: [
    &quot;NemotronHForCausalLM&quot;
  ],
  &quot;dtype&quot;: &quot;bfloat16&quot;,
  &quot;auto_map&quot;: {
    &quot;AutoConfig&quot;: &quot;configuration_nemotron_h.NemotronHConfig&quot;,
    &quot;AutoModelForCausalLM&quot;: &quot;modeling_nemotron_h.NemotronHForCausalLM&quot;
  },
  &quot;quantization_config_summary&quot;: {}
}</pre></details><details id="S028"><summary>S028 · nvidia--NVIDIA-Nemotron-3-Super-120B-A12B-BF16/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:55.138727+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>10f93eabcb9b1602fbb991d6308e787ce1df28ee9cd7a1c6d1e8c3f338b957bc</code><br>Computed UTF-8 SHA-256: <code>10f93eabcb9b1602fbb991d6308e787ce1df28ee9cd7a1c6d1e8c3f338b957bc</code><br>Recorded bytes 177209; embedded UTF-8 bytes 177209</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;PreTrainedTokenizerFast&quot;,
  &quot;model_max_length&quot;: 262144,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S029"><summary>S029 · allenai--Olmo-3-1125-32B/metadata.json · exact</summary><p><a href="https://huggingface.co/api/models/allenai/Olmo-3-1125-32B" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/models/allenai/Olmo-3-1125-32B</a></p><p class="small">Retrieved 2026-10-08T08:49:53.557487+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>ed15b578f3bd1e853f059c8b7317e22b5e5a5e4c8865fb3b27fbcbe1997aff7f</code><br>Computed UTF-8 SHA-256: <code>ed15b578f3bd1e853f059c8b7317e22b5e5a5e4c8865fb3b27fbcbe1997aff7f</code><br>Recorded bytes 9208; embedded UTF-8 bytes 9208</p><p><b>JSON $.id, $.sha, $.private, $.gated, $.disabled, $.cardData.license, $.siblings</b></p><pre>{
  &quot;id&quot;: &quot;allenai/Olmo-3-1125-32B&quot;,
  &quot;sha&quot;: &quot;c2b61dae89a1ad10e4ad5653d0e46b590902607b&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;license_label&quot;: &quot;apache-2.0&quot;,
  &quot;files&quot;: [
    &quot;.gitattributes&quot;,
    &quot;README.md&quot;,
    &quot;config.json&quot;,
    &quot;generation_config.json&quot;,
    &quot;merges.txt&quot;,
    &quot;model-00001-of-00014.safetensors&quot;,
    &quot;model-00002-of-00014.safetensors&quot;,
    &quot;model-00003-of-00014.safetensors&quot;,
    &quot;model-00004-of-00014.safetensors&quot;,
    &quot;model-00005-of-00014.safetensors&quot;,
    &quot;model-00006-of-00014.safetensors&quot;,
    &quot;model-00007-of-00014.safetensors&quot;,
    &quot;model-00008-of-00014.safetensors&quot;,
    &quot;model-00009-of-00014.safetensors&quot;,
    &quot;model-00010-of-00014.safetensors&quot;,
    &quot;model-00011-of-00014.safetensors&quot;,
    &quot;model-00012-of-00014.safetensors&quot;,
    &quot;model-00013-of-00014.safetensors&quot;,
    &quot;model-00014-of-00014.safetensors&quot;,
    &quot;model.safetensors.index.json&quot;,
    &quot;olmo-base.png&quot;,
    &quot;special_tokens_map.json&quot;,
    &quot;tokenizer.json&quot;,
    &quot;tokenizer_config.json&quot;,
    &quot;vocab.json&quot;
  ]
}</pre></details><details id="S030"><summary>S030 · allenai--Olmo-3-1125-32B/README.md · exact</summary><p><a href="https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md</a></p><p class="small">Retrieved 2026-10-08T08:49:54.076540+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>8f227fd377065f4f93b50f9e14606224a021517f2315bcd537a22bb56ba8680f</code><br>Computed UTF-8 SHA-256: <code>8f227fd377065f4f93b50f9e14606224a021517f2315bcd537a22bb56ba8680f</code><br>Recorded bytes 14957; embedded UTF-8 bytes 14957</p><p><b>content lines 149-159</b></p><pre># Model Card for Olmo 3 32B

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach.

Olmo is a series of **O**pen **l**anguage **mo**dels designed to enable the science of language models. 
These models are trained on the Dolma 3 dataset. We are releasing all code, checkpoints, and associated training details. 

| Size   | Training Tokens | Layers | Hidden Size | Q Heads | KV Heads | Context Length |
|--------|-----------------|--------|-------------|---------|----------|----------------|
| [OLMo 3 7B](https://huggingface.co/allenai/Olmo-3-1025-7B) | 5.93 Trillion | 32 | 4096 | 32 | 32 | 65,536 |
| [OLMo 3 32B](https://huggingface.co/allenai/Olmo-3-1125-32B) | 5.50 Trillion | 64 | 5120 | 40 | 8 | 65,536 |</pre><p><b>content lines 172-192</b></p><pre>## Installation

Olmo 3 is supported in transformers v4.57.0 or higher:
```bash
pip install transformers&gt;=4.57.0
```

## Inference

You can use OLMo with the standard HuggingFace transformers library:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
olmo = AutoModelForCausalLM.from_pretrained(&quot;allenai/Olmo-3-1125-32B&quot;)
tokenizer = AutoTokenizer.from_pretrained(&quot;allenai/Olmo-3-1125-32B&quot;)
message = [&quot;Language modeling is &quot;]
inputs = tokenizer(message, return_tensors=&#x27;pt&#x27;, return_token_type_ids=False)
# optional verifying cuda
# inputs = {k: v.to(&#x27;cuda&#x27;) for k,v in inputs.items()}
# olmo = olmo.to(&#x27;cuda&#x27;)
response = olmo.generate(**inputs, max_new_tokens=100, do_sample=True, top_k=0, temperature=1.0, top_p=0.7)
print(tokenizer.batch_decode(response, skip_special_tokens=True)[0])</pre><p><b>content lines 207-219</b></p><pre>We have released checkpoints for these models. For pretraining, the naming convention is `stage1-stepXXX`.  The conventions for midtraining and long context are `stage2-ingredientY-stepXXX` and `stage3-stepXXX`, respectively.


To load a specific model revision with HuggingFace, simply add the argument `revision`:
```bash
olmo = AutoModelForCausalLM.from_pretrained(&quot;allenai/Olmo-3-1125-32B&quot;, revision=&quot;stage1-step10000&quot;)
```

Or, you can access all the revisions for the models via the following code snippet:
```python
from huggingface_hub import list_repo_refs
out = list_repo_refs(&quot;allenai/Olmo-3-1125-32B&quot;)
branches = [b.name for b in out.branches]</pre><p><b>content lines 245-253</b></p><pre>### Model Sources

- **Project Page:** https://allenai.org/olmo
- **Repositories:** 
    - Core repo (training, inference, fine-tuning etc.): https://github.com/allenai/OLMo-core
    - Evaluation code: https://github.com/allenai/OLMo-Eval
    - Further fine-tuning code: https://github.com/allenai/open-instruct
- **W&amp;B Report:** https://wandb.ai/ai2-llm/Olmo-3-1125-32B/reports/Olmo-3-32B-November-2025--VmlldzoxNTA4NzAxMw
- **Paper:** https://allenai.org/papers/olmo3</pre><p><b>content lines 278-306</b></p><pre>#### Stage 1: Initial Pretraining
- Dataset: [dolma3_mix-5.5T-1125](https://huggingface.co/datasets/allenai/dolma3_mix-5.5T-1125)
- 5.50T tokens
- Coverage: 94.83%+ of total pretraining budget

#### Stage 2: Mid-training
- Ingredient 1
  - Dataset: [dolma3-dolmino-mix-1125](https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1125)
  - 100B tokens
  - Mix composition: web pages, code, math/QA/thinking/instruction/PDFs
- Ingredient 2
  - Dataset: [dolma3-dolmino-mix-1125](https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1125)
  - 100B tokens
  - Mix composition: web pages, code, math/QA/thinking/instruction/PDFs

#### Stage 3: Long Context
- Dataset: [dolma3-longmino-mix-1125](https://huggingface.co/datasets/allenai/dolma3_longmino_mix-100B-1125)
- 100B tokens
- Mix composition: midtraining data and PDFs

#### Model Merging
- 7B Model: No merging
- 32B Model: 2 versions on 100B mix, merged before starting long context run.  Final checkpoint is merged 4 final checkpoints.

## Bias, Risks, and Limitations
Like any base language model or fine-tuned model without safety filtering, these models can easily be prompted by users to generate harmful and sensitive content. Such content may also be produced unintentionally, especially in cases involving bias, so we recommend that users consider the risks when applying this technology. Additionally, many statements from OLMo or any LLM are often inaccurate, so facts should be verified.

## License
This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with [Ai2&#x27;s Responsible Use Guidelines](https://allenai.org/responsible-use).</pre></details><details id="S031"><summary>S031 · allenai--Olmo-3-1125-32B/config.json · exact</summary><p><a href="https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:54.602731+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>72efc1f0ed1161e4205a54d94110f61d21c4c4315fde9ef17163bbc33c9a6149</code><br>Computed UTF-8 SHA-256: <code>72efc1f0ed1161e4205a54d94110f61d21c4c4315fde9ef17163bbc33c9a6149</code><br>Recorded bytes 2399; embedded UTF-8 bytes 2399</p><p><b>JSON selected configuration fields</b></p><pre>{
  &quot;model_type&quot;: &quot;olmo3&quot;,
  &quot;architectures&quot;: [
    &quot;Olmo3ForCausalLM&quot;
  ],
  &quot;dtype&quot;: &quot;bfloat16&quot;,
  &quot;quantization_config_summary&quot;: {}
}</pre></details><details id="S032"><summary>S032 · allenai--Olmo-3-1125-32B/tokenizer_config.json · exact</summary><p><a href="https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/tokenizer_config.json" target="_blank" rel="noopener noreferrer">https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/tokenizer_config.json</a></p><p class="small">Retrieved 2026-10-08T08:49:54.997152+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>ffa5b2cfee6e73d5dc0a7d3b7e02eb6ed60a491dd782f2bd8ad2b4659b2aafc6</code><br>Computed UTF-8 SHA-256: <code>ffa5b2cfee6e73d5dc0a7d3b7e02eb6ed60a491dd782f2bd8ad2b4659b2aafc6</code><br>Recorded bytes 4308; embedded UTF-8 bytes 4308</p><p><b>JSON selected tokenizer fields</b></p><pre>{
  &quot;tokenizer_class&quot;: &quot;GPT2Tokenizer&quot;,
  &quot;model_max_length&quot;: 65536,
  &quot;has_embedded_chat_template&quot;: false
}</pre></details><details id="S033"><summary>S033 · mistral-large-4/release.html · exact</summary><p><a href="https://mistral.ai/news/mistral-large-4/" target="_blank" rel="noopener noreferrer">https://mistral.ai/news/mistral-large-4/</a></p><p class="small">Retrieved 2026-10-08T08:51:37.392041+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>8abfb479b3f63bdaf5c869b4b53a2ffa91c9804df683d4792089534268101d69</code><br>Computed UTF-8 SHA-256: <code>8abfb479b3f63bdaf5c869b4b53a2ffa91c9804df683d4792089534268101d69</code><br>Recorded bytes 299586; embedded UTF-8 bytes 299586</p><p><b>HTML visible text lines 55-75</b></p><pre>Introducing
Mistral Large 4
Back to Blog
11 min read
October 6, 2026
By Mistral
Share this post
Copy url to clipboard
Copied
Le Chonk
Today, we’re launching a
public preview
of Mistral Large 4. Unofficially ML4, very officially:
le Chonk
. ML4 pushes the frontier of open-weight performance. You can try the preview API today on
Mistral Studio
. Weights drop end of this month.
Frontier performance
ML4 is a 1 trillion-parameter natively multimodal model with 52 billion active parameters. It is our largest and most capable model to date, and it continues to improve rapidly as we refine it.
The model demonstrates exceptional performance across coding, agentic workflows, and multimodal understanding. It already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe. On critical enterprise workloads, including cybersecurity, finance and law, we find it to be state-of-the-art among open models. In some domains such as visual grounding, it goes further still, surpassing even frontier closed models.
We will release the weights by the end of the month. Until then, we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.</pre><p><b>HTML visible text lines 84-90</b></p><pre>ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe. The public preview is served on that same infrastructure. It is a significant milestone in our long-term investment across infrastructure, research, and product development: state-of-the-art performance in critical verticals, delivered through open weights, designed to give customers control over their AI.
This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.
The model will be available across multiple regions worldwide, including a European deployment that Mistral operates end-to-end, independently of other digital service providers and under European law. Fun fact: a significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union.
We’ve been working closely with leading enterprises across the world in finance, engineering, manufacturing, logistics, pharmaceuticals, science, shipping, public sector, and other mission-critical industries to train ML4. In fact, the model uses the same training, customization, and RL environment we offer our customers through
Mistral Forge.
Try it today
There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.</pre><p><b>HTML visible text lines 168-176</b></p><pre>The improvements are not specific to the environments we train on; they transfer to downstream evals, and the final model owes them to both post-training stages (supervised fine-tuning and RL), as shown in the charts.
What comes next
This is only the beginning. ML4 is the first milestone on the roadmap funded by our €3 billion Series D — the largest equity round ever raised by a European technology company. That capital is already being put to work: we are significantly scaling up our compute capacity in our own European datacenters, and much more is coming online in the months ahead.
More compute means more training. The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation — there is substantial headroom ahead. As we scale up training on our expanded infrastructure, we expect large and rapid improvements in the weeks and months to come.
We will release the weights by the end of the month, along with more details on the architecture, additional benchmarks, and our post-training methodology. And ML4 is only the foundation: it will serve as the base for a new generation of specialized and optimized Mistral models, built for the industries and workloads our customers care about most.
The pace of progress from here will be fast. Stay tuned.
Mistral Large 4
New
Open-weight hybrid instruct-and-reasoning MoE with multimodal input; unifies instruction, reasoning, and agentic capabilities in a single model, state-of-the-art among open weights on cybersecurity, finance, and manufacturing, natively fluent in 160+ languages.</pre></details><details id="S034"><summary>S034 · https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125 · no_content</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125</a></p><p class="small">Retrieved 2026-10-08T08:51:38.073312+00:00 · Recorded HTTP None<br>Recorded SHA-256: <code>None</code><br>Computed UTF-8 SHA-256: <code>None</code><br>Recorded bytes None; embedded UTF-8 bytes None</p><p>small-file limit exceeded</p></details><details id="S035"><summary>S035 · https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125 · no_content</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125</a></p><p class="small">Retrieved 2026-10-08T08:51:37.392918+00:00 · Recorded HTTP None<br>Recorded SHA-256: <code>None</code><br>Computed UTF-8 SHA-256: <code>None</code><br>Recorded bytes None; embedded UTF-8 bytes None</p><p>small-file limit exceeded</p></details><details id="S036"><summary>S036 · google--gemma-4-31B-it/license-page.html · exact</summary><p><a href="https://ai.google.dev/gemma/docs/gemma_4_license" target="_blank" rel="noopener noreferrer">https://ai.google.dev/gemma/docs/gemma_4_license</a></p><p class="small">Retrieved 2026-10-08T08:51:37.392512+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>03e776581e97dbadfa7b0f6442e240d20ec24bee65ab43be57b864a042a1e1c4</code><br>Computed UTF-8 SHA-256: <code>03e776581e97dbadfa7b0f6442e240d20ec24bee65ab43be57b864a042a1e1c4</code><br>Recorded bytes 259673; embedded UTF-8 bytes 259673</p><p><b>HTML visible text lines 237-310</b></p><pre>Apache License 2.0
Apache
License
Version
2.0
,
January
2004
http
:
//www.apache.org/licenses/
TERMS
AND
CONDITIONS
FOR
USE
,
REPRODUCTION
,
AND
DISTRIBUTION
1
.
Definitions
.
&quot;License&quot;
shall
mean
the
terms
and
conditions
for
use
,
reproduction
,
and
distribution
as
defined
by
Sections
1
through
9
of
this
document
.
&quot;Licensor&quot;
shall
mean
the
copyright
owner
or
entity
authorized
by
the
copyright
owner
that
is
granting
the
License
.
&quot;Legal Entity&quot;
shall
mean
the
union</pre></details><details id="S037"><summary>S037 · datasets/nvidia--Nemotron-Cascade-RL-SWE.json · exact</summary><p><a href="https://huggingface.co/api/datasets/nvidia/Nemotron-Cascade-RL-SWE" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/nvidia/Nemotron-Cascade-RL-SWE</a></p><p class="small">Retrieved 2026-10-08T08:51:39.202209+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>697dd6f74c7a445da219297eb077f7319674ff474df8acf02a8f083ad6239ece</code><br>Computed UTF-8 SHA-256: <code>697dd6f74c7a445da219297eb077f7319674ff474df8acf02a8f083ad6239ece</code><br>Recorded bytes 1504; embedded UTF-8 bytes 1504</p><p><b>Complete bounded dataset metadata JSON; reduced fields for S075–S077</b></p><pre>{
  &quot;_id&quot;: &quot;6940965afc6d2d1efc196962&quot;,
  &quot;id&quot;: &quot;nvidia/Nemotron-Cascade-RL-SWE&quot;,
  &quot;author&quot;: &quot;nvidia&quot;,
  &quot;sha&quot;: &quot;2c3c6c7cfcc1424160fe2227ad982c5acca34bc2&quot;,
  &quot;lastModified&quot;: &quot;2025-12-16T18:41:44.000Z&quot;,
  &quot;private&quot;: false,
  &quot;gated&quot;: false,
  &quot;disabled&quot;: false,
  &quot;tags&quot;: [
    &quot;language:en&quot;,
    &quot;license:cc-by-4.0&quot;,
    &quot;size_categories:100K&lt;n&lt;1M&quot;,
    &quot;format:json&quot;,
    &quot;modality:text&quot;,
    &quot;library:datasets&quot;,
    &quot;library:dask&quot;,
    &quot;library:polars&quot;,
    &quot;library:mlcroissant&quot;,
    &quot;arxiv:2512.13607&quot;,
    &quot;region:us&quot;,
    &quot;nvidia&quot;,
    &quot;reasoning&quot;,
    &quot;code&quot;,
    &quot;reinforcement learning&quot;
  ],
  &quot;description&quot;: &quot;\n\t\n\t\t\n\t\n\t\n\t\tDataset Description:\n\t\n\nThe Nemotron-Cascade-RL-SWE dataset is the RL training data for SWE code repairing task, consisting of SWE-Bench-Train, SWE-reBench, SWE-Smith, R2E-Gym/R2E-Gym-Subset and SWE-Fixer-Train. \nWe select the training data for SFT and RL stages based on its difficulty. \nAlso, to avoid data contamination, we exclude all instances originating from repositories present in the SWE-Bench_Verified evaluation dataset. \nWe create the prompts following the agentless mini… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-SWE.&quot;,
  &quot;downloads&quot;: 465,
  &quot;likes&quot;: 34,
  &quot;cardData&quot;: {
    &quot;license&quot;: &quot;cc-by-4.0&quot;,
    &quot;tags&quot;: [
      &quot;nvidia&quot;,
      &quot;reasoning&quot;,
      &quot;code&quot;,
      &quot;reinforcement learning&quot;
    ],
    &quot;language&quot;: [
      &quot;en&quot;
    ]
  },
  &quot;siblings&quot;: [
    {
      &quot;rfilename&quot;: &quot;.gitattributes&quot;
    },
    {
      &quot;rfilename&quot;: &quot;README.md&quot;
    },
    {
      &quot;rfilename&quot;: &quot;train_16k.jsonl&quot;
    },
    {
      &quot;rfilename&quot;: &quot;train_24k.jsonl&quot;
    },
    {
      &quot;rfilename&quot;: &quot;train_32k.jsonl&quot;
    }
  ],
  &quot;createdAt&quot;: &quot;2025-12-15T23:14:34.000Z&quot;,
  &quot;usedStorage&quot;: 33172765773
}</pre></details><details id="S038"><summary>S038 · nvidia--NVIDIA-Nemotron-3-Super-120B-A12B-BF16/license-page.html · exact (corrected line endings)</summary><p><a href="https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/" target="_blank" rel="noopener noreferrer">https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/</a></p><p class="small">Retrieved 2026-10-08T08:51:37.392838+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>33011b6a3b8865af22514e38be507780a2ff00761df5b924f9f398e56db1000e</code><br>Computed SHA-256: <code>33011b6a3b8865af22514e38be507780a2ff00761df5b924f9f398e56db1000e</code><br>Recorded bytes 300192; corrected UTF-8 bytes 300192<br>Source: supplied snapshot-integrity-corrections.json. Only line endings changed.</p><p><b>HTML visible text lines 667-722</b></p><pre>Download PDF
NVIDIA Nemotron Open Model License
Last Modified: December 15, 2025
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
NVIDIA Works released under this License are intended to be used permissively and enable the further development of AI technologies. Subject to the terms of this License, NVIDIA confirms that:
Works are commercially usable.
You are free to create and distribute Derivative Works.
NVIDIA does not claim ownership to any outputs generated using the Works or Derivative Works.
By using, reproducing, modifying, distributing, performing or displaying any portion or element of the Works or Derivative Works, or otherwise accepting the terms of this License, you agree to be bound by this License.
1. Definitions
.
“
License
” shall mean the terms and conditions for use, reproduction, and distribution as defined by Sections 1 through 10 of this document.
“
Legal Entity
” shall mean the union of the acting entity and all other entities that control, are controlled by, or are under common control with that entity. For the purposes of this definition, &quot;
control
&quot; means (i) the power, direct or indirect, to cause the direction or management of such entity, whether by contract or otherwise, or (ii) ownership of fifty percent (50%) or more of the outstanding shares, or (iii) beneficial ownership of such entity.
“
You
” (or “
Your
”) shall mean an individual or Legal Entity exercising permissions granted by this License.
“
Work
” shall mean the work of authorship, including machine learning model, software, checkpoints, learnt weights, algorithms, parameters, configuration files and documentation, made available under the License, as indicated by a copyright notice that is included in or attached to the work (an example is provided in the Appendix below).
“
Derivative Works
” shall mean any work, whether in source or object form, that is based on (or derived from) the Work and for which the editorial revisions, annotations, elaborations, or other modifications represent, as a whole, an original work of authorship. For the purposes of this License, Derivative Works shall not include works that remain separable from, or merely link (or bind by name) to the interfaces of, the Work and Derivative Works thereof.
2. Grant of License
. Subject to the terms and conditions of this License, NVIDIA hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable license to reproduce, prepare Derivative Works of, publicly display, publicly perform, sublicense, and distribute the Work and such Derivative Works in source or object form.
If You institute patent or copyright litigation against any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Work or an output from the Work constitutes direct or contributory patent or copyright infringement, then any licenses granted to You under this License for that Work shall terminate as of the date such litigation is filed.
3. Redistribution
. You may reproduce and distribute copies of the Work or Derivative Works thereof in any medium, with or without modifications, and in source or object form, provided that You meet the following conditions:
a. You must give any other recipients of the Work a copy of this License; and
b. You must retain, in the source form of any Derivative Works that You distribute, all copyright, patent, trademark, and attribution notices from the source form of the Work, excluding those notices that do not pertain to any part of the Derivative Works; and
c. If the Work includes a &quot;
NOTICE
&quot; text file as part of its distribution, then any Derivative Works that You distribute must include a readable copy of the following attribution notice within a “Notice” text file with such copies and the following statement: “Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.”
You may add Your own copyright statement to Your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of Your modifications, or for any such Derivative Works as a whole, provided Your use, reproduction, and distribution of the Work otherwise complies with the conditions stated in this License.
4. Trademarks
. This License does not grant permission to use the trade names, trademarks, service marks, or product names of NVIDIA, except as required for reasonable and customary use in describing the origin of the Work and reproducing the content of the NOTICE file.
5. Disclaimer of Warranty. Unless required by applicable law or agreed to in writing, NVIDIA provides the Work on an “AS IS” BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied, including, without limitation, any warranties or conditions of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A PARTICULAR PURPOSE. You are solely responsible for determining the appropriateness of using or redistributing the Work and assume any risks associated with Your exercise of permissions under this License.
6. Limitation of Liability. In no event and under no legal theory, whether in tort (including negligence), contract, or otherwise, unless required by applicable law (such as deliberate and grossly negligent acts) or agreed to in writing, shall NVIDIA be liable to You for damages, including any direct, indirect, special, incidental, or consequential damages of any character arising as a result of this License or out of the use or inability to use the Work, a Derivative Work or an output from the Work or Derivative Work (including but not limited to damages for loss of goodwill, work stoppage, computer failure or malfunction, or any and all other commercial damages or losses), even if NVIDIA has been advised of the possibility of such damages.
7.
Accepting Warranty or Additional Liability
. While redistributing the Work or Derivative Works thereof, You may choose to offer, and charge a fee for, acceptance of support, warranty, indemnity, or other liability obligations and/or rights consistent with this License. However, in accepting such obligations, You may act only on Your own behalf and on Your sole responsibility, not on behalf of NVIDIA, and only if You agree to indemnify, defend, and hold NVIDIA harmless for any liability incurred by, or claims asserted against, NVIDIA by reason of your accepting any such warranty or additional liability. You will indemnify and hold harmless NVIDIA from and against any claim by any third party arising out of or related to your use or distribution of the Works, Derivative Works thereof, or output from the Works or Derivative Works.
8. Feedback
. NVIDIA appreciates your feedback, and You agree that NVIDIA may use it without restriction or compensation to You.
9. Governing Law
. This Agreement will be governed in all respects by the laws of the United States and the laws of the State of Delaware, without regard to conflict of laws principles or the United Nations Convention on Contracts for the International Sale of Goods. The state and federal courts residing in Santa Clara County, California will have exclusive jurisdiction over any dispute or claim arising out of or related to this Agreement, and the parties irrevocably consent to personal jurisdiction and venue in those courts; except that, either party may apply for injunctive remedies or an equivalent type of urgent legal relief in any jurisdiction.
10. Trade and Compliance
. You agree to comply with all applicable export, import, trade and economic sanctions laws and regulations, as amended, including without limitation U.S. Export Administration Regulations and Office of Foreign Assets Control regulations. These laws include restrictions on destinations, end-users and end-use.
(v. December 15, 2025)
END OF TERMS AND CONDITIONS</pre></details><details id="S039"><summary>S039 · https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125 · no_content</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125</a></p><p class="small">Retrieved 2026-10-08T08:51:39.120038+00:00 · Recorded HTTP None<br>Recorded SHA-256: <code>None</code><br>Computed UTF-8 SHA-256: <code>None</code><br>Recorded bytes None; embedded UTF-8 bytes None</p><p>small-file limit exceeded</p></details><details id="S040"><summary>S040 · github/NVIDIA-NeMo--Evaluator/metadata.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Evaluator" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Evaluator</a></p><p class="small">Retrieved 2026-10-08T08:51:40.874052+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>7955fe2f705d90bc8c7d0b11e5f14a76caa0a465e4aea949de833c76d7fb7cc9</code><br>Computed UTF-8 SHA-256: <code>7955fe2f705d90bc8c7d0b11e5f14a76caa0a465e4aea949de833c76d7fb7cc9</code><br>Recorded bytes 6478; embedded UTF-8 bytes 6478</p></details><details id="S041"><summary>S041 · github/NVIDIA-NeMo--Nemotron/metadata.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Nemotron" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Nemotron</a></p><p class="small">Retrieved 2026-10-08T08:51:40.873876+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>f5461230db490a0176e2f08b74916e14df77f26c25b9e793c07c57fb130cf808</code><br>Computed UTF-8 SHA-256: <code>f5461230db490a0176e2f08b74916e14df77f26c25b9e793c07c57fb130cf808</code><br>Recorded bytes 6631; embedded UTF-8 bytes 6631</p></details><details id="S042"><summary>S042 · github/allenai--OLMo-Eval/metadata.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-Eval" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-Eval</a></p><p class="small">Retrieved 2026-10-08T08:51:40.873795+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>25e1da3b0e74c33880c367c5b992b9b7db2dc904a72cdc87b3f9933d0c1e29ee</code><br>Computed UTF-8 SHA-256: <code>25e1da3b0e74c33880c367c5b992b9b7db2dc904a72cdc87b3f9933d0c1e29ee</code><br>Recorded bytes 6115; embedded UTF-8 bytes 6115</p></details><details id="S043"><summary>S043 · github/allenai--OLMo-core/metadata.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-core" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-core</a></p><p class="small">Retrieved 2026-10-08T08:51:40.873699+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>047e681797a342d8920f9737ed778fc0c8b52f002fd52e75ab54e351212b5239</code><br>Computed UTF-8 SHA-256: <code>047e681797a342d8920f9737ed778fc0c8b52f002fd52e75ab54e351212b5239</code><br>Recorded bytes 6212; embedded UTF-8 bytes 6212</p></details><details id="S044"><summary>S044 · github/allenai--OLMo-core/commit.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-core/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-core/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:41.677506+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>f2cd23460210987d94fd59cd570707e1623a02ea8fc43d6c12d5c85e8f4220b9</code><br>Computed UTF-8 SHA-256: <code>f2cd23460210987d94fd59cd570707e1623a02ea8fc43d6c12d5c85e8f4220b9</code><br>Recorded bytes 47826; embedded UTF-8 bytes 47826</p></details><details id="S045"><summary>S045 · github/NVIDIA-NeMo--Nemotron/commit.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Nemotron/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Nemotron/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:41.648219+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>d1a7b2e9b45049bf9d17bcb12c629f994e56074463bfc06a856d3a72537e9015</code><br>Computed UTF-8 SHA-256: <code>d1a7b2e9b45049bf9d17bcb12c629f994e56074463bfc06a856d3a72537e9015</code><br>Recorded bytes 21443; embedded UTF-8 bytes 21443</p></details><details id="S046"><summary>S046 · github/NVIDIA-NeMo--Evaluator/commit.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Evaluator/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Evaluator/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:41.648017+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>a41091f855a734c550bd99b554ae6af8a4b61878adf846db39dc35e1b352c960</code><br>Computed UTF-8 SHA-256: <code>a41091f855a734c550bd99b554ae6af8a4b61878adf846db39dc35e1b352c960</code><br>Recorded bytes 21684; embedded UTF-8 bytes 21684</p></details><details id="S047"><summary>S047 · github/allenai--OLMo-Eval/commit.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-Eval/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-Eval/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:41.671577+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>7339e24bad39cf031d0e7e484b10852502ef28159380f003309be50dfd119465</code><br>Computed UTF-8 SHA-256: <code>7339e24bad39cf031d0e7e484b10852502ef28159380f003309be50dfd119465</code><br>Recorded bytes 64134; embedded UTF-8 bytes 64134</p></details><details id="S048"><summary>S048 · github/allenai--OLMo-core/tree.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:42.366282+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>e5410bbac4714f60b1bcd62963c6a5764a1ec93dfb6638657de7e6e3f42af156</code><br>Computed UTF-8 SHA-256: <code>e5410bbac4714f60b1bcd62963c6a5764a1ec93dfb6638657de7e6e3f42af156</code><br>Recorded bytes 211744; embedded UTF-8 bytes 211744</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;5f6f58a133e7ef577d596295f2c8db4651c27857&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 866,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;9b259bdfcf9022e7f4999a318ab7261300644e11&quot;,
      &quot;size&quot;: 11359,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/9b259bdfcf9022e7f4999a318ab7261300644e11&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ff43284c45b391ae7f512b8325616625acf5ff6b&quot;,
      &quot;size&quot;: 9139,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/ff43284c45b391ae7f512b8325616625acf5ff6b&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_core/data/mixes/OLMo-longmino-mix-0925.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;cbfbf6593dfaedeb97a8914ef50e664a3aa7d907&quot;,
      &quot;size&quot;: 66216,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/cbfbf6593dfaedeb97a8914ef50e664a3aa7d907&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_core/data/mixes/OLMo-midtraining-mix-0925-ingredient1-100B.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;f7521598e8d99d2a38d15375a193df2f702833ec&quot;,
      &quot;size&quot;: 2286778,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/f7521598e8d99d2a38d15375a193df2f702833ec&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_core/data/mixes/OLMo-midtraining-mix-0925-ingredient2-100B.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;f41d9ff33e42e2aa2813502ee9691f7ee35b55a0&quot;,
      &quot;size&quot;: 2286778,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/f41d9ff33e42e2aa2813502ee9691f7ee35b55a0&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_core/data/mixes/OLMo-mix-0925-official.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;5bdeee234a6f9bb9dcc4609b8c6aaffc40416f06&quot;,
      &quot;size&quot;: 110287,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/5bdeee234a6f9bb9dcc4609b8c6aaffc40416f06&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_core/data/mixes/OLMo-mix-0925.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a21ba8cd7d79697689f0cd4f2461b333b0a74d3a&quot;,
      &quot;size&quot;: 109962,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/a21ba8cd7d79697689f0cd4f2461b333b0a74d3a&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-32B-long-context.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;4cca05767922c206dffcde8007c00a57f299f255&quot;,
      &quot;size&quot;: 5467,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/4cca05767922c206dffcde8007c00a57f299f255&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-32B-midtrain-ingredient-1.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8d6b91d3bfa5d847edcf88569015655f01325229&quot;,
      &quot;size&quot;: 4975,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/8d6b91d3bfa5d847edcf88569015655f01325229&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-32B-midtrain-ingredient-2.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;eac7959fd2fa4f05f0997f16e21fd06cbfe4b63d&quot;,
      &quot;size&quot;: 4983,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/eac7959fd2fa4f05f0997f16e21fd06cbfe4b63d&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-32B-pretrain.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0edd0576f399a22efa48614fad562c556db1197c&quot;,
      &quot;size&quot;: 7151,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/0edd0576f399a22efa48614fad562c556db1197c&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-32B.csv&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8890d114e38e6ea14dc2966f9fab13debe18349f&quot;,
      &quot;size&quot;: 85002,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/8890d114e38e6ea14dc2966f9fab13debe18349f&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-7B-long-context.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;6abace2e52eb77effff46a89e179f105a09b0ac6&quot;,
      &quot;size&quot;: 5073,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/6abace2e52eb77effff46a89e179f105a09b0ac6&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-7B-midtrain.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;9dfa8bc721dcc9c3a49a6bf9391b14ea44da1349&quot;,
      &quot;size&quot;: 4864,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/9dfa8bc721dcc9c3a49a6bf9391b14ea44da1349&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-7B-pretrain-1.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;96bb51120ef29db485e15b2328814bef0b9234b8&quot;,
      &quot;size&quot;: 7004,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/96bb51120ef29db485e15b2328814bef0b9234b8&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-7B-pretrain-2.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0961ec3781a6f07689768049ef833b27c3692fbc&quot;,
      &quot;size&quot;: 7418,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/0961ec3781a6f07689768049ef833b27c3692fbc&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/OLMo-3-1025-7B.csv&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ca65edade87a5b971b0651066bcc1ab6d357f94a&quot;,
      &quot;size&quot;: 165430,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/ca65edade87a5b971b0651066bcc1ab6d357f94a&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/official/OLMo3/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2999e8a51555f4c593d150007c11c530e850888f&quot;,
      &quot;size&quot;: 8646,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/2999e8a51555f4c593d150007c11c530e850888f&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo-3-190M.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;066afa4093ff0d96e090e368671f36493f555196&quot;,
      &quot;size&quot;: 6748,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/066afa4093ff0d96e090e368671f36493f555196&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-32B-long-context.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2cf809e258aaffcd1de73b837bcff2cc74c9ae34&quot;,
      &quot;size&quot;: 6768,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/2cf809e258aaffcd1de73b837bcff2cc74c9ae34&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-32B-midtraining.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;1b7558e3390ab26fd488dc109b4e3459c392aeb5&quot;,
      &quot;size&quot;: 4976,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/1b7558e3390ab26fd488dc109b4e3459c392aeb5&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-32B.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;df9b83e2fb17792e2149f3069af05142f4d9e4f6&quot;,
      &quot;size&quot;: 6255,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/df9b83e2fb17792e2149f3069af05142f4d9e4f6&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-7B-anneal.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;6815073c522f1efb0e2ba14df5f64e9b07293269&quot;,
      &quot;size&quot;: 5798,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/6815073c522f1efb0e2ba14df5f64e9b07293269&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-7B-long-context.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;def8606c76fd243df23a1a6d2fc3b7a27c54918e&quot;,
      &quot;size&quot;: 4739,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/def8606c76fd243df23a1a6d2fc3b7a27c54918e&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-7B-midtraining.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ab51ec7403f970774b887655f938891cda308e3d&quot;,
      &quot;size&quot;: 4214,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/ab51ec7403f970774b887655f938891cda308e3d&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/OLMo3-7B.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;9e4d95d0e1618571805c53f3a088d1476afee801&quot;,
      &quot;size&quot;: 5654,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/9e4d95d0e1618571805c53f3a088d1476afee801&quot;
    },
    {
      &quot;path&quot;: &quot;src/scripts/train/OLMo3/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;35f467d94f9a1e15ca167034ac4926562f7a532e&quot;,
      &quot;size&quot;: 209,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/Olmo-core/git/blobs/35f467d94f9a1e15ca167034ac4926562f7a532e&quot;
    }
  ]
}</pre></details><details id="S049"><summary>S049 · github/NVIDIA-NeMo--Evaluator/tree.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/trees/c87f9b21769cd1a7f3b85267264101a7bcc77df6?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/trees/c87f9b21769cd1a7f3b85267264101a7bcc77df6?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:42.442151+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>315787c07d9fad7109cc30de92bd08e9828bc2eb10e4ec816fe1bd47cf624f2b</code><br>Computed UTF-8 SHA-256: <code>315787c07d9fad7109cc30de92bd08e9828bc2eb10e4ec816fe1bd47cf624f2b</code><br>Recorded bytes 154133; embedded UTF-8 bytes 154133</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;c87f9b21769cd1a7f3b85267264101a7bcc77df6&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 612,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;36b3651ff57fb67c692d40d57fb39badf8e3cb96&quot;,
      &quot;size&quot;: 10769,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/36b3651ff57fb67c692d40d57fb39badf8e3cb96&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;7543ce20baf4422b901c0b46fd8ccf55691c1d82&quot;,
      &quot;size&quot;: 7703,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/7543ce20baf4422b901c0b46fd8ccf55691c1d82&quot;
    },
    {
      &quot;path&quot;: &quot;packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/local_nemotron-3-super-120b-a12b-base.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;dad3c528f5fed54ca50c081f7f63d97fabdac042&quot;,
      &quot;size&quot;: 3795,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/dad3c528f5fed54ca50c081f7f63d97fabdac042&quot;
    },
    {
      &quot;path&quot;: &quot;packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/local_nemotron-3-super-120b-a12b.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;76a698a7a3bcd5e1f77e68a2f3378d48558b9672&quot;,
      &quot;size&quot;: 10942,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/76a698a7a3bcd5e1f77e68a2f3378d48558b9672&quot;
    },
    {
      &quot;path&quot;: &quot;packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/local_nemotron-3-super-120b-a12b_low_budget.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;41a5d7dd32ff1b7045e600f7f44986807e23c2f7&quot;,
      &quot;size&quot;: 5931,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/41a5d7dd32ff1b7045e600f7f44986807e23c2f7&quot;
    },
    {
      &quot;path&quot;: &quot;packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/local_nemotron-3-super-120b-a12b_tools.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8aa41df012581bab75eade27da6d2838461cc64f&quot;,
      &quot;size&quot;: 6543,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/8aa41df012581bab75eade27da6d2838461cc64f&quot;
    },
    {
      &quot;path&quot;: &quot;packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a998c4c7f9bebbd3ef3538c19209ad895c1f930d&quot;,
      &quot;size&quot;: 20663,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/blobs/a998c4c7f9bebbd3ef3538c19209ad895c1f930d&quot;
    }
  ]
}</pre></details><details id="S050"><summary>S050 · github/NVIDIA-NeMo--Nemotron/tree.json · exact</summary><p><a href="https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ca8c409f08a9a5a5648d427383adea741b61966a?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ca8c409f08a9a5a5648d427383adea741b61966a?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:42.431405+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>93f8ffd04b8e9312f4c0fc9a0cffc20982ab39237a57980481d6c5e5aa01d2ae</code><br>Computed UTF-8 SHA-256: <code>93f8ffd04b8e9312f4c0fc9a0cffc20982ab39237a57980481d6c5e5aa01d2ae</code><br>Recorded bytes 555167; embedded UTF-8 bytes 555167</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;ca8c409f08a9a5a5648d427383adea741b61966a&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 2154,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;d645695673349e3947e8e5ae42332d0ac3164cd7&quot;,
      &quot;size&quot;: 11358,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/d645695673349e3947e8e5ae42332d0ac3164cd7&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;b0d81964d646babddc430c36b395d324949bd8ce&quot;,
      &quot;size&quot;: 42400,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/b0d81964d646babddc430c36b395d324949bd8ce&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;84d9c5b58c0eca70ff280459607bead44a40c882&quot;,
      &quot;size&quot;: 8508,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/84d9c5b58c0eca70ff280459607bead44a40c882&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/__init__.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;7b89ac868c25d74775b95f0b7db6f3a43b4880c0&quot;,
      &quot;size&quot;: 641,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/7b89ac868c25d74775b95f0b7db6f3a43b4880c0&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;c15f10e56d83486753e3dc44836e5e42d0e3de40&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/c15f10e56d83486753e3dc44836e5e42d0e3de40&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;896ce47cf0671513019547399cf53f9037ec68d8&quot;,
      &quot;size&quot;: 8028,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/896ce47cf0671513019547399cf53f9037ec68d8&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/__init__.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;,
      &quot;size&quot;: 610,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;4d0382e7d8c83ab930036445bdf3fdd69d06df82&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/4d0382e7d8c83ab930036445bdf3fdd69d06df82&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;b78a171af33c3c5712486b6b5e8ea607f34ebc91&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/b78a171af33c3c5712486b6b5e8ea607f34ebc91&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/data_blend_cache_test.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;,
      &quot;size&quot;: 16,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/data_blend_raw_long_context.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ccfd07feda35bbed123b0ab496965e3a4f874490&quot;,
      &quot;size&quot;: 6683,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ccfd07feda35bbed123b0ab496965e3a4f874490&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/data_blend_raw_phase1.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a2b1e8906a0ddd9cbe1608b99756a1e63d5be3bb&quot;,
      &quot;size&quot;: 7044,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/a2b1e8906a0ddd9cbe1608b99756a1e63d5be3bb&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/data_blend_raw_phase2.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;66946adf9ce26b6a125a1331b7be1372572f7b48&quot;,
      &quot;size&quot;: 6602,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/66946adf9ce26b6a125a1331b7be1372572f7b48&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/data_blend_raw_small.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;,
      &quot;size&quot;: 16,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;1abd23356c5d5a1552c9230f9b6eb0962f7c99d8&quot;,
      &quot;size&quot;: 2030,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/1abd23356c5d5a1552c9230f9b6eb0962f7c99d8&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/long_context.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;3ac6c41a201ca2dddd762ea1c1e7ad24dff48cec&quot;,
      &quot;size&quot;: 1959,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/3ac6c41a201ca2dddd762ea1c1e7ad24dff48cec&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/phase1.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;79c8db3c0de295626fc0580425b992b61c2c13f3&quot;,
      &quot;size&quot;: 1948,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/79c8db3c0de295626fc0580425b992b61c2c13f3&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/phase2.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;016f64b2bcb252ebf3946a165d2a56513f3cb659&quot;,
      &quot;size&quot;: 2045,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/016f64b2bcb252ebf3946a165d2a56513f3cb659&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/data_prep/tiny.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;070a06322888fdac055fe35387f388d6689f690e&quot;,
      &quot;size&quot;: 1957,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/070a06322888fdac055fe35387f388d6689f690e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;4a28dfa253828c88b4933126e7f00fd24ad1db50&quot;,
      &quot;size&quot;: 318,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/4a28dfa253828c88b4933126e7f00fd24ad1db50&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/long_context_1m.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a3aa429354b8f2d793bb073c06a20345b98753c9&quot;,
      &quot;size&quot;: 1796,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/a3aa429354b8f2d793bb073c06a20345b98753c9&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/long_context_mixed.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;d15d175ecfc587c89a771947d1a555e84f7f2d7f&quot;,
      &quot;size&quot;: 1622,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/d15d175ecfc587c89a771947d1a555e84f7f2d7f&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/phase1.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;cb6d58e2d3a3a23d7b5ee307770ceb0b33b31faa&quot;,
      &quot;size&quot;: 1182,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/cb6d58e2d3a3a23d7b5ee307770ceb0b33b31faa&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/phase2.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2b2877e486ff17d7020334f8d1ab968c5a44b4fa&quot;,
      &quot;size&quot;: 1185,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/2b2877e486ff17d7020334f8d1ab968c5a44b4fa&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/config/test.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;dcdfc4785021b31c81fba052d67a5edac5318e8e&quot;,
      &quot;size&quot;: 714,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/dcdfc4785021b31c81fba052d67a5edac5318e8e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c0555157b6189b4af5f7c38265c78c72a25fb166&quot;,
      &quot;size&quot;: 12741,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c0555157b6189b4af5f7c38265c78c72a25fb166&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/test_train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;284c76d9492e3c055f707d32a48853c8ee4d4549&quot;,
      &quot;size&quot;: 3947,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/284c76d9492e3c055f707d32a48853c8ee4d4549&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage0_pretrain/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;4289c75ee31b650b2aba3593de886df8a70b0bec&quot;,
      &quot;size&quot;: 7837,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/4289c75ee31b650b2aba3593de886df8a70b0bec&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;dabcaad4e6e9347fea9b7ceb89294358d46dd368&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/dabcaad4e6e9347fea9b7ceb89294358d46dd368&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;1d4115fa9bc3cd5cb31f6173982b3c9fc96549ca&quot;,
      &quot;size&quot;: 8234,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/1d4115fa9bc3cd5cb31f6173982b3c9fc96549ca&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/__init__.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;,
      &quot;size&quot;: 610,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;18c617b8640b05ab9a44a2e3683661a4c6ecd992&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/18c617b8640b05ab9a44a2e3683661a4c6ecd992&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;8a1ba341556e55cf0c71dc88dc45e5f40ba6fbca&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/8a1ba341556e55cf0c71dc88dc45e5f40ba6fbca&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/data_prep/data_blend_raw.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2b359964a47816bdf9bd46a866c97d1cd44cc7fc&quot;,
      &quot;size&quot;: 3787,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/2b359964a47816bdf9bd46a866c97d1cd44cc7fc&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/data_prep/data_blend_tiny.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;,
      &quot;size&quot;: 16,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c8fb48831beae0cba68c1b4442e76d1c56489bd4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e7f8eaf861962d9d41326d69620cc4f6060944f8&quot;,
      &quot;size&quot;: 2824,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/e7f8eaf861962d9d41326d69620cc4f6060944f8&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/data_prep/tiny.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;51ac7651fc8a1e2ebdddf4890f8b357f14774a03&quot;,
      &quot;size&quot;: 2797,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/51ac7651fc8a1e2ebdddf4890f8b357f14774a03&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;cad8e48c35384c350ed222d6d3a9958a5e0fe0ef&quot;,
      &quot;size&quot;: 1465,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/cad8e48c35384c350ed222d6d3a9958a5e0fe0ef&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/config/test.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0696842dfff58e2e0bfd819ef789c5b32b50a5e4&quot;,
      &quot;size&quot;: 1022,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/0696842dfff58e2e0bfd819ef789c5b32b50a5e4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;860ce1632658aeca08740831a5b95c942ac82c63&quot;,
      &quot;size&quot;: 18676,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/860ce1632658aeca08740831a5b95c942ac82c63&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/test_train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ea47f4249bea879ad735744f276ece859752eb9d&quot;,
      &quot;size&quot;: 3840,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ea47f4249bea879ad735744f276ece859752eb9d&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage1_sft/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0540ec284d8cc633a0086cc188c22ec55b7c1221&quot;,
      &quot;size&quot;: 17078,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/0540ec284d8cc633a0086cc188c22ec55b7c1221&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;bce6cc326463f3073ba1282b8048fa63051eca00&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/bce6cc326463f3073ba1282b8048fa63051eca00&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0c6171986469d18925a5243686d0234ad3afd2aa&quot;,
      &quot;size&quot;: 7187,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/0c6171986469d18925a5243686d0234ad3afd2aa&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/__init__.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;,
      &quot;size&quot;: 610,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/341a77c5bc66dee5d2ba0edf888f91e5bf225e3c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/_data_prep_base.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e38040a44c2d702f62b11d1cb3798c75c03058c4&quot;,
      &quot;size&quot;: 8020,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/e38040a44c2d702f62b11d1cb3798c75c03058c4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;0b039b7dd5db011617e58b13db1c0d26a0eef380&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/0b039b7dd5db011617e58b13db1c0d26a0eef380&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;5403df698e3eda830f4b056fb838c39d282ef2a4&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/5403df698e3eda830f4b056fb838c39d282ef2a4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/data_prep/data_blend_raw.json&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2388dfc986d0bdb884138f51ccec76a551a4e947&quot;,
      &quot;size&quot;: 169,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/2388dfc986d0bdb884138f51ccec76a551a4e947&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;69c10f01c6faa68bbb401bfc37fba5d5da5cd8a3&quot;,
      &quot;size&quot;: 1221,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/69c10f01c6faa68bbb401bfc37fba5d5da5cd8a3&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/data_prep/tiny.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;b999c0710bc066b1ff1ba76cec57a89e5e3f7a01&quot;,
      &quot;size&quot;: 593,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/b999c0710bc066b1ff1ba76cec57a89e5e3f7a01&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e3327654fdb1018250deece0d650a72307221867&quot;,
      &quot;size&quot;: 11675,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/e3327654fdb1018250deece0d650a72307221867&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/test.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c95397b33a74928b1595d80bd6673ee39d3b0601&quot;,
      &quot;size&quot;: 942,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c95397b33a74928b1595d80bd6673ee39d3b0601&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/config/tiny.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0c4c0b2d4681b1f32c7ae21a611c946c91b5bc47&quot;,
      &quot;size&quot;: 787,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/0c4c0b2d4681b1f32c7ae21a611c946c91b5bc47&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;542d0745e789872f185d023885321543111022ce&quot;,
      &quot;size&quot;: 16547,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/542d0745e789872f185d023885321543111022ce&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;2f7c75d6cbb9c7fa21e967724c74d38bc8416b52&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/2f7c75d6cbb9c7fa21e967724c74d38bc8416b52&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;d9d2d9e4a4bb8b91db455c639f04fe2dcfc30f24&quot;,
      &quot;size&quot;: 7097,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/d9d2d9e4a4bb8b91db455c639f04fe2dcfc30f24&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;ecd0ea373ffb95da991726c357248daa271382ce&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ecd0ea373ffb95da991726c357248daa271382ce&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;ad1932ac374c5043921fce28332a9c9bbf8f0375&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ad1932ac374c5043921fce28332a9c9bbf8f0375&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/data_prep/rlvr1.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;678ca3a0c65beed79fe56d775caf01f78699982b&quot;,
      &quot;size&quot;: 853,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/678ca3a0c65beed79fe56d775caf01f78699982b&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/data_prep/rlvr2.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;b4d00f3efffce4a11f75a80edbcadbaa465b4bd6&quot;,
      &quot;size&quot;: 794,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/b4d00f3efffce4a11f75a80edbcadbaa465b4bd6&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/data_prep/rlvr3.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;20d2a184a08e1a713c45bae4a15769907b1d4517&quot;,
      &quot;size&quot;: 779,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/20d2a184a08e1a713c45bae4a15769907b1d4517&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ecf742832a75e6c8dcb37f22df6cf603fea1952c&quot;,
      &quot;size&quot;: 17611,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ecf742832a75e6c8dcb37f22df6cf603fea1952c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/rlvr2.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;deb88dcedd415d20ef481292e3a7ff7d3f992505&quot;,
      &quot;size&quot;: 102,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/deb88dcedd415d20ef481292e3a7ff7d3f992505&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/rlvr3.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;07becdd8959fe3a6eb90e78570292eaa482e2c0b&quot;,
      &quot;size&quot;: 102,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/07becdd8959fe3a6eb90e78570292eaa482e2c0b&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/small.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;410e0cde66f992eb0072fd08c2798b9aaba0973e&quot;,
      &quot;size&quot;: 1038,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/410e0cde66f992eb0072fd08c2798b9aaba0973e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/config/test.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;c387c77127c27e2e63d0b7c6159704b3e6a62d12&quot;,
      &quot;size&quot;: 919,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/c387c77127c27e2e63d0b7c6159704b3e6a62d12&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ee99c16b8d4a585060f63b7f412f9a69d2472a9c&quot;,
      &quot;size&quot;: 2431,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ee99c16b8d4a585060f63b7f412f9a69d2472a9c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage1_rlvr/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8f0a6fa5f13cdc4500d15e0df01f420bd0d6ee17&quot;,
      &quot;size&quot;: 16286,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/8f0a6fa5f13cdc4500d15e0df01f420bd0d6ee17&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;bc9fce22197ab10245be3b527aa07213e2976fee&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/bc9fce22197ab10245be3b527aa07213e2976fee&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2a1b46b5f0a992662f54181b6dd96bb106dba22e&quot;,
      &quot;size&quot;: 4216,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/2a1b46b5f0a992662f54181b6dd96bb106dba22e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;e6d48ef44293c2759c86bbf58d627d29288457a5&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/e6d48ef44293c2759c86bbf58d627d29288457a5&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;f3a80253695ec0c0202c37c91e07d58278ad0681&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/f3a80253695ec0c0202c37c91e07d58278ad0681&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;4ca2f6973f329acad0bfb389031e8b3d783ab023&quot;,
      &quot;size&quot;: 755,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/4ca2f6973f329acad0bfb389031e8b3d783ab023&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;f7c69ae244ebde70fab896329de33b1b8bc6e570&quot;,
      &quot;size&quot;: 9584,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/f7c69ae244ebde70fab896329de33b1b8bc6e570&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/config/small.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;3002d3cd1f298257735c6772f336c2510d0ecdb2&quot;,
      &quot;size&quot;: 697,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/3002d3cd1f298257735c6772f336c2510d0ecdb2&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ead23afd4c19ed5986720bc16a14bb507b2e50f9&quot;,
      &quot;size&quot;: 2130,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ead23afd4c19ed5986720bc16a14bb507b2e50f9&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe1/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;995309a71a7380039e4877495592692353437479&quot;,
      &quot;size&quot;: 15745,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/995309a71a7380039e4877495592692353437479&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;e5b3c27933c19794b7860bf5021342af1fc414a0&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/e5b3c27933c19794b7860bf5021342af1fc414a0&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;17809745c4513eb74fa5a6f32569df25b901edcf&quot;,
      &quot;size&quot;: 5511,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/17809745c4513eb74fa5a6f32569df25b901edcf&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;12ba3a7b0d1bfdea36088778f8e2f11980427581&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/12ba3a7b0d1bfdea36088778f8e2f11980427581&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;655d0e178aed9fde9eb9b49d8c338e2786062e09&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/655d0e178aed9fde9eb9b49d8c338e2786062e09&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;be4cbfca77a929b85afadbebb8a5c63ace76a259&quot;,
      &quot;size&quot;: 754,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/be4cbfca77a929b85afadbebb8a5c63ace76a259&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8423a5bbcda176bed3ae02d6581d61fb9676b6de&quot;,
      &quot;size&quot;: 10504,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/8423a5bbcda176bed3ae02d6581d61fb9676b6de&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;f188224a5e8878ad5eba55df31a34fee237b343f&quot;,
      &quot;size&quot;: 2144,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/f188224a5e8878ad5eba55df31a34fee237b343f&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage2_swe2/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;aef20d49543929c85a2164cabcc2ab90d1f0e124&quot;,
      &quot;size&quot;: 15745,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/aef20d49543929c85a2164cabcc2ab90d1f0e124&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;f5a9b534c348507f7af2d3bddb5752e8720485fe&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/f5a9b534c348507f7af2d3bddb5752e8720485fe&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;964fbe9ece945bdc13cb05a232bc1e291a5cdd2d&quot;,
      &quot;size&quot;: 3666,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/964fbe9ece945bdc13cb05a232bc1e291a5cdd2d&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;2f002cd87cb727b4a8f8b77405de367a3e950df4&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/2f002cd87cb727b4a8f8b77405de367a3e950df4&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/config/data_prep&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;f6fba74a36d3eab380991ae984be0b2c3418911d&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/f6fba74a36d3eab380991ae984be0b2c3418911d&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/config/data_prep/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;44d971be048f2de8e681271007062be84c20b3bb&quot;,
      &quot;size&quot;: 747,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/44d971be048f2de8e681271007062be84c20b3bb&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ef46758a27a71960d422fb0e81b68ebd3cb3164c&quot;,
      &quot;size&quot;: 12324,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/ef46758a27a71960d422fb0e81b68ebd3cb3164c&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/config/small.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e663a70d38f11e94c99eee6eabf1ae71f154d52e&quot;,
      &quot;size&quot;: 592,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/e663a70d38f11e94c99eee6eabf1ae71f154d52e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/data_prep.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;deecd0266cad6916aa7d8c1f12dcb4ff4de67a4e&quot;,
      &quot;size&quot;: 2116,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/deecd0266cad6916aa7d8c1f12dcb4ff4de67a4e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/stage3_rlhf/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;7ed89c28d9855d535ea35c0bd7ce24b8a3c3ce9d&quot;,
      &quot;size&quot;: 15748,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/7ed89c28d9855d535ea35c0bd7ce24b8a3c3ce9d&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/test_train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;938e10a6a4eaa0d9a8838c7c29858ed5a653eaf5&quot;,
      &quot;size&quot;: 14204,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/938e10a6a4eaa0d9a8838c7c29858ed5a653eaf5&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage2_rl/train.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;133601320904137a98aa772934caaab047ed1742&quot;,
      &quot;size&quot;: 15687,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/133601320904137a98aa772934caaab047ed1742&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage3_eval&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;885b94d3db2fb34bc3636e6ab43e41004514bd05&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/885b94d3db2fb34bc3636e6ab43e41004514bd05&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage3_eval/README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;3dc07683281d91cefe6fd0514ea7154734050c2b&quot;,
      &quot;size&quot;: 3154,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/3dc07683281d91cefe6fd0514ea7154734050c2b&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage3_eval/config&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;ee49163303ed57506534a15ee5ac101bd89c3a8e&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ee49163303ed57506534a15ee5ac101bd89c3a8e&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/stage3_eval/config/default.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;2505feaff042d732eec751cde3532646162eb3a7&quot;,
      &quot;size&quot;: 6493,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/2505feaff042d732eec751cde3532646162eb3a7&quot;
    },
    {
      &quot;path&quot;: &quot;src/nemotron/recipes/super3/tiny_model.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;228c37424c8445a53b02614e8c65e5d4b93add7c&quot;,
      &quot;size&quot;: 4781,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/228c37424c8445a53b02614e8c65e5d4b93add7c&quot;
    },
    {
      &quot;path&quot;: &quot;tests/recipes/super3/__init__.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e69de29bb2d1d6434b8b29ae775ad8c2e48c5391&quot;,
      &quot;size&quot;: 0,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391&quot;
    },
    {
      &quot;path&quot;: &quot;tests/recipes/super3/run_integration_test.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;654ed3713223a33f8e7e26c1463cbf434315e95c&quot;,
      &quot;size&quot;: 6951,
      &quot;url&quot;: &quot;https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/blobs/654ed3713223a33f8e7e26c1463cbf434315e95c&quot;
    }
  ]
}</pre></details><details id="S051"><summary>S051 · github/allenai--OLMo-Eval/tree.json · exact</summary><p><a href="https://api.github.com/repos/allenai/OLMo-Eval/git/trees/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/allenai/OLMo-Eval/git/trees/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:42.595357+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>413670285549f7fd2d76f6775f96c2b81992cf6b4d10ce74f926883d43a041c6</code><br>Computed UTF-8 SHA-256: <code>413670285549f7fd2d76f6775f96c2b81992cf6b4d10ce74f926883d43a041c6</code><br>Recorded bytes 206401; embedded UTF-8 bytes 206401</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 840,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;261eeb9e9f8b2b4b0d119366dda99c6fd7d35c64&quot;,
      &quot;size&quot;: 11357,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/261eeb9e9f8b2b4b0d119366dda99c6fd7d35c64&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;535b12866bc53e883399b1a3ebf69d4318f73c2b&quot;,
      &quot;size&quot;: 70331,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/535b12866bc53e883399b1a3ebf69d4318f73c2b&quot;
    },
    {
      &quot;path&quot;: &quot;scripts/baselines/launch_olmo3_baselines_0426.sh&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;b2a81b9e1dfba5cd55fd64f53728ccda8301f33b&quot;,
      &quot;size&quot;: 22536,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/b2a81b9e1dfba5cd55fd64f53728ccda8301f33b&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_eval/evals/suites/olmobase.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0e4e6e594c51bf5458ed174339b9fdb4f32a8950&quot;,
      &quot;size&quot;: 8172,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/0e4e6e594c51bf5458ed174339b9fdb4f32a8950&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_eval/inference/patches/olmo3_tool_parser_patch.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;851dfa134ab2a199aed4e859c7ed340b3d560b6d&quot;,
      &quot;size&quot;: 5887,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/851dfa134ab2a199aed4e859c7ed340b3d560b6d&quot;
    },
    {
      &quot;path&quot;: &quot;src/olmo_eval/inference/templates/tool_chat_template_olmo3.jinja&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;cb78f5f2c912ccccf6ce204f495364e9406ba62b&quot;,
      &quot;size&quot;: 3282,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/cb78f5f2c912ccccf6ce204f495364e9406ba62b&quot;
    },
    {
      &quot;path&quot;: &quot;tests/evals/suites/test_olmobase.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;f70a0b59049d002749a7c775979e6605e6fbca9b&quot;,
      &quot;size&quot;: 2423,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/f70a0b59049d002749a7c775979e6605e6fbca9b&quot;
    },
    {
      &quot;path&quot;: &quot;tests/inference/patches/test_olmo3_tool_parser_patch.py&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;9b846bc4df8c11e3430c0f8de6dfb9dbb6b42f32&quot;,
      &quot;size&quot;: 5613,
      &quot;url&quot;: &quot;https://api.github.com/repos/allenai/olmo-eval/git/blobs/9b846bc4df8c11e3430c0f8de6dfb9dbb6b42f32&quot;
    }
  ]
}</pre></details><details id="S052"><summary>S052 · github/allenai--OLMo-core/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:43.293197+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>caa8f395bfaeb213fe94ca04084cc9e6536fc319ee35415677969263579d3006</code><br>Computed UTF-8 SHA-256: <code>caa8f395bfaeb213fe94ca04084cc9e6536fc319ee35415677969263579d3006</code><br>Recorded bytes 9139; embedded UTF-8 bytes 9139</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 1-40</b></p><pre>&lt;div align=&quot;center&quot;&gt;
  &lt;!-- &lt;img src=&quot;https://github.com/allenai/OLMo/assets/8812459/774ac485-a535-4768-8f7c-db7be20f5cc3&quot; width=&quot;300&quot;/&gt; --&gt;
  &lt;img src=&quot;https://huggingface.co/datasets/allenai/blog-images/resolve/main/olmo2/olmo.png&quot; alt=&quot;OLMo Logo&quot; width=&quot;280&quot; style=&quot;margin-left:&#x27;auto&#x27; margin-right:&#x27;auto&#x27; display:&#x27;block&#x27;&quot;/&gt;
  &lt;br&gt;
  &lt;h1&gt;Olmo-core&lt;/h1&gt;
  &lt;h4&gt;Building blocks for OLMo modeling and training&lt;/h4&gt;
&lt;/div&gt;
&lt;p align=&quot;center&quot;&gt;
  &lt;a href=&quot;https://olmo-core.readthedocs.io/en/latest/&quot;&gt;
    &lt;img alt=&quot;Docs&quot; src=&quot;https://img.shields.io/badge/API-docs-red&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://github.com/allenai/Olmo-core/tree/main/src/examples&quot;&gt;
    &lt;img alt=&quot;Examples&quot; src=&quot;https://img.shields.io/badge/API-examples-994B00&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://github.com/allenai/Olmo-core/releases/tag/v1.9.0&quot;&gt;
    &lt;img alt=&quot;Pypi&quot; src=&quot;https://img.shields.io/pypi/v/ai2-olmo-core.svg&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://github.com/allenai/Olmo-core/blob/main/LICENSE&quot;&gt;
    &lt;img alt=&quot;GitHub License&quot; src=&quot;https://img.shields.io/github/license/allenai/OLMo&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://arxiv.org/pdf/2501.00656.pdf&quot;&gt;
    &lt;img alt=&quot;Paper URL&quot; src=&quot;https://img.shields.io/badge/arxiv-2402.00838-orange&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://playground.allenai.org&quot;&gt;
    &lt;img alt=&quot;Playground&quot; src=&quot;https://img.shields.io/badge/Ai2-Playground-F0529C&quot;&gt;
  &lt;/a&gt;
  &lt;a href=&quot;https://discord.gg/sZq3jTNVNG&quot;&gt;
    &lt;img alt=&quot;Discord&quot; src=&quot;https://img.shields.io/badge/Discord%20-%20blue?style=flat&amp;logo=discord&amp;label=Ai2&amp;color=%235B65E9&quot;&gt;
  &lt;/a&gt;
&lt;/p&gt;

## Installation

First install [PyTorch](https://pytorch.org) according to the instructions specific to your operating system and hardware.

For development, we recommend installing from source:

```bash
git clone https://github.com/allenai/Olmo-core.git
cd Olmo-core</pre></details><details id="S053"><summary>S053 · github/NVIDIA-NeMo--Evaluator/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/c87f9b21769cd1a7f3b85267264101a7bcc77df6/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/c87f9b21769cd1a7f3b85267264101a7bcc77df6/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:43.337245+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>268ad059874e48298a32caf2754569e6cace4bb6a0df2d22fca73f3b738d16f8</code><br>Computed UTF-8 SHA-256: <code>268ad059874e48298a32caf2754569e6cace4bb6a0df2d22fca73f3b738d16f8</code><br>Recorded bytes 7703; embedded UTF-8 bytes 7703</p><p class="small">Git blob SHA-1 matches tree: True</p></details><details id="S054"><summary>S054 · github/NVIDIA-NeMo--Nemotron/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:43.373455+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>78498606809ad92f35b6686e1733ed858d011b5cbd89d06f98308ca54f9d91b8</code><br>Computed UTF-8 SHA-256: <code>78498606809ad92f35b6686e1733ed858d011b5cbd89d06f98308ca54f9d91b8</code><br>Recorded bytes 42400; embedded UTF-8 bytes 42400</p><p class="small">Git blob SHA-1 matches tree: True</p></details><details id="S055"><summary>S055 · github/allenai--OLMo-Eval/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-Eval/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-Eval/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:43.479037+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>50912cdb95b57717602f3c1abf921e57551741cc9741e6ace77ade90ddc4b6c8</code><br>Computed UTF-8 SHA-256: <code>50912cdb95b57717602f3c1abf921e57551741cc9741e6ace77ade90ddc4b6c8</code><br>Recorded bytes 70331; embedded UTF-8 bytes 70331</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 19-63</b></p><pre>## Quick Start

This project uses [uv](https://docs.astral.sh/uv/) with a checked-in `uv.lock`
for reproducible builds. To get started, sync the repo with `uv`, browse the
available tasks and suites, and preview a run with the built-in `mock` provider.

### Run Your First Eval

```bash
# Install uv if not already installed
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install Python 3.12 if your machine does not already have it
uv python install 3.12

# Install dependencies + the package (editable) from the lockfile.
# The default groups (`dev` + `vllm`) are installed automatically, which
# pulls in storage, beaker, hf, and the vLLM inference provider. vLLM
# deps are marked Linux-only via PEP 508 markers, so this works on macOS
# too — no extra flags needed.
uv sync --frozen

# Install pre-commit hooks
make setup

# To update the lockfile after changing pyproject.toml
uv lock

# Add an optional extra on top of the defaults (e.g. agents, litellm)
uv sync --frozen --extra agents

# `openhands` conflicts with vllm — opt out of the vllm group when using it
uv sync --frozen --no-group vllm --extra openhands

# Browse a few suites
uv run olmo-eval suite inspect mmlu
uv run olmo-eval suite inspect gpqa
uv run olmo-eval suite inspect olmobase:code

# Preview a run without loading a model
uv run olmo-eval run -m mock -t gsm8k --dry-run

# Preview another run with a different task spec
uv run olmo-eval run -m mock -t humaneval:3shot:bpb --dry-run
```</pre><p><b>content lines 78-124</b></p><pre>### Tasks

Tasks define how to load data, format prompts, and score outputs. Register with `@register`:

```python
from olmo_eval.evals.tasks.common import Task, register
from olmo_eval.data import DataSource

@register(&quot;my_task&quot;)
class MyTask(Task):
    # DataSource specifies path, subset (optional), and split
    data_source = DataSource(path=&quot;cais/mmlu&quot;, subset=&quot;abstract_algebra&quot;, split=&quot;test&quot;)
    ...
```

**Variants** can also act as named evaluation presets (for example, few-shot settings):

```python
from olmo_eval.evals.tasks.common import register_variant

register_variant(&quot;my_task&quot;, &quot;3shot&quot;, num_fewshot=3, fewshot_seed=42)
# Built-in example: uv run olmo-eval run -m llama3.1-8b -t humaneval:3shot:bpb
```

**Runtime Dependencies** allow tasks to specify packages installed at job startup:

```python
@register(&quot;code_eval&quot;)
class CodeEvalTask(Task):
    data_source = DataSource(path=&quot;my-org/code-dataset&quot;, split=&quot;test&quot;)
    dependencies = [&quot;code-sandbox==1.0&quot;, &quot;git+https://github.com/user/repo@v2.0&quot;]
    ...
```

### Suites

Suites group multiple tasks for batch evaluation:

```python
from olmo_eval.evals.suites import Suite, register

register(Suite(
    name=&quot;my_suite&quot;,
    tasks=(&quot;task_a:3shot&quot;, &quot;task_b:3shot&quot;, &quot;task_c:3shot&quot;),
))
```
</pre><p><b>content lines 230-243</b></p><pre>### Model Presets

Pre-configured model settings in `olmo_eval/common/constants/models.py`:

```python
from olmo_eval.common.constants import get_model_presets

# Returns dict of preset name -&gt; ModelConfig
presets = get_model_presets()
# {
#     &quot;llama3.1-8b&quot;: ModelConfig(model=&quot;meta-llama/Meta-Llama-3.1-8B&quot;),
#     &quot;olmo-2-7b&quot;: ModelConfig(model=&quot;allenai/OLMo-2-1124-7B&quot;),
#     ...
# }</pre><p><b>content lines 1041-1057</b></p><pre>## Advanced Usage

### Multi-GPU and Tool-Augmented Evaluation

```bash
# Basic evaluation
uv run olmo-eval run -m llama3.1-8b -t mmlu -t gsm8k -t arc_easy

# Large models with multi-GPU tensor parallelism
uv run olmo-eval run -m llama3.1-70b -t mmlu --num-gpus 4

# Refresh Hugging Face cache before loading a remote model
uv run olmo-eval run -m allenai/OLMo-2-1124-7B -t mmlu --force-download-model

# Tool-augmented evaluation with harness
uv run olmo-eval run -m llama3.1-8b -t simpleqa:judge --harness dr_tulu
```</pre></details><details id="S056"><summary>S056 · github/zai-org--GLM-5/metadata.json · exact</summary><p><a href="https://api.github.com/repos/zai-org/GLM-5" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/zai-org/GLM-5</a></p><p class="small">Retrieved 2026-10-08T08:51:43.864328+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>1a381b038d9999b7500d8d3503d5fd627887022470cc0faf171e209cbd5357e3</code><br>Computed UTF-8 SHA-256: <code>1a381b038d9999b7500d8d3503d5fd627887022470cc0faf171e209cbd5357e3</code><br>Recorded bytes 6021; embedded UTF-8 bytes 6021</p></details><details id="S057"><summary>S057 · github/MoonshotAI--Kimi-K3/metadata.json · exact</summary><p><a href="https://api.github.com/repos/MoonshotAI/Kimi-K3" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/MoonshotAI/Kimi-K3</a></p><p class="small">Retrieved 2026-10-08T08:51:43.833250+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>c6c21b407d42efd5246c7008f63e08c3244f0c8058e00a387a8afb0ce0d2a184</code><br>Computed UTF-8 SHA-256: <code>c6c21b407d42efd5246c7008f63e08c3244f0c8058e00a387a8afb0ce0d2a184</code><br>Recorded bytes 6167; embedded UTF-8 bytes 6167</p></details><details id="S058"><summary>S058 · github/QwenLM--Qwen3.8/metadata.json · exact</summary><p><a href="https://api.github.com/repos/QwenLM/Qwen3.8" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/QwenLM/Qwen3.8</a></p><p class="small">Retrieved 2026-10-08T08:51:43.928142+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>8554e6f8f5ed2d685ad3208c518994ffc96baa873e1d08a9f000ceaa116d4fd2</code><br>Computed UTF-8 SHA-256: <code>8554e6f8f5ed2d685ad3208c518994ffc96baa873e1d08a9f000ceaa116d4fd2</code><br>Recorded bytes 6023; embedded UTF-8 bytes 6023</p></details><details id="S059"><summary>S059 · github/MoonshotAI--Kimi-K3/commit.json · exact</summary><p><a href="https://api.github.com/repos/MoonshotAI/Kimi-K3/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/MoonshotAI/Kimi-K3/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:44.547132+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>56b5c0b187530f0905113d50f3028905cc2190be7ded0f41473d60dd1db6f255</code><br>Computed UTF-8 SHA-256: <code>56b5c0b187530f0905113d50f3028905cc2190be7ded0f41473d60dd1db6f255</code><br>Recorded bytes 3805; embedded UTF-8 bytes 3805</p></details><details id="S060"><summary>S060 · github/zai-org--GLM-5/commit.json · exact</summary><p><a href="https://api.github.com/repos/zai-org/GLM-5/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/zai-org/GLM-5/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:44.538095+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>956538bb7898584f25d07f797cfb81cfd0f0929c231545a116f7f96bc3051c59</code><br>Computed UTF-8 SHA-256: <code>956538bb7898584f25d07f797cfb81cfd0f0929c231545a116f7f96bc3051c59</code><br>Recorded bytes 5865; embedded UTF-8 bytes 5865</p></details><details id="S061"><summary>S061 · github/QwenLM--Qwen3.8/commit.json · exact</summary><p><a href="https://api.github.com/repos/QwenLM/Qwen3.8/commits/main" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/QwenLM/Qwen3.8/commits/main</a></p><p class="small">Retrieved 2026-10-08T08:51:44.698365+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>a76b5834fc66860f21c9a9bb8a170bf0853a16de75d37fe793f1ee32309652be</code><br>Computed UTF-8 SHA-256: <code>a76b5834fc66860f21c9a9bb8a170bf0853a16de75d37fe793f1ee32309652be</code><br>Recorded bytes 25960; embedded UTF-8 bytes 25960</p></details><details id="S062"><summary>S062 · github/zai-org--GLM-5/tree.json · exact</summary><p><a href="https://api.github.com/repos/zai-org/GLM-5/git/trees/c8ad661c6cf4cb0a78064987bc42f97e14355929?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/zai-org/GLM-5/git/trees/c8ad661c6cf4cb0a78064987bc42f97e14355929?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:45.233589+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>7d05add59d05ffc06f0b973384eaf1d616b90f60f515750efdd521476a300fc1</code><br>Computed UTF-8 SHA-256: <code>7d05add59d05ffc06f0b973384eaf1d616b90f60f515750efdd521476a300fc1</code><br>Recorded bytes 6428; embedded UTF-8 bytes 6428</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;c8ad661c6cf4cb0a78064987bc42f97e14355929&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 28,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;.github&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;3a5a2d8b645b979a7f86b579b5fb9950a15a5a59&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/3a5a2d8b645b979a7f86b579b5fb9950a15a5a59&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;19df11d3e36c8c554e1ed8f0559c1db6b93b6f18&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/19df11d3e36c8c554e1ed8f0559c1db6b93b6f18&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE/bug_report.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;7df03d1bd143b491f75ee968bd8c9bea4af6cf6b&quot;,
      &quot;size&quot;: 2638,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/7df03d1bd143b491f75ee968bd8c9bea4af6cf6b&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE/feature-request.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;789db9c1a5ec16451d18edc98fc22f836d357624&quot;,
      &quot;size&quot;: 1189,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/789db9c1a5ec16451d18edc98fc22f836d357624&quot;
    },
    {
      &quot;path&quot;: &quot;.github/PULL_REQUEST_TEMPLATE.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;722039181469e84ce33a5287efc9567ca2a75da6&quot;,
      &quot;size&quot;: 1235,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/722039181469e84ce33a5287efc9567ca2a75da6&quot;
    },
    {
      &quot;path&quot;: &quot;.gitignore&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;7fe3d435e261532e04d35183920a2dcdb68147cc&quot;,
      &quot;size&quot;: 1697,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/7fe3d435e261532e04d35183920a2dcdb68147cc&quot;
    },
    {
      &quot;path&quot;: &quot;.pre-commit-config.yaml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0c44522cd53d6c2d5cba80039e38c857da80aa91&quot;,
      &quot;size&quot;: 429,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/0c44522cd53d6c2d5cba80039e38c857da80aa91&quot;
    },
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;ca04a51e70398d8fe19b23a8c317374dfe25408b&quot;,
      &quot;size&quot;: 11343,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/ca04a51e70398d8fe19b23a8c317374dfe25408b&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;b50da8f1ed05d96f27184200496acd59f3b345fb&quot;,
      &quot;size&quot;: 15185,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/b50da8f1ed05d96f27184200496acd59f3b345fb&quot;
    },
    {
      &quot;path&quot;: &quot;README_zh.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;46d05024df3bd6dca172cb3b5776d5aa15c682a6&quot;,
      &quot;size&quot;: 14621,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/46d05024df3bd6dca172cb3b5776d5aa15c682a6&quot;
    },
    {
      &quot;path&quot;: &quot;example&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;baaba216430c920e160d6a3f5321ef44936ecffa&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/baaba216430c920e160d6a3f5321ef44936ecffa&quot;
    },
    {
      &quot;path&quot;: &quot;example/ascend.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;caed7f0d0f7d336db2998a35b57b1c613a6e328c&quot;,
      &quot;size&quot;: 3147,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/caed7f0d0f7d336db2998a35b57b1c613a6e328c&quot;
    },
    {
      &quot;path&quot;: &quot;requirements.txt&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;593729185e7076df0f7ac6bd6ee162ed8a17abf9&quot;,
      &quot;size&quot;: 58,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/593729185e7076df0f7ac6bd6ee162ed8a17abf9&quot;
    },
    {
      &quot;path&quot;: &quot;resources&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;bf5a290eb422ce48723f96b94cb2342b27ab0a8c&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/bf5a290eb422ce48723f96b94cb2342b27ab0a8c&quot;
    },
    {
      &quot;path&quot;: &quot;resources/WECHAT.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;4cc2da66b0b6a9ea10dd8b5c35ab4e7a564ecc28&quot;,
      &quot;size&quot;: 179,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/4cc2da66b0b6a9ea10dd8b5c35ab4e7a564ecc28&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0cb2f3d0a3683ecbf422ce3ebb64307d8e691519&quot;,
      &quot;size&quot;: 1891638,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/0cb2f3d0a3683ecbf422ce3ebb64307d8e691519&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench_51.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;cd447649fe2118a8688f4e8463593edb171af15e&quot;,
      &quot;size&quot;: 1430884,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/cd447649fe2118a8688f4e8463593edb171af15e&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench_52.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;be4838d31334ac90dbed58853bcc307ef7cca4b2&quot;,
      &quot;size&quot;: 725694,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/be4838d31334ac90dbed58853bcc307ef7cca4b2&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench_52_lh.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8233b0b10a226bede40b5478c2ad5d9eae768528&quot;,
      &quot;size&quot;: 3550923,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/8233b0b10a226bede40b5478c2ad5d9eae768528&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench_53.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;44a96a603eff33eb235e37a559ba453fb4887601&quot;,
      &quot;size&quot;: 1127922,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/44a96a603eff33eb235e37a559ba453fb4887601&quot;
    },
    {
      &quot;path&quot;: &quot;resources/bench_53_2.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e82f6a168e7b6be019f8e5d46dd8c42b6ad39e5e&quot;,
      &quot;size&quot;: 1102398,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/e82f6a168e7b6be019f8e5d46dd8c42b6ad39e5e&quot;
    },
    {
      &quot;path&quot;: &quot;resources/logo.svg&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;3c9d247911fc31d2e29ef4f40c017c16d78d7a38&quot;,
      &quot;size&quot;: 11036,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/3c9d247911fc31d2e29ef4f40c017c16d78d7a38&quot;
    },
    {
      &quot;path&quot;: &quot;resources/realworld_bench.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e9a7955a690f02140b9b492be9223ced8f1661f5&quot;,
      &quot;size&quot;: 758457,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/e9a7955a690f02140b9b492be9223ced8f1661f5&quot;
    },
    {
      &quot;path&quot;: &quot;resources/vending_bench.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;857bdb47f46ed9e2aa7c453a66e04e36cc2ff6c3&quot;,
      &quot;size&quot;: 225180,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/857bdb47f46ed9e2aa7c453a66e04e36cc2ff6c3&quot;
    },
    {
      &quot;path&quot;: &quot;resources/wechat.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e23ee84e43415c6ec428fbd7f41dc9e0af0cb1a3&quot;,
      &quot;size&quot;: 4724,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/e23ee84e43415c6ec428fbd7f41dc9e0af0cb1a3&quot;
    },
    {
      &quot;path&quot;: &quot;skills&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;e24014aa1e99b8d681d37807de055e802fa6795c&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/e24014aa1e99b8d681d37807de055e802fa6795c&quot;
    },
    {
      &quot;path&quot;: &quot;skills/glm-master-skill&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;7bcedffc7da8b1afbc4296f104ab0f08047b5e40&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/trees/7bcedffc7da8b1afbc4296f104ab0f08047b5e40&quot;
    },
    {
      &quot;path&quot;: &quot;skills/glm-master-skill/SKILL.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;8f8f612098d859d86b6fde788eef289bc9fbbf1a&quot;,
      &quot;size&quot;: 6498,
      &quot;url&quot;: &quot;https://api.github.com/repos/zai-org/GLM-5/git/blobs/8f8f612098d859d86b6fde788eef289bc9fbbf1a&quot;
    }
  ]
}</pre></details><details id="S063"><summary>S063 · github/MoonshotAI--Kimi-K3/tree.json · exact</summary><p><a href="https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:45.233452+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>dcc0a0fd81e0a987c7dc486d30befadeecdf6a31000b7b74aaf9bb4d41eacf7f</code><br>Computed UTF-8 SHA-256: <code>dcc0a0fd81e0a987c7dc486d30befadeecdf6a31000b7b74aaf9bb4d41eacf7f</code><br>Recorded bytes 1287; embedded UTF-8 bytes 1287</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;3cb39dfd32e51c3328e2e4b4af21341247d06c43&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 5,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;0de9f0698195e20305982b95f47246bfd85602e5&quot;,
      &quot;size&quot;: 3065,
      &quot;url&quot;: &quot;https://api.github.com/repos/MoonshotAI/Kimi-K3/git/blobs/0de9f0698195e20305982b95f47246bfd85602e5&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;21fe1f27c13dd41d8e45da6ea628e98abc105b60&quot;,
      &quot;size&quot;: 45004,
      &quot;url&quot;: &quot;https://api.github.com/repos/MoonshotAI/Kimi-K3/git/blobs/21fe1f27c13dd41d8e45da6ea628e98abc105b60&quot;
    },
    {
      &quot;path&quot;: &quot;assets&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;44cb3966a9d011df3c3587e6e005fd245aa0b8f2&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/44cb3966a9d011df3c3587e6e005fd245aa0b8f2&quot;
    },
    {
      &quot;path&quot;: &quot;assets/kimi-logo.png&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;870b8be6e07cc2c46f7173e800fbaff8af0af5d1&quot;,
      &quot;size&quot;: 87988,
      &quot;url&quot;: &quot;https://api.github.com/repos/MoonshotAI/Kimi-K3/git/blobs/870b8be6e07cc2c46f7173e800fbaff8af0af5d1&quot;
    },
    {
      &quot;path&quot;: &quot;k3_tech_report.pdf&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a869c3eba43ba7055f367761330252abcbac66db&quot;,
      &quot;size&quot;: 1795670,
      &quot;url&quot;: &quot;https://api.github.com/repos/MoonshotAI/Kimi-K3/git/blobs/a869c3eba43ba7055f367761330252abcbac66db&quot;
    }
  ]
}</pre></details><details id="S064"><summary>S064 · github/QwenLM--Qwen3.8/tree.json · exact</summary><p><a href="https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1" target="_blank" rel="noopener noreferrer">https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1</a></p><p class="small">Retrieved 2026-10-08T08:51:45.387443+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>05fddb816468c1c3b8e72c70e8ae006ce88f993327d3459305976f834fbe91ad</code><br>Computed UTF-8 SHA-256: <code>05fddb816468c1c3b8e72c70e8ae006ce88f993327d3459305976f834fbe91ad</code><br>Recorded bytes 1955; embedded UTF-8 bytes 1955</p><p><b>JSON $.tree (all paths for S062–S064; relevant paths for larger trees)</b></p><pre>{
  &quot;sha&quot;: &quot;2ea10dc725823bf7c3e21ce8557cbe15245132ae&quot;,
  &quot;truncated&quot;: false,
  &quot;full_entry_count&quot;: 8,
  &quot;entries&quot;: [
    {
      &quot;path&quot;: &quot;.github&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;17120e45eead37f07bd65b06b93edeef6c6acb35&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/17120e45eead37f07bd65b06b93edeef6c6acb35&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;02065ddd1b34a6658106012d767cd7f0eaf75c7f&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/02065ddd1b34a6658106012d767cd7f0eaf75c7f&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE/bug_report.yml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e724d81303873205288fee90a06d183d1257fa8e&quot;,
      &quot;size&quot;: 3730,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/blobs/e724d81303873205288fee90a06d183d1257fa8e&quot;
    },
    {
      &quot;path&quot;: &quot;.github/ISSUE_TEMPLATE/config.yml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;e795534bfdb7ceb354ce050fd9e6a484230356d8&quot;,
      &quot;size&quot;: 516,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/blobs/e795534bfdb7ceb354ce050fd9e6a484230356d8&quot;
    },
    {
      &quot;path&quot;: &quot;.github/workflows&quot;,
      &quot;mode&quot;: &quot;040000&quot;,
      &quot;type&quot;: &quot;tree&quot;,
      &quot;sha&quot;: &quot;cd515e4c1ac7980ddf6f4afa597370192776b997&quot;,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/cd515e4c1ac7980ddf6f4afa597370192776b997&quot;
    },
    {
      &quot;path&quot;: &quot;.github/workflows/inactive.yml&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;68f851e90ae0a94a49d3ceb2d53e3fe89737996c&quot;,
      &quot;size&quot;: 1460,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/blobs/68f851e90ae0a94a49d3ceb2d53e3fe89737996c&quot;
    },
    {
      &quot;path&quot;: &quot;LICENSE&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;261eeb9e9f8b2b4b0d119366dda99c6fd7d35c64&quot;,
      &quot;size&quot;: 11357,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/blobs/261eeb9e9f8b2b4b0d119366dda99c6fd7d35c64&quot;
    },
    {
      &quot;path&quot;: &quot;README.md&quot;,
      &quot;mode&quot;: &quot;100644&quot;,
      &quot;type&quot;: &quot;blob&quot;,
      &quot;sha&quot;: &quot;a7896ca771e13546c9eb62b67d6e65c3dba680bf&quot;,
      &quot;size&quot;: 14334,
      &quot;url&quot;: &quot;https://api.github.com/repos/QwenLM/Qwen3.8/git/blobs/a7896ca771e13546c9eb62b67d6e65c3dba680bf&quot;
    }
  ]
}</pre></details><details id="S065"><summary>S065 · github/zai-org--GLM-5/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:45.821216+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>78ebb1936e94a4c4412bca2abc1c9cbe599455df12fb1373e75c495cb043651e</code><br>Computed UTF-8 SHA-256: <code>78ebb1936e94a4c4412bca2abc1c9cbe599455df12fb1373e75c495cb043651e</code><br>Recorded bytes 15185; embedded UTF-8 bytes 15185</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 17-24</b></p><pre>
### GLM-5.3 &amp; GLM-5.3-Flash

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

+ Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and
  Agents&#x27; Last Exam.
+ Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.</pre><p><b>content lines 55-63</b></p><pre>GLM-5.1, by contrast, is built to stay effective on agentic tasks over much longer horizons. We&#x27;ve found that the model handles ambiguous problems with better judgment and stays productive over longer sessions. It breaks complex problems down, runs experiments, reads results, and identifies blockers with real precision. By revisiting its reasoning and revising its strategy through repeated iteration, GLM-5.1 sustains optimization over hundreds of rounds and thousands of tool calls. The longer it runs, the better the result.

### GLM-5

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

Reinforcement learning aims to bridge the gap between competence and excellence in pre-trained models. However, deploying it at scale for LLMs is a challenge due to the RL training inefficiency. To this end, we developed [slime](https://github.com/THUDM/slime), a novel **asynchronous RL infrastructure** that substantially improves training throughput and efficiency, enabling more fine-grained post-training iterations. With advances in both pre-training and post-training, GLM-5 delivers significant improvement compared to GLM-4.7 across a wide range of academic benchmarks and achieves best-in-class performance among all open-source models in the world on reasoning, coding, and agentic tasks,  closing the gap with frontier models.

![bench](resources/bench.png)</pre></details><details id="S066"><summary>S066 · github/MoonshotAI--Kimi-K3/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/MoonshotAI/Kimi-K3/3cb39dfd32e51c3328e2e4b4af21341247d06c43/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/MoonshotAI/Kimi-K3/3cb39dfd32e51c3328e2e4b4af21341247d06c43/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:45.846867+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>849a303d849486aac61d1f0c253e3a1148ca59e801be807151c2dc3cce57ccb2</code><br>Computed UTF-8 SHA-256: <code>849a303d849486aac61d1f0c253e3a1148ca59e801be807151c2dc3cce57ccb2</code><br>Recorded bytes 45004; embedded UTF-8 bytes 45004</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 611-622</b></p><pre>- [SGLang](https://github.com/sgl-project/sglang) — see [cookbook](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3)
- [TokenSpeed](https://lightseek.org/tokenspeed) — see [recipes](https://lightseek.org/tokenspeed/recipes/models#kimi-k3)

---
## 6. Model Usage

Kimi K3 always has thinking enabled, and will return `reasoning_content`. Thinking effort is configured with the top-level `reasoning_effort` request field, which supports `&quot;low&quot;`, `&quot;high&quot;`, and `&quot;max&quot;` (default `&quot;max&quot;`).

Kimi K3 was trained in the preserved thinking history mode. For multi-turn conversations and tool calls, Kimi K3 requires the complete assistant message returned by the API to be passed back to `messages` as-is — including `reasoning_content` and `tool_calls`, not just `content`:

```python
import openai</pre></details><details id="S067"><summary>S067 · github/QwenLM--Qwen3.8/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md</a></p><p class="small">Retrieved 2026-10-08T08:51:46.113012+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>a71ec46607f81d6056336fb0a8431a26a1c7d8db6ac568a0c021c36f1ed3c92e</code><br>Computed UTF-8 SHA-256: <code>a71ec46607f81d6056336fb0a8431a26a1c7d8db6ac568a0c021c36f1ed3c92e</code><br>Recorded bytes 14334; embedded UTF-8 bytes 14334</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 15-25</b></p><pre>## Introduction

### Qwen3.8

For the first time, Qwen3.8 brings a Qwen-Max-class model to open release. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to carry complex, multi-step tasks through to completion with greater reliability.

Qwen3.8 features the following enhancements:
- **Core Capabilities**: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
- **Agent Execution**: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
- **Downstream Compatibility**: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
- **Flexible Thinking Control**: Reasoning depth can be tuned with `reasoning_effort`, and reasoning context from historical messages is retained via `preserve_thinking`.</pre><p><b>content lines 34-44</b></p><pre>### Qwen3.5

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

Qwen3.5 features the following enhancements:

- **Unified Vision-Language Foundation**: Early fusion training on trillions of multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
- **Efficient Hybrid Architecture**: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.
- **Scalable RL Generalization**: Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
- **Global Linguistic Coverage**: Expanded support to 201 languages and dialects, enabling inclusive, worldwide deployment with nuanced cultural and regional understanding.
- **Next-Generation Training Infrastructure**: Near-100% multimodal training efficiency compared to text-only training and asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.</pre><p><b>content lines 118-145</b></p><pre>### Local Use

#### Hugging Face Transformers

[`transformers`](https://huggingface.co/docs/transformers) acts as the model-definition framework in the current open-weight LLM landscape.
It also includes functionalities for LLM inference and training. The addition of serving capabilities in `transformers` makes it much easier to integrate new models in your development.

To launch a server, simply use the `transformers serve` command:
```shell
transformers serve Qwen/Qwen3.8-27B --port 8000 --continuous-batching
```
An OpenAI-compatible API will be available at `http://localhost:8000/v1`.
See [the Serve CLI guide](https://huggingface.co/docs/transformers/serve-cli/serving) for more information.

#### llama.cpp

[`llama.cpp`](https://github.com/ggml-org/llama.cpp) enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware.
llama.cpp supports the Qwen3.5 open model series (text &amp; vision).
Look for models ending with GGUF on Hugging Face Hub.

#### MLX (Apple Silicon)

If you are running on Apple Silicon, both [`mlx-lm`](https://github.com/ml-explore/mlx-lm) (text-only) and [`mlx-vlm`](https://github.com/Blaizzy/mlx-vlm) (vision + text) support the Qwen3.5 open model series. Look for models ending with MLX on Hugging Face Hub.

#### Unsloth

[Unsloth](https://unsloth.ai) contains a local UI to run and train LLMs and diffusion models, including Qwen3.8 and more.
See [the Qwen3.8 guide](https://unsloth.ai/docs/models/qwen3.8) for running Qwen3.8 quants with Unsloth.</pre><p><b>content lines 191-197</b></p><pre>### Finetuning

We advise you to use training frameworks, including [Unsloth](https://github.com/unslothai/unsloth), [Swift](https://github.com/modelscope/swift), [Llama-Factory](https://github.com/hiyouga/LLaMA-Factory), to finetune your models with SFT, DPO, GRPO, etc.

## License Agreement

Please find the license file released with the model weights on Hugging Face Hub or ModelScope.</pre></details><details id="S068"><summary>S068 · github/allenai--OLMo-core/src/scripts/train/OLMo3/OLMo3-32B.py · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B.py" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B.py</a></p><p class="small">Retrieved 2026-10-08T09:04:25.813586+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>0cb62bdd7e580c7123579d7ba2dc8807c86ed80d4884b03576cf20a67f3d40a0</code><br>Computed UTF-8 SHA-256: <code>0cb62bdd7e580c7123579d7ba2dc8807c86ed80d4884b03576cf20a67f3d40a0</code><br>Recorded bytes 6255; embedded UTF-8 bytes 6255</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 40-81</b></p><pre>SEQUENCE_LENGTH = 8 * 1024
GLOBAL_BATCH_SIZE = 8 * 1024 * 1024


def build_model_config(common: CommonComponents) -&gt; TransformerConfig:
    return TransformerConfig.olmo3_32B(vocab_size=common.tokenizer.padded_vocab_size())


def build_train_module_config(common: CommonComponents) -&gt; TransformerTrainModuleConfig:
    rank_microbatch_size = SEQUENCE_LENGTH
    if common.launch is not None:
        gpus = {CLUSTER_TO_GPU_TYPE.get(c, &quot;unknown&quot;) for c in common.launch.clusters}
        if all(&quot;B200&quot; in g for g in gpus):
            rank_microbatch_size *= 2

    return TransformerTrainModuleConfig(
        rank_microbatch_size=rank_microbatch_size,
        max_sequence_length=common.max_sequence_length,
        optim=SkipStepAdamWConfig(
            lr=6e-4,
            weight_decay=0.1,
            betas=(0.9, 0.95),
            group_overrides=[
                OptimGroupOverride(params=[&quot;embeddings.weight&quot;], opts=dict(weight_decay=0.0))
            ],
        ),
        compile_model=True,
        dp_config=TransformerDataParallelConfig(
            name=DataParallelType.hsdp,
            param_dtype=DType.bfloat16,
            reduce_dtype=DType.float32,
            wrapping_strategy=TransformerDataParallelWrappingStrategy.full,
            shard_degree=64,
        ),
        ac_config=TransformerActivationCheckpointingConfig(
            mode=TransformerActivationCheckpointingMode.budget, activation_memory_budget=0.5
        ),
        float8_config=Float8Config(enabled=False),
        z_loss_multiplier=1e-5,
        max_grad_norm=1.0,
        scheduler=CosWithWarmup(warmup_steps=2000),
    )</pre><p><b>content lines 88-108</b></p><pre>) -&gt; DataComponents:
    dataset_config = NumpyFSLDatasetConfig.from_data_mix(
        DataMix.OLMo_mix_0925,
        tokenizer=common.tokenizer,
        mix_base_dir=common.root_dir,
        work_dir=common.work_dir,
        sequence_length=common.max_sequence_length,
        max_target_sequence_length=max(common.max_sequence_length, 8192),
        generate_doc_lengths=intra_document_masking,
        instance_filter_config=None
        if not include_instance_filter
        else InstanceFilterConfig(
            repetition_max_period=13, repetition_min_period=1, repetition_max_count=32
        ),
    )

    data_loader_config = NumpyDataLoaderConfig(
        global_batch_size=common.global_batch_size, seed=34521, num_workers=8
    )

    return DataComponents(dataset=dataset_config, data_loader=data_loader_config)</pre><p><b>content lines 119-128</b></p><pre>
    return (
        TrainerConfig(
            save_folder=f&quot;gs://ai2-llm/checkpoints/{common.run_name}/&quot;,
            save_overwrite=True,
            metrics_collect_interval=50,
            cancel_check_interval=cancel_check_interval,
            max_duration=Duration.epochs(1),
        )
        .with_callback(</pre></details><details id="S069"><summary>S069 · github/allenai--OLMo-core/src/scripts/train/OLMo3/OLMo3-32B-midtraining.py · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-midtraining.py" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-midtraining.py</a></p><p class="small">Retrieved 2026-10-08T09:04:25.814136+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>0a67600eafa7eedbb380609da8d7d6d3cb0c9d48cd920fbaefc42453db23b2a2</code><br>Computed UTF-8 SHA-256: <code>0a67600eafa7eedbb380609da8d7d6d3cb0c9d48cd920fbaefc42453db23b2a2</code><br>Recorded bytes 4976; embedded UTF-8 bytes 4976</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 20-24</b></p><pre>SEQ_LENGTH = 8192
GLOBAL_BATCH_SIZE = 4 * 1024 * 1024  # ~4M tokens
MAX_TOKENS = 100_000_000_000  # 100B
LR = 0.0002071235285
SEED = 1337</pre><p><b>content lines 55-98</b></p><pre>    tokenizer_config = TokenizerConfig.dolma2()
    model_config = TransformerConfig.olmo3_32B(
        vocab_size=tokenizer_config.padded_vocab_size(),
    )

    train_module_config: TransformerTrainModuleConfig = cookbook.configure_train_module(
        max_sequence_length=SEQ_LENGTH,
        rank_microbatch_size=SEQ_LENGTH,
        learning_rate=LR,
        scheduler=LinearWithWarmup(units=SchedulerUnits.steps, warmup=0, alpha_f=0.0),
        activation_memory_budget=0.5,
        dp_shard_degree=64,
    )

    source_list = SourceMixtureList.from_yaml(
        &quot;src/olmo_core/data/source_mixtures/OLMo3-32B-midtraining-modelnamefilter.yaml&quot;
    )
    source_list.validate()
    dataset_config = NumpyFSLDatasetConfig.from_src_mix(
        src_mix=SourceMixtureDatasetConfig(
            source_list=source_list,
            requested_tokens=MAX_TOKENS,
            global_batch_size=GLOBAL_BATCH_SIZE,
            processes=16,
            seed=SEED,
        ),
        tokenizer=tokenizer_config,
        work_dir=work_dir,
        sequence_length=SEQ_LENGTH,
        instance_filter_config=InstanceFilterConfig(
            repetition_max_period=13, repetition_min_period=1, repetition_max_count=32
        ),
    )

    data_loader_config = NumpyDataLoaderConfig(
        global_batch_size=GLOBAL_BATCH_SIZE, seed=SEED, num_workers=4
    )

    trainer_config = cookbook.configure_trainer(
        load_path=&quot;gs://ai2-llm/checkpoints/stego32-highlr-filter3/step656000&quot;,
        load_trainer_state=False,
        load_optim_state=True,
        max_duration=Duration.tokens(MAX_TOKENS),
        checkpoint_dir=save_dir,</pre></details><details id="S070"><summary>S070 · github/allenai--OLMo-core/src/scripts/train/OLMo3/OLMo3-32B-long-context.py · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-long-context.py" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-long-context.py</a></p><p class="small">Retrieved 2026-10-08T09:04:25.814191+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>8b9551a8001cfb0fe37c1d7056bf78bff3c200b91d03c593b860362cb58e6348</code><br>Computed UTF-8 SHA-256: <code>8b9551a8001cfb0fe37c1d7056bf78bff3c200b91d03c593b860362cb58e6348</code><br>Recorded bytes 6768; embedded UTF-8 bytes 6768</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 42-45</b></p><pre>SEQUENCE_LENGTH = 65536
GLOBAL_BATCH_SIZE = 8 * 1024 * 1024  # ~8M tokens, 2**23
MAX_TOKENS = 100_000_000_000  # 100B
LR = 0.0002071235285  # same as midtraining</pre><p><b>content lines 98-110</b></p><pre>    run_name = f&quot;{common.run_name}-{datetime.now().astimezone().strftime(&#x27;%Y%m%dT%H%M%z&#x27;)}&quot;

    return (
        TrainerConfig(
            load_path=&quot;gs://ai2-llm/checkpoints/stego32-midtraining-runs-merged-step23842-resharded16&quot;,
            load_strategy=LoadStrategy.always,
            load_trainer_state=False,
            load_optim_state=True,
            save_folder=f&quot;gs://ai2-llm/checkpoints/{common.run_name}/&quot;,
            save_overwrite=True,
            metrics_collect_interval=50,
            cancel_check_interval=cancel_check_interval,
            max_duration=Duration.tokens(MAX_TOKENS),</pre><p><b>content lines 145-162</b></p><pre>    Default dataset and data loader configurations. Constructs a simple FSL dataset and data loader
    configuration with default settings.
    &quot;&quot;&quot;
    dataset_config = NumpyPackedFSLDatasetConfig.glob(
        &quot;gs://ai2-llm/preprocessed/tylerr/lc-reshard-final-cleaned/v0.1/allenai/dolma2-tokenizer/*.npy&quot;,
        tokenizer=common.tokenizer,
        work_dir=common.work_dir,
        sequence_length=common.max_sequence_length,
        generate_doc_lengths=intra_document_masking,  # enables intra-document masking
        source_group_size=8,
        source_permutation_seed=123,
        instance_filter_config=None
        if not include_instance_filter
        else InstanceFilterConfig(
            repetition_max_period=13, repetition_min_period=1, repetition_max_count=32
        ),
    )
</pre></details><details id="S071"><summary>S071 · github/allenai--OLMo-core/src/scripts/official/OLMo3/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/official/OLMo3/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/official/OLMo3/README.md</a></p><p class="small">Retrieved 2026-10-08T09:04:25.814285+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>24d48f74755914d53b776d6273778d4a67bf38c90563e15c7cd48b6d321ab165</code><br>Computed UTF-8 SHA-256: <code>24d48f74755914d53b776d6273778d4a67bf38c90563e15c7cd48b6d321ab165</code><br>Recorded bytes 8646; embedded UTF-8 bytes 8646</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 1-10</b></p><pre># Olmo 3 Model Training

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The base models were trained using a staged training approach.

Olmo is a series of **O**pen **l**anguage **mo**dels designed to enable the science of language models. These models are trained on the Dolma 3 dataset. We are releasing all code, checkpoints, logs (coming soon), and associated training details.

| Size   | Training Tokens | Layers | Hidden Size | Q Heads | KV Heads | Context Length |
|--------|-----------------|--------|-------------|---------|----------|----------------|
| [OLMo 3 7B](https://huggingface.co/allenai/Olmo-3-1025-7B) | 5.93 Trillion | 32 | 4096 | 32 | 32 | 65,536 |
| [OLMo 3 32B](https://huggingface.co/allenai/Olmo-3-1125-32B) | 5.50 Trillion | 64 | 5120 | 40 | 8 | 65,536 |</pre><p><b>content lines 21-47</b></p><pre>## Training Data

Olmo 3 7B pretraining follows a three-stage procedure.
In the first stage, we train on large amounts of mostly web-based data: [dolma3](https://huggingface.co/datasets/allenai/dolma3).
In the second stage, we train on a smaller amount of high-quality, targeted data: [dolma3-dolmino](https://huggingface.co/datasets/allenai/dolma3_dolmino).
And in the third stage, we train on high-quality data consisting of a portion of longer documents: [dolma3-longmino](https://huggingface.co/datasets/allenai/dolma3_longmino).

For further details please refer to the [dolma3](https://github.com/allenai/dolma3) repo.

Versions of these datasets that have been pre-tokenized with [allenai/dolma3-tokenizer](https://huggingface.co/allenai/dolma2-tokenizer) (same as `allenai/dolma2-tokenizer`) are available from https://olmo-data.org/, with manifests defined in the mixes below:

| Model | Stage | Data Mix |
|-------|-------|-----|
| Olmo 3 7B | stage 1 (pretraining) | [OLMo-mix-0625-official.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-mix-0625-official.txt) |
| Olmo 3 7B | stage 2 (midtraining) | [OLMo-midtraining-mix-0625-100B.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-midtraining-mix-0625-100B.txt) |
| Olmo 3 7B | stage 3 (long-context) | [OLMo-longmino-mix-0625.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-longmino-mix-0625.txt) |
| Olmo 3 32B | stage 1 (pretraining) | dolma3 -&gt; [OLMo-mix-0925-official.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-mix-0925-official.txt) |
| Olmo 3 32B | stage 2 (midtraining) | dolma3-dolmino -&gt; [OLMo-midtraining-mix-0925-ingredient1-100B.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-midtraining-mix-0925-ingredient1-100B.txt) &lt;br&gt; [OLMo-midtraining-mix-0925-ingredient2-100B.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-midtraining-mix-0925-ingredient2-100B.txt) |
| Olmo 3 32B | stage 3 (long-context) | dolma3-longmino -&gt; [OLMo-longmino-mix-0925.txt](https://github.com/allenai/Olmo-core/blob/main/src/olmo_core/data/mixes/OLMo-longmino-mix-0925.txt) |

In general, we recommend the mixes defined for Olmo 3 32B as they are slightly more refined.

For example, a numpy file containing tokenized data could be retrieved with:

```bash
wget https://olmo-data.org/preprocessed/dolma3-0625/v0.1-official/allenai/dolma3-tokenizer/olmocr_science_pdfs/science_math_and_technology/000000.npy
```</pre><p><b>content lines 69-87</b></p><pre>## Olmo 3 32B Model Training

Official training scripts, checkpoints, and monitoring logs for the Olmo 3 32B pretraining process can be found in the table below. Unlike for Olmo 3 7B, we use model merging (&quot;souping&quot;) at multiple points during pretraining. In particular, we soup (with simple averaging of parameters) the outputs of two separate midtraining runs and we soup the final three checkpoints produced by the long-context stage.

| Stage | Tokens  | GPUs | Script | Monitoring |
|-------|-----------|------|--------|------------|
| stage 1 (pretraining) | 5.50 Trillion | 1024 H100s | [OLMo-3-1025-32B-pretrain.py](https://github.com/allenai/Olmo-core/blob/main/src/scripts/official/OLMo3/OLMo-3-1025-32B-pretrain.py) | [wandb.ai/Olmo3-32B](https://wandb.ai/ai2-llm/Olmo-3-1125-32B/reports/Olmo-3-32B-November-2025--VmlldzoxNTA4NzAxMw) |
| stage 2 (midtraining) | 100 Billion x2 | 512 H100s | [OLMo-3-1025-32B-midtrain-ingredient-1.py](https://github.com/allenai/Olmo-core/blob/main/src/scripts/official/OLMo3/OLMo-3-1025-32B-midtrain-ingredient-1.py) &lt;br&gt; [OLMo-3-1025-32B-midtrain-ingredient-2.py](https://github.com/allenai/Olmo-core/blob/main/src/scripts/official/OLMo3/OLMo-3-1025-32B-midtrain-ingredient-2.py) | [wandb.ai/Olmo3-32B](https://wandb.ai/ai2-llm/Olmo-3-1125-32B/reports/Olmo-3-32B-November-2025--VmlldzoxNTA4NzAxMw) |
| stage 3 (long-context) | 100 Billion | 1024 H100s | [OLMo-3-1025-32B-long-context.py](https://github.com/allenai/Olmo-core/blob/main/src/scripts/official/OLMo3/OLMo-3-1025-32B-long-context.py) | [wandb.ai/Olmo3-32B](https://wandb.ai/ai2-llm/Olmo-3-1125-32B/reports/Olmo-3-32B-November-2025--VmlldzoxNTA4NzAxMw) |

A full list of Olmo-core format checkpoints for Olmo 3 32B can be found in [OLMo-3-1025-32B.csv](https://github.com/allenai/Olmo-core/blob/main/src/scripts/official/OLMo3/OLMo-3-1025-32B.csv).

A full list of HF format checkpoints for Olmo 3 32B can be found by enumerating the HF repo refs:

```python
from huggingface_hub import list_repo_refs
out = list_repo_refs(&quot;allenai/Olmo-3-1125-32B&quot;)
branches = [b.name for b in out.branches]
```</pre></details><details id="S072"><summary>S072 · github/NVIDIA-NeMo--Nemotron/src/nemotron/recipes/super3/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/README.md</a></p><p class="small">Retrieved 2026-10-08T09:04:26.508518+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>e75f4531f9ff9e55f1495e1a5a8aca7ef695ea0d64b2c59e6153567b3689e438</code><br>Computed UTF-8 SHA-256: <code>e75f4531f9ff9e55f1495e1a5a8aca7ef695ea0d64b2c59e6153567b3689e438</code><br>Recorded bytes 8508; embedded UTF-8 bytes 8508</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 1-15</b></p><pre># Nemotron 3 Super Training Recipe

A complete 4-stage training pipeline for Nemotron 3 Super, a high-capacity hybrid Mamba-Transformer-MoE model with multi-token prediction and DeepEP.

## Model Overview

Nemotron 3 Super is a high-capacity hybrid Mamba-Transformer with sparse MoE, featuring multi-token prediction (MTP), shared experts, and DeepEP for efficient expert parallelism.

| Property | Value |
|----------|-------|
| Architecture | Hybrid Mamba-Transformer with sparse MoE |
| Multi-Token Prediction | Yes (MTP layers with 0.3 loss scaling) |
| Shared Experts | Yes |
| DeepEP | Supported |
| Training Stages | 4 (Pretrain → SFT → RL → Eval) |</pre><p><b>content lines 51-66</b></p><pre>| Stage | Purpose | Framework | Output |
|-------|---------|-----------|--------|
| [Stage 0: Pretrain](./stage0_pretrain/) | Train on large text corpus | Megatron-Bridge | Base model checkpoint |
| [Stage 1: SFT](./stage1_sft/) | Instruction tuning | Megatron-Bridge | Instruction-following model |
| [Stage 2: RL](./stage2_rl/) | Alignment with GRPO | NeMo-RL | Final aligned model |
| [Stage 3: Eval](./stage3_eval/) | Model evaluation | NeMo-Evaluator | Benchmark results |

## Prerequisites

### v0 Requirements

&gt; **Slurm Only**: This initial release has been tested exclusively with Slurm execution. Support for additional NeMo-Run executors (local, Docker, SkyPilot, DGX Cloud) is planned for future releases.

- **Slurm cluster**: GPU nodes (B200 recommended, 4 nodes for training)
- **Weights &amp; Biases**: Required for experiment tracking and artifact lineage (future versions will be backend-agnostic)
- **Container images**: NeMo containers with Megatron-Bridge and NeMo-RL</pre><p><b>content lines 91-110</b></p><pre>## Quick Start

### Full Pipeline

```bash
# Stage 0: Data prep + Pretraining
uv run nemotron super3 data prep pretrain --run YOUR-CLUSTER
uv run nemotron super3 pretrain --run YOUR-CLUSTER

# Stage 1: Data prep + SFT
uv run nemotron super3 data prep sft --run YOUR-CLUSTER
uv run nemotron super3 sft --run YOUR-CLUSTER

# Stage 2: Data prep + RL
uv run nemotron super3 data prep rl --run YOUR-CLUSTER
uv run nemotron super3 rl --run YOUR-CLUSTER

# Stage 3: Evaluation
uv run nemotron super3 eval --run YOUR-CLUSTER
```</pre></details><details id="S073"><summary>S073 · github/NVIDIA-NeMo--Nemotron/src/nemotron/recipes/super3/stage0_pretrain/README.md · exact</summary><p><a href="https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md</a></p><p class="small">Retrieved 2026-10-08T09:04:26.509142+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>bb9ec9703bc8eeb59eddbe406153bce9870892a621f49cb881c9f2a05523bf29</code><br>Computed UTF-8 SHA-256: <code>bb9ec9703bc8eeb59eddbe406153bce9870892a621f49cb881c9f2a05523bf29</code><br>Recorded bytes 8028; embedded UTF-8 bytes 8028</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 1-28</b></p><pre># Stage 0: Pretraining

Pretrain Nemotron 3 Super on a large text corpus using Megatron-Bridge.

## Overview

This stage tokenizes raw text data and trains the base language model from scratch. It produces a pretrained checkpoint that serves as the foundation for subsequent instruction tuning (SFT) and alignment (RL) stages.

| Component | Description |
|-----------|-------------|
| `data_prep.py` | Tokenizes raw text into Megatron bin/idx format |
| `train.py` | Runs pretraining using Megatron-Bridge |
| `config/` | Configuration files for data prep and training |

## Training Pipeline

Pretraining follows a **4-phase curriculum** from the tech report:

| Phase | Config | Internal Scale | Focus |
|-------|--------|----------------|-------|
| Phase 1 | `phase1` | 20T (80%) | Diversity — broad coverage, WSD warmup + stable LR |
| Phase 2 | `phase2` | 5T (20%) | Quality — high-quality sources, WSD minus_sqrt decay |
| LC Stage 1 | `long_context_1m` | 34B | 1M context extension, constant LR 4.5e-6 |
| LC Stage 2 | `long_context_mixed` | 17B | Alternating 1M/4K to recover math benchmarks |

&gt; **Note**: The open-sourced data covers an estimated 8–10T tokens (~40–50% of the
&gt; internal 25T blend). Missing categories (code, nemotron-cc-code, crawl++, academic)
&gt; used internal-only data. Adjust `train_iters` and blend weights to match your data.</pre><p><b>content lines 30-62</b></p><pre>## Quick Start

### Using nemotron CLI (Recommended)

```bash
# 1. Prepare data for each phase
uv run nemotron super3 data prep pretrain -c phase1 --run YOUR-CLUSTER
uv run nemotron super3 data prep pretrain -c phase2 --run YOUR-CLUSTER
uv run nemotron super3 data prep pretrain -c long_context --run YOUR-CLUSTER

# 2. Run pretraining phases sequentially
uv run nemotron super3 pretrain -c phase1 --run YOUR-CLUSTER
uv run nemotron super3 pretrain -c phase2 --run YOUR-CLUSTER     # resumes from phase1 checkpoint
uv run nemotron super3 pretrain -c long_context_1m --run YOUR-CLUSTER  # resumes from phase2
uv run nemotron super3 pretrain -c long_context_mixed --run YOUR-CLUSTER  # resumes from lc1

# Quick test with tiny config
uv run nemotron super3 pretrain -c tiny --run YOUR-CLUSTER
```

### Direct Script Execution

Inside a container on a compute node:

```bash
# Data preparation
python data_prep.py --config config/data_prep/phase1.yaml

# Training (single node)
python train.py --config config/tiny.yaml

# Training (distributed)
torchrun --nproc_per_node=8 train.py --config config/phase1.yaml</pre></details><details id="S074"><summary>S074 · github/NVIDIA-NeMo--Evaluator/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md · exact</summary><p><a href="https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/c87f9b21769cd1a7f3b85267264101a7bcc77df6/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md" target="_blank" rel="noopener noreferrer">https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/c87f9b21769cd1a7f3b85267264101a7bcc77df6/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md</a></p><p class="small">Retrieved 2026-10-08T09:04:26.508750+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>87b6e057980f949158cfc6ad5ae0a925b634975b3b8597d82653bfd6a9a28699</code><br>Computed UTF-8 SHA-256: <code>87b6e057980f949158cfc6ad5ae0a925b634975b3b8597d82653bfd6a9a28699</code><br>Recorded bytes 20663; embedded UTF-8 bytes 20663</p><p class="small">Git blob SHA-1 matches tree: True</p><p><b>content lines 1-3</b></p><pre># NVIDIA Nemotron 3 Super 120B A12B — Reproducing Model Card Evaluation Results

This tutorial demonstrates how to reproduce the evaluation results for the [**NVIDIA Nemotron 3 Super 120B A12B**](https://build.nvidia.com/nvidia/nemotron-3-super-120b-a12b) model using the NeMo Evaluator Launcher.</pre><p><b>content lines 34-69</b></p><pre>---

## Prerequisites

### 1. Install the NeMo Evaluator Launcher

```bash
pip install nemo-evaluator-launcher
```

Or install from source:

```bash
git clone https://github.com/NVIDIA-NeMo/Evaluator.git
cd Evaluator/packages/nemo-evaluator-launcher
pip install -e .
```

### 2. Required API Keys

Set the following environment variables before running the evaluation:

```bash
# Required: NGC API key for accessing the model endpoint
export NGC_API_KEY=&quot;your-ngc-api-key&quot;

# Required: HuggingFace token for accessing datasets
export HF_TOKEN=&quot;your-huggingface-token&quot;

# Required for HLE and Multi-Challenge benchmarks: OpenAI/Azure API key for GPT-4o judge
export JUDGE_API_KEY=&quot;your-judge-api-key&quot;

# Required for Omniscience benchmark: Gemini API key for Gemini-3-Flash-Preview judge
export GEMINI_API_KEY=&quot;your-gemini-api-key&quot;

# Required for AA LCR benchmark: API key with access to non-reasoning Qwen3-235B-A22B judge</pre><p><b>content lines 95-115</b></p><pre>|-------------|---------|
| [`local_nemotron-3-super-120b-a12b.yaml`](./local_nemotron-3-super-120b-a12b.yaml) | Main evaluation suite |
| [`local_nemotron-3-super-120b-a12b_tools.yaml`](./local_nemotron-3-super-120b-a12b_tools.yaml) | Tool usage evaluation |
| [`local_nemotron-3-super-120b-a12b_low_budget.yaml`](./local_nemotron-3-super-120b-a12b_low_budget.yaml) | Low-budget / low-effort thinking evaluation |

### 2. Run the Evaluation

```bash
nemo-evaluator-launcher run \
  --config local_nemotron-3-super-120b-a12b.yaml
```

### 3. Dry Run (Preview Configuration)

To preview the configuration without running the evaluation:

```bash
nemo-evaluator-launcher run \
  --config local_nemotron-3-super-120b-a12b.yaml \
  --dry-run
```</pre><p><b>content lines 140-185</b></p><pre>| `nemo_skills.ns_mmlu_pro` | MMLU-Pro (10-choice format) |
| `nemo_skills.ns_aime2025` | AIME 2025 |
| `nemo_skills.ns_aime2026` | AIME 2026 |
| `ns_hmmt_feb2025` | HMMT Feb 2025 |
| `AA_math_test_500` | Math 500 |
| `nemo_skills.ns_critpt` | CritPt |
| `ns_scicode` | SciCode |
| `ns_livecodebench_v5` | LiveCodeBench v5 (Jul–Dec 2024) |
| `ns_livecodebench` | LiveCodeBench v6 (Aug 2024–May 2025) |
| `ns_ifbench` | IFBench |
| `nemo_skills.ns_bfcl_v4` | BFCL v4 |
| `tau2_bench_telecom` | Tau2 Bench Telecom |
| `tau2_bench_retail` | Tau2 Bench Retail |
| `tau2_bench_airline` | Tau2 Bench Airline |
| `ns_arena_hard_v2` | Arena Hard v2 |
| `ns_aa_lcr` | AA LCR |
| `ns_omniscience` | AA-Omniscience |
| `ns_mmlu_prox` | MMLU-Pro X |
| `ns_wmt24pp_comet` | WMT24++ (COMET score) |

---

## Configuration Details

### Model Endpoint

The evaluation uses the NVIDIA API endpoint:

```yaml
target:
  api_endpoint:
    model_id: nvidia/nemotron-super-120b-a12b
    url: https://integrate.api.nvidia.com/v1/chat/completions
    api_key_name: NGC_API_KEY
```

### Default Parameters

The configuration uses the following default inference parameters:

| Parameter | Value | Description |
|-----------|-------|-------------|
| `max_new_tokens` | 131072 | Maximum tokens to generate |
| `temperature` | 1.0 | Sampling temperature |
| `top_p` | 0.95 | Nucleus sampling threshold |
| `parallelism` | 1 | Concurrent requests |</pre><p><b>content lines 346-363</b></p><pre>   - Verify the key names match exactly: `NGC_API_KEY`, `HF_TOKEN`, `JUDGE_API_KEY`, `GEMINI_API_KEY`, `QWEN_API_KEY`, `ARTIFICIAL_ANALYSIS_API_KEY`

2. **Timeout Errors**
   - The model generates long reasoning chains; timeouts are set to 3600s
   - Reduce `parallelism` if hitting rate limits

3. **HLE Judge Failures**
   - Verify `JUDGE_API_KEY` has access to GPT-4o
   - Check the judge endpoint URL is accessible

4. **WMT24++ COMET Scoring**
   - Download the XCOMET-XXL checkpoint from HuggingFace and update `model_path` in the config

5. **HuggingFace Dataset Access**
   - Some datasets require accepting terms on HuggingFace Hub
   - Ensure your `HF_TOKEN` has appropriate permissions

### Getting Help</pre></details><details id="S075"><summary>S075 · datasets/allenai--dolma3_mix-5.5T-1125-summary.json · exact</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified</a></p><p class="small">Retrieved 2026-10-08T09:04:26.507687+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>e3000b26dd2d82f573d03a6da830c1a7a5c843ff6e5339492a52eddd78b607c0</code><br>Computed UTF-8 SHA-256: <code>e3000b26dd2d82f573d03a6da830c1a7a5c843ff6e5339492a52eddd78b607c0</code><br>Recorded bytes 682; embedded UTF-8 bytes 682</p><p><b>Complete bounded dataset metadata JSON; reduced fields for S075–S077</b></p><pre>{
  &quot;_id&quot;: &quot;6923dbbd0304ddc9b6172334&quot;,
  &quot;id&quot;: &quot;allenai/dolma3_mix-6T&quot;,
  &quot;sha&quot;: &quot;689a3ea2d8217e64d73a5058913fa43ad15e81aa&quot;,
  &quot;gated&quot;: false,
  &quot;cardData&quot;: {
    &quot;license&quot;: &quot;odc-by&quot;,
    &quot;task_categories&quot;: [
      &quot;text-generation&quot;
    ],
    &quot;language&quot;: [
      &quot;en&quot;
    ],
    &quot;configs&quot;: [
      {
        &quot;config_name&quot;: &quot;default&quot;,
        &quot;data_files&quot;: [
          {
            &quot;split&quot;: &quot;train&quot;,
            &quot;path&quot;: &quot;data/**/*.jsonl.zst&quot;
          }
        ],
        &quot;features&quot;: [
          {
            &quot;name&quot;: &quot;id&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;text&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;metadata&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;source&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;version&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;created&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;added&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;doc&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;attributes&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          }
        ]
      }
    ]
  },
  &quot;lastModified&quot;: &quot;2026-01-15T05:36:27.000Z&quot;
}</pre></details><details id="S076"><summary>S076 · datasets/allenai--dolma3_dolmino_mix-100B-1125-summary.json · exact</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified</a></p><p class="small">Retrieved 2026-10-08T09:04:27.068347+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>9ac3a385fcfa109f6cd4e970c0ea369fcfedc4e43002908d467d99747711ad32</code><br>Computed UTF-8 SHA-256: <code>9ac3a385fcfa109f6cd4e970c0ea369fcfedc4e43002908d467d99747711ad32</code><br>Recorded bytes 649; embedded UTF-8 bytes 649</p><p><b>Complete bounded dataset metadata JSON; reduced fields for S075–S077</b></p><pre>{
  &quot;_id&quot;: &quot;691cf26d01ad6c17d5c16c98&quot;,
  &quot;id&quot;: &quot;allenai/dolma3_dolmino_mix-100B-1125&quot;,
  &quot;sha&quot;: &quot;f23aa129fda8335ba9760057bcc1f0c02f3d068b&quot;,
  &quot;gated&quot;: false,
  &quot;cardData&quot;: {
    &quot;license&quot;: &quot;odc-by&quot;,
    &quot;language&quot;: [
      &quot;en&quot;
    ],
    &quot;configs&quot;: [
      {
        &quot;config_name&quot;: &quot;default&quot;,
        &quot;data_files&quot;: [
          {
            &quot;split&quot;: &quot;train&quot;,
            &quot;path&quot;: &quot;data/**/*&quot;
          }
        ],
        &quot;features&quot;: [
          {
            &quot;name&quot;: &quot;id&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;text&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;metadata&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;source&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;version&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;created&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;added&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;doc&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          },
          {
            &quot;name&quot;: &quot;attributes&quot;,
            &quot;dtype&quot;: &quot;string&quot;
          }
        ]
      }
    ]
  },
  &quot;lastModified&quot;: &quot;2026-02-23T19:03:37.000Z&quot;
}</pre></details><details id="S077"><summary>S077 · datasets/allenai--dolma3_longmino_mix-100B-1125-summary.json · exact</summary><p><a href="https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified" target="_blank" rel="noopener noreferrer">https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125?expand%5B%5D=sha&amp;expand%5B%5D=gated&amp;expand%5B%5D=cardData&amp;expand%5B%5D=lastModified</a></p><p class="small">Retrieved 2026-10-08T09:04:27.068701+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>9e6c8bdff5506a231083f09ca5f06ed7c60dac0562ca9121b7a7c203f59bb53c</code><br>Computed UTF-8 SHA-256: <code>9e6c8bdff5506a231083f09ca5f06ed7c60dac0562ca9121b7a7c203f59bb53c</code><br>Recorded bytes 324; embedded UTF-8 bytes 324</p><p><b>Complete bounded dataset metadata JSON; reduced fields for S075–S077</b></p><pre>{
  &quot;_id&quot;: &quot;691cf13196892f484b29fe91&quot;,
  &quot;id&quot;: &quot;allenai/dolma3_longmino_mix-100B-1125&quot;,
  &quot;sha&quot;: &quot;28fea4330d8f8e27221010d42c4bc53ba9ec3236&quot;,
  &quot;gated&quot;: false,
  &quot;cardData&quot;: {
    &quot;license&quot;: &quot;odc-by&quot;,
    &quot;language&quot;: [
      &quot;en&quot;
    ],
    &quot;configs&quot;: [
      {
        &quot;config_name&quot;: &quot;default&quot;,
        &quot;data_files&quot;: [
          {
            &quot;split&quot;: &quot;train&quot;,
            &quot;path&quot;: &quot;data/**/*&quot;
          }
        ]
      }
    ]
  },
  &quot;lastModified&quot;: &quot;2026-02-24T02:07:54.000Z&quot;
}</pre></details><details id="S078"><summary>S078 · deepseek-ai--DeepSeek-V4-Pro/inference/README.md · exact</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/inference/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/inference/README.md</a></p><p class="small">Retrieved 2026-10-08T09:04:27.377171+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>68dba94f8676578cddff2b0e8861586ef89d1857c6ad29e40bf5f17610b03bdf</code><br>Computed UTF-8 SHA-256: <code>68dba94f8676578cddff2b0e8861586ef89d1857c6ad29e40bf5f17610b03bdf</code><br>Recorded bytes 951; embedded UTF-8 bytes 951</p><p><b>content lines 1-26</b></p><pre># Inference code for DeepSeek models

First convert huggingface model weight files to the format of this project.
```bash
export EXPERTS=384
export MP=8
export CONFIG=config.json
python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} --n-experts ${EXPERTS} --model-parallel ${MP}
```

Then chat with DeepSeek model at will!
```bash
torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --interactive
```

Or batch inference from file.
```bash
torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}
```

Or multi nodes inference.
```bash
torchrun --nnodes ${NODES} --nproc-per-node $((MP / NODES)) --node-rank $RANK --master-addr $ADDR generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}
```

If you want to use fp8, just remove `&quot;expert_dtype&quot;: &quot;fp4&quot;` in `config.json` and specify `--expert-dtype fp8` in `convert.py`.</pre></details><details id="S079"><summary>S079 · deepseek-ai--DeepSeek-V4-Pro/encoding/README.md · exact</summary><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/encoding/README.md" target="_blank" rel="noopener noreferrer">https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/encoding/README.md</a></p><p class="small">Retrieved 2026-10-08T09:04:27.550272+00:00 · Recorded HTTP 200<br>Recorded SHA-256: <code>605363e9e43ee91beba88ea96c7806ce6ecdb2924e481459c9d16e1526470c10</code><br>Computed UTF-8 SHA-256: <code>605363e9e43ee91beba88ea96c7806ce6ecdb2924e481459c9d16e1526470c10</code><br>Recorded bytes 8118; embedded UTF-8 bytes 8118</p><p><b>content lines 1-42</b></p><pre># DeepSeek-V4 Encoding

This document describes the prompt encoding format used by DeepSeek-V4 series models. The encoding handles multi-turn conversations, tool calling, extended thinking (reasoning), and quick instruction tasks.

A self-contained reference implementation is provided in `encoding_dsv4.py`.

## Quick Start

```python
from encoding_dsv4 import encode_messages, parse_message_from_completion_text

# Encode a conversation
messages = [
    {&quot;role&quot;: &quot;system&quot;, &quot;content&quot;: &quot;You are a helpful assistant.&quot;},
    {&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;What is 2+2?&quot;},
]
prompt = encode_messages(messages, thinking_mode=&quot;thinking&quot;)
# =&gt; &quot;&lt;｜begin▁of▁sentence｜&gt;You are a helpful assistant.&lt;｜User｜&gt;What is 2+2?&lt;｜Assistant｜&gt;&lt;think&gt;&quot;

# Parse model output back to structured message
completion = &quot;Simple arithmetic.&lt;/think&gt;2 + 2 = 4.&lt;｜end▁of▁sentence｜&gt;&quot;
parsed = parse_message_from_completion_text(completion, thinking_mode=&quot;thinking&quot;)
# =&gt; {&quot;role&quot;: &quot;assistant&quot;, &quot;reasoning_content&quot;: &quot;Simple arithmetic.&quot;, &quot;content&quot;: &quot;2 + 2 = 4.&quot;, &quot;tool_calls&quot;: []}
```

&gt; **Note:** The `parse_message_from_completion_text` function is designed to handle well-formatted model output only. It does not attempt to correct or recover from malformed output that the model might occasionally generate. For production use, additional error handling is recommended.

## Message Format

### Special Tokens

| Token | Purpose |
|-------|---------|
| `&lt;｜begin▁of▁sentence｜&gt;` | Beginning of sequence (BOS) |
| `&lt;｜end▁of▁sentence｜&gt;` | End of assistant turn (EOS) |
| `&lt;｜User｜&gt;` | User turn prefix |
| `&lt;｜Assistant｜&gt;` | Assistant turn prefix |
| `&lt;｜latest_reminder｜&gt;` | Latest reminder (date, locale, etc.) |
| `&lt;think&gt;` / `&lt;/think&gt;` | Reasoning block delimiters |
| `｜DSML｜` | DSML markup token |

### Roles</pre></details><script>document.getElementById('filter').addEventListener('input',function(){const q=this.value.toLowerCase();document.querySelectorAll('tr[data-model]').forEach(r=>r.hidden=!r.dataset.model.includes(q));});document.querySelectorAll('a[href^="#S"]').forEach(a=>a.addEventListener('click',()=>{const d=document.getElementById(a.getAttribute('href').slice(1));if(d)d.open=true;}));</script></main></html>