﻿audit_date,evidence_basis,model,repository,revision,snapshot_gated,listed_weight_shards,weights_status,weights_judgment,weights_sources,weights_urls,configuration_status,configuration_judgment,configuration_sources,configuration_urls,tokenizer_status,tokenizer_judgment,tokenizer_sources,tokenizer_urls,inference_status,inference_judgment,inference_sources,inference_urls,training_library_status,training_library_judgment,training_library_sources,training_library_urls,training_recipe_status,training_recipe_judgment,training_recipe_sources,training_recipe_urls,data_description_status,data_description_judgment,data_description_sources,data_description_urls,training_data_status,training_data_judgment,training_data_sources,training_data_urls,evaluation_status,evaluation_judgment,evaluation_sources,evaluation_urls,license_status,license_judgment,license_sources,license_urls,reproduction_gaps
2026-10-08,"User-supplied snapshots, not fetched in-session",Mistral Large 4,,,,,R,October 6 announcement promises weights by month-end (October 2026); preview API is available. No weight listing supplied.,S033,https://mistral.ai/news/mistral-large-4/,U,No configuration file or model repository listing in the supplied evidence; architecture details promised.,S033,https://mistral.ai/news/mistral-large-4/,U,No tokenizer files or listing supplied.,S033,https://mistral.ai/news/mistral-large-4/,U,Preview API described; downloadable model-specific inference code not verified.,S033,https://mistral.ai/news/mistral-large-4/,U,Mistral Forge named as the training/customization/RL environment; source release not verified.,S033,https://mistral.ai/news/mistral-large-4/,U,"No original training pipeline inspected. Further post-training methodology is promised, not a promise of training code.",S033,https://mistral.ai/news/mistral-large-4/,P,Broad description: substantial multilingual data covering 160+ languages; no exact corpus manifest or mixture.,S033,https://mistral.ai/news/mistral-large-4/,U,No downloadable training dataset evidenced; no dataset repository snapshot supplied.,S033,https://mistral.ai/news/mistral-large-4/,P,"Benchmarks and selected methodological descriptions, but no runnable evaluation package inspected; further benchmarks promised.",S033,https://mistral.ai/news/mistral-large-4/,U,No license for the promised weight release verified. Marketing use of open-weight does not establish terms.,S033,https://mistral.ai/news/mistral-large-4/,"First obtain the promised weights and their license, config/tokenizer and compatible runtime. Training reproduction additionally needs the corpus, preprocessing, full pre/post-training code and settings, checkpoints, and evaluation assets. An API preview cannot supply these."
2026-10-08,"User-supplied snapshots, not fetched in-session",DeepSeek-V4-Pro,deepseek-ai/DeepSeek-V4-Pro,b5968e9190ef611bbf34a7229255be88a0e937c1,False,64,A,"64 safetensors shards and index listed; public, gated=false. Config specifies FP8 quantization, not an all-BF16 checkpoint.",S001 S004,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/config.json,A,config.json read; generation_config.json listed. deepseek_v4 architecture.,S001 S004,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/config.json,A,tokenizer.json listed and tokenizer_config.json read. No Jinja template by design: custom encoding scripts/tests and documentation are supplied.,S001 S005 S003 S079,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/tokenizer_config.json | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/encoding/README.md,A,"inference/model.py, kernel.py, convert.py and generate.py listed; README provides conversion and torchrun commands. Source bodies of those Python files were not supplied.",S001 S078,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/inference/README.md,N,"No training implementation in the inspected model repository listing; Muon/GRPO are method descriptions, not released training code.",S001 S003,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md,N,"No original pretraining/post-training recipe in the inspected release. Card describes domain experts, SFT/RL and distillation only.",S001 S003,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md,P,"Card reports >32T diverse high-quality tokens and post-training stages; exact sources, filtering and mixtures unspecified in inspected text.",S003,https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md,N,No corpus shards or downloadable corpus release identified in the inspected card/listing; availability elsewhere remains unknown.,S001 S003,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md,P,Benchmark tables with shot counts/metrics and encoding instructions; no full benchmark runner/config package in the listing.,S001 S003 S079,https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Pro | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/encoding/README.md,A,MIT text inspected; card explicitly applies it to repository and weights. License hash matches after exact CRLF reconstruction.,S002 S003,https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/LICENSE | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/b5968e9190ef611bbf34a7229255be88a0e937c1/README.md,"Inference materials are substantial, but need compatible hardware/runtime, full shard acquisition, and the custom encoder. Retraining still needs the >32T corpus and processing, Muon and distributed-training configuration, domain SFT/GRPO data and rewards, consolidation/distillation recipe and teachers, run state, and exact evaluation harnesses."
2026-10-08,"User-supplied snapshots, not fetched in-session",Qwen3.8-27B,Qwen/Qwen3.8-27B,1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0,False,18,A,"18 safetensors shards and index listed; public, gated=false. This is the post-trained vision-language model; a coming-soon hosted service does not make the weights promised.",S006 S008,https://huggingface.co/api/models/Qwen/Qwen3.8-27B | https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md,A,config.json read; image/video preprocessing and generation configurations listed. Model reuses qwen3_5 architecture identifiers.,S006 S009,https://huggingface.co/api/models/Qwen/Qwen3.8-27B | https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/config.json,A,"tokenizer.json, vocab.json, merges.txt and chat_template.jinja listed; tokenizer_config.json read.",S006 S010,https://huggingface.co/api/models/Qwen/Qwen3.8-27B | https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/tokenizer_config.json,P,"Public usage/serving examples and links to Transformers, SGLang, vLLM and TokenSpeed; no self-contained inference implementation in either inspected release tree. Linked runtime source not inspected.",S008 S064 S067,https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md | https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1 | https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md,P,"Generic Unsloth, Swift and Llama-Factory fine-tuning recommendations; linked libraries not inspected. They are not the original pipeline.",S067,https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md,N,No original model training recipe in the inspected HF/GitHub trees. Generic fine-tuning guidance is excluded.,S006 S064 S067,https://huggingface.co/api/models/Qwen/Qwen3.8-27B | https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1 | https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md,U,Exact 3.8-27B training corpus/mixture not described in inspected material. The family README's trillions-of-tokens discussion is under Qwen3.5 and is not evidence of this release's exact data.,S008 S067,https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md | https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md,N,No training corpus release in inspected model/project listings or cards; other sources unverified.,S006 S064 S067,https://huggingface.co/api/models/Qwen/Qwen3.8-27B | https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1 | https://raw.githubusercontent.com/QwenLM/Qwen3.8/2ea10dc725823bf7c3e21ce8557cbe15245132ae/README.md,P,"Benchmark footnotes identify prompts, harnesses, judges and some corrected labels; no complete runnable evaluation package or corrected-label files inspected.",S008 S064,https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md | https://api.github.com/repos/QwenLM/Qwen3.8/git/trees/2ea10dc725823bf7c3e21ce8557cbe15245132ae?recursive=1,A,Apache-2.0 LICENSE text inspected; recorded hash reproduced with CRLF. Model card also requires CRLF reconstruction for its recorded hash.,S007 S008,https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE | https://huggingface.co/Qwen/Qwen3.8-27B/raw/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/README.md,"Acquire all weights and multimodal preprocessing/runtime dependencies first. Retraining needs exact vision-language data, processing/mixtures, pretraining and post-training run configs, teacher/reward/agent data and environments. Score reproduction needs the corrected benchmark annotations and precise harness/judge versions; generic SFT examples do not fill these gaps."
2026-10-08,"User-supplied snapshots, not fetched in-session",GLM-5.3,zai-org/GLM-5.3,aca966e4e02791568aa6a4ced368624b3d897f42,False,141,A,"141 safetensors shards and index listed; public, gated=false. Config includes FP8 quantization.",S011 S014,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/config.json,A,config.json read; glm_moe_dsa architecture and generation config listed.,S011 S014,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/config.json,A,tokenizer.json and chat_template.jinja listed; tokenizer_config.json read.,S011 S015,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/tokenizer_config.json,P,"Serving instructions link SGLang, vLLM and TokenSpeed; no standalone implementation in the inspected GLM-5 project tree. Linked engines not inspected.",S013 S062 S065,https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md | https://api.github.com/repos/zai-org/GLM-5/git/trees/c8ad661c6cf4cb0a78064987bc42f97e14355929?recursive=1 | https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md,P,Family README links slime asynchronous RL infrastructure for GLM-5; slime source is not in this snapshot and does not establish the GLM-5.3 recipe.,S065,https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md,N,No exact GLM-5.3 original pipeline in inspected trees. Card says it shares GLM-5.2's base and gains come from post-training.,S011 S013 S062 S065,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md | https://api.github.com/repos/zai-org/GLM-5/git/trees/c8ad661c6cf4cb0a78064987bc42f97e14355929?recursive=1 | https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md,U,Exact GLM-5.3 data not verified. GLM-5 and GLM-5.3-Flash corpus figures in the family README are not attributed to this exact release.,S013 S065,https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md | https://raw.githubusercontent.com/zai-org/GLM-5/c8ad661c6cf4cb0a78064987bc42f97e14355929/README.md,N,No training corpus found in inspected release/project listings; no independent corpus inventory supplied.,S011 S062,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://api.github.com/repos/zai-org/GLM-5/git/trees/c8ad661c6cf4cb0a78064987bc42f97e14355929?recursive=1,P,"Detailed benchmark footnotes specify sampling, harness versions, timeouts and modifications; result YAMLs listed. Full patched harnesses/judges not inspected.",S011 S013,https://huggingface.co/api/models/zai-org/GLM-5.3 | https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/README.md,A,"Custom GLM-5.3 license, not MIT: model-as-a-service operators with aggregate affiliate revenue >$10B over any consecutive 12 months require Z.AI security review before commercial use.",S012,https://huggingface.co/zai-org/GLM-5.3/raw/aca966e4e02791568aa6a4ced368624b3d897f42/LICENSE,"Need the GLM-5.2 base run/data lineage plus the actual 5.3 SFT/RL data, reward functions, orchestration and run settings. The slime link is insufficient. Evaluation needs exact modified containers/harnesses and judge access; usage must account for the custom license condition."
2026-10-08,"User-supplied snapshots, not fetched in-session",Kimi-K3,moonshotai/Kimi-K3,f831ab66814297da540d832a5235f8e904f29d06,False,96,A,"96 safetensors shards and index listed; public, gated=false. Card describes native MXFP4 weights/MXFP8 activations with QAT from SFT onward.",S016 S018,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md,A,"config.json read; custom configuration, processor and vision-processing Python files listed.",S016 S019,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/config.json,A,"tiktoken.model, tokenization_kimi.py and encoding_k3.py listed; tokenizer_config.json read.",S016 S020,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/tokenizer_config.json,A,Model implementation and processing/tokenization Python files listed; vLLM/SGLang/TokenSpeed deployment links provided. Python file bodies not supplied.,S016 S018 S019,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/config.json,N,No original training implementation in inspected model/project trees; architecture/inference modules do not constitute a trainer.,S016 S063,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1,N,No original pretraining/SFT/RL/QAT pipeline in inspected trees. Technical-report PDF is listed but its contents were not supplied.,S016 S018 S063,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md | https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1,U,Exact training data descriptions not verified from supplied card. Listed k3_tech_report.pdf was not inspected.,S018 S063,https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md | https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1,N,No training corpus files in the inspected model/project listings; availability elsewhere unknown.,S016 S063,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://api.github.com/repos/MoonshotAI/Kimi-K3/git/trees/3cb39dfd32e51c3328e2e4b4af21341247d06c43?recursive=1,P,Sampling settings and benchmark/harness notes supplied; result YAMLs listed. In-house tasks and adjusted GPU-task branches prevent claiming a full reproducible suite.,S016 S018,https://huggingface.co/api/models/moonshotai/Kimi-K3 | https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/README.md,A,Custom Kimi K3 license: >$20M aggregate revenue over any consecutive 12 months for MaaS operators triggers separate agreement; specified large commercial products require Kimi K3 UI attribution. Sections 2–3 have internal-use and official-product/certified-partner exceptions.,S017,https://huggingface.co/moonshotai/Kimi-K3/raw/f831ab66814297da540d832a5235f8e904f29d06/LICENSE,"Need the training corpus and its exact processing, full pretraining/post-training/QAT recipe, teachers/rewards, run state and compatible kernels. Evaluation needs Kimi Code setup, in-house tasks and the H20-calibrated task branch. Check the separate-agreement and attribution conditions before affected commercial use."
2026-10-08,"User-supplied snapshots, not fetched in-session",Gemma 4 31B-IT,google/gemma-4-31B-it,842da3794eaa0b77d5f08bae87a17459d91ff475,False,2,A,"2 safetensors shards and index listed; public, gated=false in this snapshot. Do not infer a gate from older Gemma releases.",S021,https://huggingface.co/api/models/google/gemma-4-31B-it,A,gemma4 config.json read; processor_config.json and generation config listed.,S021 S023,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/config.json,A,tokenizer.json and chat_template.jinja listed; tokenizer_config.json read.,S021 S024,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/tokenizer_config.json,P,Transformers loading/generation examples are public in the card. Runtime implementation is external and not supplied as source.,S022,https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,U,Original training implementation or relevant library source not inspected; no separate training repository snapshot supplied.,S021 S022,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,N,No original pretraining/instruction-tuning pipeline in the inspected model repo/card.,S021 S022,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,P,"Card describes web/code/math/images and multilingual data, January 2025 cutoff, and safety/quality preprocessing; no exact source manifest or mixture.",S022,https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,N,"Data descriptions supplied, not downloadable training shards. No corpus in the inspected model listing/card.",S021 S022,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,P,"Benchmark results and safety-evaluation approach supplied, plus one result YAML listing; full runnable evaluation recipe not inspected.",S021 S022,https://huggingface.co/api/models/google/gemma-4-31B-it | https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md,A,Apache-2.0 confirmed by the model card and the linked official Gemma 4 license page; that page's stored hash matches.,S022 S036,https://huggingface.co/google/gemma-4-31B-it/raw/842da3794eaa0b77d5f08bae87a17459d91ff475/README.md | https://ai.google.dev/gemma/docs/gemma_4_license,"Need exact multimodal corpus versions, preprocessing/filtering and mixture rules, original trainer and run settings, plus instruction-tuning/teacher/reward data and code. The published loading example supports inference, not retraining. Benchmark and safety results lack a fully inspected reproduction package."
2026-10-08,"User-supplied snapshots, not fetched in-session",NVIDIA Nemotron 3 Super 120B-A12B-BF16,nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16,2dc98e2afe4face0e4ce40972a915c45368bd34a,False,50,A,"50 safetensors shards and index listed; public, gated=false. This exact checkpoint is BF16; NVFP4 pretraining does not change its release identity.",S025 S027 S026,https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/config.json | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md,A,config.json read; configuration_nemotron_h.py and generation config listed.,S025 S027,https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/config.json,A,"tokenizer.json, special_tokens_map.json and chat_template.jinja listed; tokenizer_config.json read.",S025 S028,https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/tokenizer_config.json,A,modeling_nemotron_h.py and reasoning parser listed; card supplies vLLM/SGLang/TRT-LLM/Transformers examples. Implementation bodies not supplied.,S025 S026,https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md,A,"Nemotron developer tree has training/data-prep code, with Megatron-Bridge, NeMo-RL and related frameworks identified; recipe README contents inspected.",S050 S072 S073,https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ca8c409f08a9a5a5648d427383adea741b61966a?recursive=1 | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/README.md | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md,P,"Model-specific pretrain→SFT→RL→eval recipe and phase configs listed, with run commands; more than a generic library. It is an adapted public recipe, not a demonstrated exact replay of the original run.",S050 S072 S073 S026,https://api.github.com/repos/NVIDIA-NeMo/Nemotron/git/trees/ca8c409f08a9a5a5648d427383adea741b61966a?recursive=1 | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/README.md | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md,A,"Card names public, crawled, synthetic and explicitly private datasets; recipe describes 20T+5T pretraining and 34B+17B long-context stages.",S026 S073,https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md,P,"Some real dataset files evidenced: Cascade-RL-SWE lists three JSONL files, ungated. Full corpus is explicitly incomplete: recipe names internal-only code/crawl/academic categories and card lists private NVIDIA/third-party data. Other collections are links, not audited inventories.",S026 S037 S073,https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md | https://huggingface.co/api/datasets/nvidia/Nemotron-Cascade-RL-SWE | https://raw.githubusercontent.com/NVIDIA-NeMo/Nemotron/ca8c409f08a9a5a5648d427383adea741b61966a/src/nemotron/recipes/super3/stage0_pretrain/README.md,A,"Model-specific reproducibility tutorial and evaluation configuration paths supplied, with launch commands and benchmark settings. No runs performed; containers, credentials and exact BF16 endpoint equivalence still need validation.",S026 S049 S074,https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md | https://api.github.com/repos/NVIDIA-NeMo/Evaluator/git/trees/c87f9b21769cd1a7f3b85267264101a7bcc77df6?recursive=1 | https://raw.githubusercontent.com/NVIDIA-NeMo/Evaluator/c87f9b21769cd1a7f3b85267264101a7bcc77df6/packages/nemo-evaluator-launcher/examples/nemotron/nemotron-3-super/reproducibility.md,P,Card identifies NVIDIA Nemotron Open Model License. The supplied license-page content was read but does not match its recorded SHA-256/byte count; detailed legal terms remain integrity-unverified.,S026 S038,https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/raw/2dc98e2afe4face0e4ce40972a915c45368bd34a/README.md | https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/,"Restore the internal-only and licensed corpus categories, full SFT/RL data and generators, precise blend/iteration settings and original run lineage. Public recipe adaptation is required for missing data. Pin containers and validate the BF16 endpoint for evaluation; obtain a hash-consistent license capture before relying on detailed terms."
2026-10-08,"User-supplied snapshots, not fetched in-session",Olmo-3-1125-32B,allenai/Olmo-3-1125-32B,c2b61dae89a1ad10e4ad5653d0e46b590902607b,False,14,A,"14 safetensors shards and index listed; public, gated=false. This is the base model, not Think/Instruct. Intermediate checkpoints are described but their branch listing was not supplied.",S029 S030,https://huggingface.co/api/models/allenai/Olmo-3-1125-32B | https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md,A,olmo3 config.json read; generation_config.json listed.,S029 S031,https://huggingface.co/api/models/allenai/Olmo-3-1125-32B | https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/config.json,A,"tokenizer.json, vocab.json, merges.txt and special_tokens_map.json listed; tokenizer_config.json read.",S029 S032,https://huggingface.co/api/models/allenai/Olmo-3-1125-32B | https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/tokenizer_config.json,P,Transformers >=4.57 loading/generation code provided; OLMo-core source tree supplied. Runtime source required to execute that example not fully inspected.,S030 S048 S052,https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1 | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/README.md,A,"OLMo-core training source tree and README inspected; actual 32B stage training scripts supplied, beyond fine-tuning examples.",S048 S052 S068 S069 S070,https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1 | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/README.md | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B.py | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-midtraining.py | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-long-context.py,P,"Official staged 32B scripts/manifests listed, selected train scripts read, and stage budgets/merge method described. Exact release reconstruction needs reconciliation: card says final four checkpoints merged; official training README says three.",S030 S048 S068 S069 S070 S071,https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1 | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B.py | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-midtraining.py | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/train/OLMo3/OLMo3-32B-long-context.py | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/official/OLMo3/README.md,A,"Stage-specific Dolma/Dolmino/Longmino links, budgets and composition provided, with tokenized-data access instructions and manifest paths.",S030 S048 S071,https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://api.github.com/repos/allenai/OLMo-core/git/trees/5f6f58a133e7ef577d596295f2c8db4651c27857?recursive=1 | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/official/OLMo3/README.md,P,"Reduced dataset metadata says gated=false and gives data-file patterns; tokenized-data host/manifests documented. No dataset shard listing or shard content supplied. Pretrain URL returns id allenai/dolma3_mix-6T, not requested 5.5T-1125; exact run-data identity unresolved.",S030 S034 S035 S039 S071 S075 S076 S077,https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125 | https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125 | https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125 | https://raw.githubusercontent.com/allenai/OLMo-core/5f6f58a133e7ef577d596295f2c8db4651c27857/src/scripts/official/OLMo3/README.md | https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified | https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified | https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified,P,"Public OLMo-Eval tree, task/suite/inference commands and card results. Snapshot README is a general current harness, not a verified exact original 1125 evaluation invocation/config.",S030 S051 S055,https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://api.github.com/repos/allenai/OLMo-Eval/git/trees/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc?recursive=1 | https://raw.githubusercontent.com/allenai/OLMo-Eval/5ef9cee1cfa4eafd8bb624bd18e2817c74cd54cc/README.md,P,Apache-2.0 declared in model metadata/card; no model LICENSE file listed or license body supplied for this model. This is weaker than inspected license text. Dataset labels say ODC-BY; their terms were not supplied.,S029 S030 S075 S076 S077,https://huggingface.co/api/models/allenai/Olmo-3-1125-32B | https://huggingface.co/allenai/Olmo-3-1125-32B/raw/c2b61dae89a1ad10e4ad5653d0e46b590902607b/README.md | https://huggingface.co/api/datasets/allenai/dolma3_mix-5.5T-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified | https://huggingface.co/api/datasets/allenai/dolma3_dolmino_mix-100B-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified | https://huggingface.co/api/datasets/allenai/dolma3_longmino_mix-100B-1125?expand%5B%5D=sha&expand%5B%5D=gated&expand%5B%5D=cardData&expand%5B%5D=lastModified,"Resolve the 5.5T-1125→6T dataset identity, enumerate/pin all stage shards and official manifests, map supplied development scripts to official release scripts, replace or obtain cloud checkpoint/data paths, and reconcile final three-vs-four checkpoint averaging. Verify optimizer states, logs and exact evaluation settings. No Think/Instruct training is needed to reproduce this base model."
