# Optotagging protocol amendment 01 — final after metadata review

Finalized 2026-10-09 before reading recording spike times or computing stimulation responses. This document supersedes the provisional optotagging-protocol-amendment-01.md and amends optotagging-protocol-and-extract-spec.md. Unchanged rules in the original protocol continue to apply.

## Evidence reviewed

Reviewed the attached optotagging-metadata.zip, SHA-256 a29c0563ec584d146fda5e3a1608c4f23b980e00fdaa2f1d349b5b43df8c9742. Inspected both metadata JSON files, remote schema inventories, stimulus and unit CSVs, electrode CSVs, and the extraction script as text. Did not execute the supplied extraction script or access remote data. The archive contains metadata and scalar unit metrics, not spike-time arrays.

The accompanying optotagging-metadata-verification.zip retains runnable verification code, software versions, input-member hashes, full per-unit eligibility audits, stimulus condition counts, exact event extraction windows and machine-readable findings. Verification applies to the supplied extract and its internal consistency; full remote NWB checksums were not independently recomputed. Source digests are preserved as reported.

## Verified recording identities

Both files identify NWB version 2.2.2 and DANDI release 0.251116.2247. Their asset IDs, subject IDs, session IDs, paths and sizes match the previously supplied DANDI 000022 listings.

| Field | SST | Wild-type |
|---|---|---|
| Session | 794812542 | 767871931 |
| Subject | 774672354 | 753795604 |
| Recording start, with source timezone | 2019-01-19T01:14:56-08:00 | 2019-01-02T15:20:27-08:00 |
| Genotype | Sst-IRES-Cre/wt;Ai32(RCL-ChR2(H134R)_EYFP)/wt | wt/wt |
| Sex; age | Female; P120D | Male; P135D |
| Species | Mus musculus | Mus musculus |
| Units in table | 2,688 | 2,685 |
| All optotagging events | 300 | 180 |

timestamps_reference_time equals session_start_time in each file. These recording dates are distinct from the November 2025 release dates and October 2026 metadata extraction timestamps. Animals differ in sex, age, recording date and stimulus levels; this two-animal comparison is not a controlled population genotype comparison.

SST asset: 64ea7bdf-dd18-4831-ab01-b9a035a3da2f; path sub-774672354/sub-774672354_ses-794812542.nwb; reported SHA-256 69e66bfe0c3a071a1fd45c4543ddedb6b18b36150450ab0a718991d3925d7d3d.
Wild-type asset: 1acc9a81-e67f-4586-8e59-561ba507b489; path sub-753795604/sub-753795604_ses-767871931.nwb; reported SHA-256 53b5b2f08ee38e08754c63b404488cba2fc67652fe8c70d4e921f6762f61653c.

## Verified stimuli and fixed level selection

The actual source table is processing/optotagging/optogenetic_stimulation. Both sessions contain single square pulses of 5 and 10 ms, one-second trains described as 2.5 ms pulses at 10 Hz, and one-second raised-cosine stimuli. The analysis selects only stimulus_name=pulse, condition=a single square pulse, duration within 0.1 ms of 10 ms.

| Group | Recorded level | 10 ms trials | Role |
|---|---:|---:|---|
| SST | 1.3 | 25 | Secondary |
| SST | 1.7 | 25 | Secondary |
| SST | 2.0 | 25 | Primary |
| Wild-type | 1.0 | 15 | Descriptive |
| Wild-type | 2.5 | 15 | Descriptive |
| Wild-type | 4.0 | 15 | Primary descriptive control |

Level units and delivered optical-power calibration are not established by this attachment. Retain these as recorded levels without assigning physical units or assuming matching power across sessions.

All stimulus IDs are unique within session; timing and level values are finite, durations positive, and stop-start agrees with duration within 1.82e-12 seconds. No other listed stimulation interval overlaps any selected event's [-100,+100) ms window. Minimum onset spacing over all events is 1.68005 s in SST and 1.69006 s in wild-type. These checks establish event isolation, not recording coverage.

Keep the original minimum of 20 usable trials PER UNIT AND LEVEL for inference. Never pool light levels, conditions or animals to reach it. SST has 25 scheduled trials per level and must retain at least 20 after any coverage/invalid-interval exclusions. If the highest level becomes ineligible, do not promote a lower level to primary.

Wild-type has only 15 trials per level and is descriptive only: retain rasters, unsmoothed 1 ms peri-event rates, before/after rates, paired differences, fractions of trials with spikes, and conditional first-spike distributions in [2,8) ms with median, interquartile range and observed trial counts. Do not assign inferential responder/nonresponder labels, response p/q values, significant onset estimates or responder percentages. Use “descriptive only: fewer than 20 usable trials,” not “no response.” Wild-type artifacts remain diagnostically useful, but a weak/absent signal cannot establish artifact absence in SST.

The primary FDR family therefore contains all eligible SST primary-level unit tests. Do not insert artificial wild-type p=1 values. Keep the separate joint family for eligible lower-level unit tests. No between-genotype significance test treats units or trials as independent animals.

## Quality audit and additional source-flag exclusion

Retain the original rules: quality=good, finite amplitude_cutoff<0.1, presence_ratio>0.95, isi_violations<0.5. No response, firing-rate or waveform-width filter is added.

All required QC fields are present with no missing values. Names match the original Allen fields, but their in-file descriptions say “no description”; a metric-definition validation against external documentation has not been performed. Use values as supplied, without percent conversion or substitution of similarly named fields.

The original conjunction passes 999 SST and 703 wild-type units. All unit peak_channel_id values map uniquely to the supplied electrode IDs. A newly observed source flag requires an explicit amendment: exclude a unit if its mapped peak electrode has valid_data=false; also exclude unmapped or unknown-validity peak electrodes if encountered later. This is a conservative source-validity rule, not proof about the timing of a recording gap. The supplied tables contain four original-QC-passing SST units and one original-QC-passing wild-type unit on false-flagged electrodes.

Final metadata-eligible candidates: 995 SST and 702 wild-type units. These are NOT responsive-unit counts and remain subject to event coverage checks. The full audit retains original flags and every exclusion. Keep all units in the spike extract when feasible, including excluded and silent units, to preserve the source audit.

## Coverage: evidence, assumption, and remaining dependency

The supplied unit-column/schema inventories have no obs_intervals. The metadata's interval-table lists have no invalid_times table; no invalid-times export is present. Describe this as “no invalid_times table recorded in the supplied inventory,” not “no acquisition gaps.” Probe metadata supplies IDs and LFP sampling rates but not acquisition start/stop bounds or interruption records. The existence of an acquisition group, session_start_time, valid_data=true, high presence_ratio, or a stimulus event does not prove continuous per-unit spike detection.

Accordingly, present coverage_status is unverified for every candidate unit/event. This attachment cannot establish uninterrupted acquisition.

The finalized assumption for later inferential work is: a quality-passing unit on a valid peak electrode is assumed observed throughout a selected [onset-0.100 s,onset+0.100 s) window if that window lies within documented acquisition coverage for the corresponding probe (or documented applicable session coverage), does not intersect known gaps/invalid intervals, and the extractor confirms complete retrieval of the requested spike slices. This remains assumed per-unit observation, not measured obs_intervals. Label qualifying observations coverage_status=assumed_from_acquisition_metadata and name the coverage evidence.

Before inferential analysis, request acquisition bounds and known interruption information plus an extraction-completeness statement. If only an extractor statement is available, distinguish complete extraction from evidence that acquisition occurred; neither substitutes automatically for the other. If supporting acquisition bounds/coverage evidence remain unavailable, retain descriptive outputs only and label coverage unverified. Do not silently manufacture intervals or interpret an unverified missing window as a zero count.

Do not estimate coverage from a unit's first/last spike, infer a gap from silence, or exclude low-response trials. Once a window is confirmed completely extracted and included under the declared coverage assumption, no spikes in that window is a valid zero. Retain per-unit/per-trial observation flags and report actual denominators. If coverage fails for some events, apply exclusions without consulting responses.

## Extraction instructions refined by verified metadata

Use the exact per-event CSVs in the verification archive. Retain absolute spike times from the UNION of the 200 ms windows, not every time between the earliest and latest window:
- SST: 75 windows, 15.0 s total retained exposure per unit; enclosing bounds 9212.8594 to 9778.89404 seconds in the NWB clock.
- Wild-type: 45 windows, 9.0 s total retained exposure per unit; enclosing bounds 9365.53193 to 9704.12674 seconds.

The enclosing bounds are extraction envelopes, NOT recording start/stop or evidence of coverage. Use all three levels separately. Preserve artifact-window spikes and units with zero spikes.

The source layout reports contiguous, uncompressed float64 spike_times datasets and int64 cumulative spike_times_index arrays, so a range-capable HDF5 reader can retrieve the index and selected per-unit time slices without downloading full 2.6 GB files. Layout information alone does not prove current server range support or validate a future extract. Keep source row/ID mapping, sorted-time checks, extraction-completeness records and checksums. No network retrieval was attempted during this review.

## Unchanged response and artifact rules

Retain the primary equal-duration baseline [-8,-2) ms and response [2,8) ms; onset mask [-2,2) ms and offset mask [8,12) ms; original response effect thresholds, paired randomization, FDR, operational SST onset definition and fixed seed. Retain the stricter [3,8) ms versus [-7,-2) ms sensitivity analysis. Show all eligible-unit rasters with excluded intervals distinguished. Sub-2 ms latencies cannot be resolved, and sorted spikes without appropriate waveform/raw-trace evidence cannot decisively exclude light artifacts or establish direct optogenetic activation.

Metadata review is complete for the supplied attachment. The amendment is finalized; acquisition coverage evidence and the spike extract remain outstanding. No spike-response analysis, response classifications or latency measurements have been performed.
