# Optotagging analysis protocol and required data extract

Status: fixed before accessing recording spike data or computing responses, 2026-10-09. No recording-level analysis has been performed. The supplied tutorial includes illustrative outputs; these are not results of this analysis. Network access to the DANDI API was denied. The supplied source guide is retained as the original session upload.

## Sources and provenance

DANDI dandiset 000022, Allen Institute - Visual Coding - Neuropixels (Functional Connectivity Stimulus Set), fixed release 0.251116.2247. Release creation: 2025-11-16T22:47:01.736619Z (from supplied listing). This is NOT the recording date.

- SST: subject 774672354, session 794812542; path sub-774672354/sub-774672354_ses-794812542.nwb; asset 64ea7bdf-dd18-4831-ab01-b9a035a3da2f; size 2,589,149,988 bytes.
- Wild-type: subject 753795604, session 767871931; path sub-753795604/sub-753795604_ses-767871931.nwb; asset 1acc9a81-e67f-4586-8e59-561ba507b489; size 2,644,247,736 bytes.

Both source listings have complete pagination (next=null). The tutorial associates the two session files with SST and wild-type, respectively; confirm using each file's subject genotype. The separate probe ecephys NWB files are not required for the spike-table analysis.

Source guide upload SHA-256: c5de1cc8eae64fc159ded46156c3597997c4d6e305804ef41f13635a03f2afac.
Tutorial: https://alleninstitute.github.io/openscope_databook/optotagging/
Recording dates, NWB format version, source checksums, genotype strings and actual stimulus settings remain unverified.

## Fixed cohort, quality and stimulus rules

Use these two sessions only, one animal per group. These are case studies, not biological replication for genotype inference. Analyze all anatomical regions with qualifying units and retain region/probe labels; do not call all units in the SST mouse SST neurons.

Include units with source quality label good, finite amplitude_cutoff < 0.1, presence_ratio > 0.95, and isi_violations < 0.5. Do not filter on firing rate, waveform width or observed light response. Missing required quality information makes a unit ineligible for the primary analysis; report the exclusion rather than substituting an undocumented metric. Retain a per-unit exclusion audit and source metric definitions. Allen's isi_violations metric must not be substituted with a similarly named percentage without documenting its equivalence.

Primary stimuli: isolated, single square 10 ms pulses (duration tolerance 0.1 ms), not pulse trains or ramps. Analyze all qualifying levels separately. The highest recorded level within each session is the prespecified primary level; all lower levels are sensitivity analyses. If that level has fewer than 20 usable trials, mark the primary analysis unavailable rather than selecting a more favorable level. Command voltage is not delivered optical power: do not assume nominal levels make the sessions power-matched.

A usable event must have known onset and offset, complete unit observation coverage for [onset-0.100 s, onset+0.100 s), and no overlap with a source invalid interval. Reject events if any other stimulation interval overlaps that window. Unit-specific observation gaps invalidate that unit/event pair. Require at least 20 usable trials per included unit. Do not infer complete observation from absence of spikes.

## Primary response definition

Use half-open windows in seconds relative to onset:
- baseline [-0.008, -0.002);
- response [0.002, 0.008).

Both exposures are 6 ms per trial. Report trial-level counts, rates in Hz, paired rate difference, response/baseline ratio (undefined when baseline is zero), and fraction of trials containing at least one spike in each window.

For each eligible unit at the primary level, test increased post-stimulation count with a paired, one-sided sign-flip randomization test of the mean trial count difference. Enumerate all sign assignments when feasible (at most 20 nonzero differences); otherwise use 99,999 assignments, seed 20261009, with (b+1)/(B+1) p-values. Apply Benjamini-Hochberg FDR at 0.05 jointly across eligible units in both sessions. The responder rule is q <= 0.05, post-stimulation rate >= 2 times baseline rate, and an absolute increase >= 10 Hz. Zero-baseline units can satisfy this rule through the absolute increase criterion; do not report infinite fold changes.

Lower-level analyses are secondary; adjust over all unit-by-lower-level tests jointly and label them separately. No genotype p-value will treat neurons or trials as independent animals.

## Latency and individual trials

Draw rasters for every qualifying unit, paginated, with trial order preserved, light interval [0,10) ms, and artifact masks clearly distinguished. Provide a wider [-100,100) ms view and a [-10,25) ms detail view. Excluded-interval spikes can appear as gray context but must never enter response statistics. Never select only attractive rasters: the overview uses up to three units per session nearest the 25th, 50th and 75th percentiles of eligible primary rate differences (ties by unit ID); all eligible units remain in the atlas.

For primary responders, quantify an operational onset using six 1 ms bins over [2,8) ms, paired with baseline bins over [-8,-2) ms. Swap the complete baseline/response six-bin vectors within trials, using 99,999 random assignments with the fixed seed, to obtain the null distribution of the maximum positive mean binwise count difference. The onset is the left edge of the earliest pair of consecutive response bins both exceeding the 99th percentile of this maximum-bin null distribution. Otherwise onset is not resolved. This is a conservative operational, per-unit onset criterion; it does not establish direct optogenetic activation. Label it conditional on responder selection, with 1 ms resolution and a 2 ms detection floor. Do not assign zero latency or impute a latency to nonresponders.

Also retain the first spike in [2,8) ms on every trial, using missing for no spike. Show their distribution and report median, interquartile range, responding-trial fraction and a 95% percentile bootstrap interval for the conditional median (2,000 trial-resampling replicates; seed 20261009). This conditional first-spike statistic can contain spontaneous spikes and is not interchangeable with population onset latency.

## Light artifacts and prespecified sensitivity

Exclude [-2,2) ms around onset and [8,12) ms around the 10 ms pulse offset from inference. The primary response ends at 8 ms to avoid the offset artifact. Preserve rather than delete those spikes in the source extract. Show onset/offset transients and WT responses as artifact diagnostics, without using them to tune the primary masks.

Repeat response tests with a stricter 3 ms onset exclusion: response [3,8) ms against baseline [-7,-2) ms, both 5 ms long; apply the same effect thresholds and FDR family. Report changed classifications and all unresolved latencies.

A clean WT comparison does not prove the absence of artifacts in SST. No-waveform analysis cannot decisively distinguish residual photoelectric contamination, sorting errors, direct activation and network-mediated activity. Explicitly qualify any positive finding. Waveform/raw-trace checks are optional supplementary evidence, not an excuse to silently change the primary rules.

## Exact small-file extract needed

Supply one ZIP containing the following files. CSV/Parquet plus JSON is sufficient; no full recordings, LFP, behavioral movies or continuous acquisitions are needed.

### sessions.json (one record per session)

Required: session_id, subject_id, exact genotype, session_start_time including timezone, timestamps_reference_time (if present), identifier, session_description, NWB format version, DANDI dandiset/version/asset_id/path, original asset size, published checksum and checksum algorithm (if available), extraction timestamp and extraction software/version.

Read recording time from NWB session_start_time; do not substitute file_create_date, DANDI created/modified dates or release dates. Include subject species/sex/age and stimulus device, wavelength, calibration and level units when present. Explicitly mark absent metadata.

### stimuli.csv (all optotagging events, all conditions and levels)

Columns: session_id, stimulus_id, start_time_s, stop_time_s, duration_s, condition, stimulus_name, level, level_units. Preserve original row identifiers, condition labels and any pulse-train parameters. Do not export only selected events: the complete table is required to detect nearby pulses and choose conditions without consulting spike responses.

The tutorial's table is processing/optotagging/optogenetic_stimulation. It notes that newer files can store these as interval presentation tables. Export the actual source table name and original column names in a schema mapping; do not assume an alternative path exists. If individual pulse edges differ from table intervals, include the true edges or stimulus waveform timebase needed to recover them.

### units.csv (all units, not just responsive units)

Columns: session_id, unit_id, quality, amplitude_cutoff, presence_ratio, isi_violations, peak_channel_id, probe_id, anatomical_structure, firing_rate. Export all available additional quality fields, e.g. snr, isolation_distance, d_prime, nn_hit_rate and nn_miss_rate; these are audit fields, not additional primary thresholds. Required QC columns absent in the source must be identified as absent.

From /units, preserve row identifiers and the mapping from unit peak channel/electrode references to probe and anatomical structure in the electrode table. Include each unit's observation intervals separately as observations.csv: session_id, unit_id, start_time_s, stop_time_s. If the source has no observation intervals, provide its documented continuous recording bounds and relevant probe availability; do not fabricate obs_intervals.

### spikes.csv or spikes.parquet

Columns: session_id, unit_id, spike_time_s. Export spikes from the UNION of [t-0.100 s, t+0.100 s) for EVERY isolated or candidate single 10 ms pulse onset t, at EVERY level, before quality filtering. Preserve absolute NWB-clock seconds at full available precision. Include units with zero spikes through units.csv; an absent spike row must not imply an absent unit. Deduplicate overlapping extraction windows. Do not blank light artifacts or apply a response filter during extraction.

The NWB storage is normally /units/spike_times with the cumulative ragged-array /units/spike_times_index. Read metadata and indices first, then retrieve only relevant per-unit spike slices via HTTP byte ranges using an HDF5-compatible reader. Binary search within sorted unit spike arrays, or read one unit at a time and discard times outside the event windows. A compressed HDF5 chunk may require transfer of more bytes than the retained extract; record bytes transferred if supported. Do not download full files as an automatic fallback.

If candidate conditions cannot be identified from metadata, first supply stimuli.csv and the metadata/QC tables; no response data are needed to resolve stimulus identity. If streaming is unavailable to the exporter, the same small extract can be made from a locally held NWB.

### invalid_intervals.csv and schema.json

Invalid intervals: session_id, start_time_s, stop_time_s, reason, affected_probe_or_unit (if present). Supply headers plus an explicit no-invalid-intervals statement if none are recorded. schema.json must describe clock origin, time units, source NWB paths/columns, array-to-unit mappings, null conventions, and any missing required fields. Include a SHA-256 manifest for every delivered extract.

Optional waveform audit: unit mean waveforms with sampling rate, time axis, electrode mapping and amplitude units; these alone do not establish peri-light waveform stability. If available, small raw voltage segments [-5,+15] ms around the first five chronological 10 ms trials at each level, on included units' peak channels and adjacent contacts, plus spike assignments and scaling, support supplementary artifact inspection. These are not required to start the table-based analysis.

## Planned retained outputs and completion checks

Keep the provided source archive, immutable source manifest, this protocol, extracted tables and checksum manifest, executable analysis code and environment versions, exclusion audit, trial-level counts/first-spike latencies, per-unit statistics, full raster atlas, latency summary figures and an interpretation report.

Before declaring completion, check unique IDs, sorted finite spike times, clock alignment, duration/offset consistency, observation coverage, non-overlap/invalid intervals, exact retained trial counts and spot-checked hand counts. Visually inspect final figures and verify every reported count against exported calculations. No response result, latency, recording date or figure is currently available.
