# Do clear average optotagging responses occur on every pulse?

**No. In the highest-level SST condition, 22 of the 31 previously identified responders had at least one trial with no detected spike in the response window.** Their individual response probabilities ranged from 40% to 100%. Across this selected cohort, 101 of 775 unit–pulse observations had no spike, despite a mean response-window firing rate of 229.0 Hz versus 11.0 Hz in the matched baseline.

These are exploratory analyses of the uploaded recordings and previously analyzed data, not a new blinded or preregistered experiment. “Response” below means at least one sorted spike in **[2,8) ms after pulse onset**, conditional on uninterrupted unit observation. It does not mean a proven light-caused action potential.

## Recordings, validation and comparison

The uploaded extract identifies Allen Visual Coding Neuropixels recordings in [DANDI 000022, release 0.251116.2247](https://dandiarchive.org/dandiset/000022/0.251116.2247):

| Recording | Session | Animal | Eligible units | Isolated 10 ms pulses per level | Recorded levels |
|---|---|---|---:|---:|---|
| SST-ChR2 | 794812542 | Female, P120D; Sst-IRES-Cre/wt;Ai32/wt | 995 | 25 | 1.3, 1.7, 2.0 |
| Wild-type | 767871931 | Male, P135D; wt/wt | 702 | 15 | 1.0, 2.5, 4.0 |

Levels are source-table values; physical units and optical-power calibration are not established. Values are not assumed comparable between recordings. Other pulse durations, trains and cosine stimuli are outside this extract’s spike-analysis scope.

All 8 outer-manifest entries and 36 source-manifest entries passed SHA-256 verification. Spike indexing, unit alignment, finite sorted timestamps, extraction-window bounds, event isolation and probe acquisition endpoints passed structural checks. Recalculated rates and spike-containing trial fractions reproduce the earlier analysis, with maximum rate differences below 3 × 10⁻¹⁴ Hz.

The original unit filters were retained: good quality, finite amplitude cutoff <0.1, presence ratio >0.95, ISI violations <0.5, and a valid peak electrode. No firing-rate or response-based exclusion was added. Prior responder labels were carried forward from the supplied reference results; their p/q calculations were not independently rerun.

Each response window was compared with **[−8,−2) ms before that same pulse**, an equally long 6 ms baseline. Counts include zero-spike trials. Matched baseline controls for local spontaneous activity but is not a randomized no-light intervention. The stricter sensitivity comparison uses [3,8) versus [−7,−2) ms.

## Reliability with the same units held fixed

To avoid changing the cohort at every light level, the principal comparison follows the **same 31 SST units classified as responders at level 2.0**:

| Recorded level | Mean response rate, Hz | Mean baseline rate, Hz | Pulses with a response spike | Pulses with a baseline spike | Zero-spike observations | Units with any zero-spike trial |
|---|---:|---:|---:|---:|---:|---:|
| 1.3 | 36.6 | 10.8 | 21.5% | 6.5% | 608/775 | 31/31 |
| 1.7 | 157.0 | 10.1 | 68.9% | 5.8% | 241/775 | 26/31 |
| 2.0 | 229.0 | 11.0 | 87.0% | 6.3% | 101/775 | 22/31 |

Probabilities are means of per-unit fractions; trial counts are equal within each level. The 775 observations are **31 units × 25 shared pulses**, not 775 independent experimental replicates. Median per-unit probabilities are 12%, 76% and 92%, respectively. At level 2.0, 9/31 units spiked on all 25 observed pulses.

The paired excess probability above baseline is 15.1, 63.1 and 80.6 percentage points. Nominal 95% pulse-bootstrap intervals are 12.6–17.5, 60.0–66.1 and 78.1–83.1 percentage points, respectively. Bootstrap sampling resamples complete pulses across all units together and keeps baseline and response paired, preserving within-pulse correlation. It assumes exchangeable pulses and does not correct serial dependence, selection bias or multiplicity. These are within-recording descriptive uncertainty intervals, not animal-population confidence intervals.

For individual SST units, pointwise Wilson 95% intervals quantify finite-trial precision under independent, stationary Bernoulli trials. Even **25/25 has an interval of approximately 86.7–100%**. The selected 23/25 example has an interval of 75.0–97.8%. Neither proves future perfect reliability. Responder selection on these same data can inflate apparent performance and invalidate nominal post-selection coverage.

The lower-level cohort results are conditional on selection at the highest level. For comparison, the earlier level-specific responder subsets contain 3, 24 and 31 units, with mean probabilities of 84.0%, 83.3% and 87.0%. Those changing subsets obscure the much larger level dependence visible when units are held fixed.

## Why an average can look stronger than trial reliability

A firing rate weights spike counts, including bursts; response probability counts a trial only once:

**Mean rate = probability of any spike × mean spikes on successful trials ÷ window duration.**

At level 2.0, the 31 responders produced 1,065 spikes in 674 successful unit–pulse observations, or 1.58 spikes per successful observation. Thus a large average can reflect both more successful trials and multiple spikes on successful trials.

Unit **951083397**, the median-reliability illustration, averaged **266.7 Hz** at level 2.0 yet responded on **23/25 trials**, leaving two zero-spike trials. At levels 1.3 and 1.7 it responded on 6/25 and 19/25 trials. Unit 951083778, the least reliable primary responder, averaged 80 Hz against zero baseline but responded on only 10/25 trials.

The raster illustrations preserve chronological pulse order and all rows, including failures. PSTHs are unsmoothed 1 ms bins. Spikes outside the counting windows remain visible as context; a cross at the right of a raster marks a trial with no [2,8) ms spike, not a spike at that plotted location.

## Timing and trial order

At SST level 2.0, the median across responders of each unit’s conditional first-spike median was **4.08 ms**. Unit medians ranged from 2.26 to 7.19 ms. The median within-unit first-spike interquartile range was **0.91 ms**, with unit IQRs ranging from 0.33 to 2.80 ms; median within-unit SD was 0.94 ms.

These distributions include successful trials only. Failures are retained separately and never imputed as zero latency or as spikes at the window endpoint. They are first *observed in-window* spike times, not necessarily the first biological spike after light onset. The 2 ms lower cutoff and 8 ms upper cutoff truncate timing; early or later responses can fall outside this definition.

For illustration, unit 951083397’s conditional median shifted from 6.41 to 4.57 to 3.65 ms across levels, with IQRs of 2.22, 1.28 and 0.78 ms. Cohort jitter summaries use only units with at least five successful trials: 13, 28 and 31 units across levels. Their median IQRs were 1.92, 1.12 and 0.91 ms. This is a changing successful subset, so it is not evidence of a uniform within-unit improvement.

Chronological success matrices reveal individual failures and clusters without sorting trials by outcome. For the fixed cohort, mean probabilities in the first 12 versus last 13 pulses were 22.0% versus 21.1%, 68.3% versus 69.5%, and 85.8% versus 88.1%. These descriptive splits do not establish stationarity or rule out unit-specific adaptation, drift or correlated failures.

## All eligible units and the wild-type recording

The selected responders are a small fraction of recorded units. Mean probabilities across **all eligible units** were:

| Recording | Level | Post-pulse spike probability | Matched baseline probability |
|---|---:|---:|---:|
| SST | 1.3 | 4.29% | 4.40% |
| SST | 1.7 | 5.54% | 4.04% |
| SST | 2.0 | 6.03% | 4.00% |
| Wild-type | 1.0 | 5.84% | 5.68% |
| Wild-type | 2.5 | 5.52% | 4.84% |
| Wild-type | 4.0 | 5.69% | 5.46% |

Wild-type remains descriptive under the earlier protocol because there are only 15 trials per level. No new wild-type responder labels, p/q values or confidence intervals are assigned. Small aggregate differences do not establish absence of a response or absence of artifacts.

The wild-type raster deliberately shows the unit with the largest high-level rate difference, **950944980**, rather than presenting a selected extreme as typical. At level 4.0 it averaged 77.8 Hz post-pulse versus 22.2 Hz baseline but had spikes on only 7/15 post-pulse trials and 2/15 baseline trials. Its other levels had 4/15 and 3/15 post-pulse successes. This selection is illustrative, not inferential evidence of optotagging.

## Sensitivity and limits

For the fixed 31-unit SST cohort, narrowing the response window to [3,8) ms and matching its baseline changes excess probabilities from 15.1, 63.1 and 80.6 percentage points to **14.7, 62.1 and 77.9**. The central observation is therefore not confined to the earliest included millisecond. The supplied reference’s stricter-window labels retained 30/31 primary responders; those labels are reported as prior results, not newly retested classifications.

Extraction completeness is verified relative to supplied spike arrays, but **biological observation completeness is assumed**. Probe timestamp endpoints enclose the windows; internal gaps, unit-specific observation intervals and uninterrupted spike detection are unavailable. A zero-spike row is retained as observed silence under that assumption, not dropped as missing. It could still reflect an unreported detection gap.

No raw voltage, peri-light waveforms or LFP signals are supplied. Highly repeatable light-locked events can include photoelectric contamination or sorting artifacts. Excluding the onset/offset guard regions and testing [3,8) ms reduces a specific sensitivity but cannot prove all remaining events are genuine spikes.

One SST-ChR2 mouse and one wild-type mouse differ in sex, age, session and light settings. Units are not independent animals. These data support **heterogeneous, level-dependent light-associated spiking in these recordings**. They do not establish universal pulse reliability, a causal probability for each spike, direct activation, SST identity for every responding unit, artifact-free responses, or a general genotype effect.

## Figure captions and reproducibility

**reliability-overview.png.** Each point is an eligible unit at the indicated level: 995 SST units and 702 wild-type units, with 25 and 15 pulses per unit respectively. Purple marks the prior SST responder subset for that level. The diagonal denotes equal baseline and post-pulse probabilities. Coincident points overlap; the complete distributions are retained in the CSV.

**sst-reliability-and-timing.png.** The same 31 primary responders, ranked once by level-2.0 reliability and unit ID. Top: response fractions with nominal Wilson intervals and paired baseline fractions. Middle: first-spike median and IQR only where at least five successful trials exist; missing marks indicate insufficient successes, not zero latency. Bottom: chronological binary outcomes for all 25 pulses at each level.

**sst-averages-and-trials.png.** Three primary responders closest to the 25th, 50th and 75th percentiles of high-level reliability, without duplicates and with ties resolved by unit ID: 951083408, 951083397 and 951083573. Each unit is shown at every level, with its PSTH above all 25 raster rows. The same rate scale is used across levels within each unit. Pale gold denotes the 0–10 ms light pulse and pale blue the 2–8 ms counting window.

**wildtype-averages-and-trials.png.** The explicitly selected high-level rate-difference extreme, unit 950944980, shown at all three recorded levels with 15 raster rows each. The colors and shading follow the SST raster figure. Descriptive only.

**optotagging-reliability-results.zip** contains all four figures, 106,215 trial rows, 5,091 unit/level rows, level summaries, unit quality audits, included pulse metadata, fixed cohort ordering, validation records and a standalone Python reproduction script. The numerical and plotting code was executed in the notebook; the packaged command-line script was syntax-checked. Counts in illustrated trials were independently checked with direct timestamp masks. All four final images passed text-overlap/canvas checks and visual review.

Input ZIP SHA-256: `5171ab9d426b283a81f1176502b495c74d10e71d81fe3b1350894982cd882d69`.
Nested source ZIP SHA-256: `c64e8857f0062acadecea190209bfbe96de16c2897862c4654f34a93e43cc154`.
Whole remote NWB checksums and supplied remote extraction-check claims were not independently reverified.
