# Optotagging cutoff sensitivity

The original result reproduces exactly: **31 of 995 quality-eligible SST units (3.12%)** are selected at the highest recorded light level, 2.0. Selection is moderately sensitive to the statistical cutoff alone and more sensitive to response window and correction method. Thirteen units remain selected throughout the main corrected sensitivity grid. These are light-associated responses, not proof of direct activation.

## Data and reference verification

The supplied archive contains Allen SST session 794812542 and wild-type session 767871931. Both checksum manifests passed. We verified unit/spike alignment, finite sorted spike times, complete supplied extraction windows, stimulus isolation, and enclosing probe acquisition bounds. Recomputed rates, p-values, adjusted p-values, and selected identities matched the supplied reference at all three levels, including its stricter window.

The unchanged quality filters are quality=good, amplitude_cutoff<0.1, presence_ratio>0.95, isi_violations<0.5, and a valid peak electrode, with finite numerical QC values. These retain 995/2688 SST units and 702/2685 wild-type units. All isolated 10 ms square pulses are retained: SST has 25 trials per level and wild-type 15.

Reference: response [2,8) ms versus baseline [-8,-2) ms, one-sided paired sign-flip test, BH adjusted p≤0.05, response rate≥2×baseline rate, and mean increase≥10 Hz. Fold requirements are implemented by multiplication, so a zero baseline is handled without an infinite ratio. The highest light level remains primary.

## One-factor comparisons at the primary SST level

Everything other than the named factor remains at the reference setting.

| Factor | Setting | Selected / 995 |
|---|---|---:|
| BH cutoff | 0.01 | 28 |
| BH cutoff | 0.05 | 31 |
| BH cutoff | 0.10 | 31 |
| Minimum increase | 5, 10, or 20 Hz | 31 at each |
| Minimum fold increase | 1.5, 2, or 3 | 31 at each |
| Response window | [2,8) ms | 31 |
| Response window | [3,8) ms | 30 |
| Response window | [2,6) ms | 24 |
| Response window | [4,8) ms | 27 |

Every combination of the tested effect requirements also selects 31 at the original window, BH correction, and 0.05 cutoff. These effect thresholds are therefore not binding for the reference-selected set. This does not mean effect size is generally irrelevant.

At cutoff 0.05 and reference effect requirements:

| Window (ms) | BH FDR | BY FDR | Holm FWER | Uncorrected |
|---|---:|---:|---:|---:|
| [2,8) | 31 | 28 | 27 | 41 |
| [3,8) | 30 | 27 | 26 | 41 |
| [2,6) | 24 | 24 | 22 | 36 |
| [4,8) | 27 | 23 | 21 | 36 |

The alternative BH windows lose 1, 7, and 4 reference-selected units, respectively, without gaining others at these fixed thresholds. Uncorrected counts are a diagnostic comparator, not the recommended headline result.

## Sensitivity grid and stable identities

We crossed four windows, three cutoffs (0.01/0.05/0.10), three absolute effects (5/10/20 Hz), and three fold requirements (1.5/2/3), with BH, BY, Holm, and uncorrected tests. Equal-duration baselines end at −2 ms: [-8,-2), [-7,-2), [-6,-2), and [-6,-2), respectively. All windows respect the reference onset and offset masks; the shorter windows probe early versus later responses within the original safe interval.

Corrections are calculated before effect-size filtering. Per window, primary SST tests form one family of 995 units; both lower levels jointly form the reference secondary family of 1990 tests. Lower levels remain separate in reported counts. There are 324 corrected settings and 108 uncorrected settings per level. A separate 324-setting analysis per level corrects across all four windows as well, giving primary and secondary families of 3980 and 7960 tests. Total evaluated settings: 2268.

For the **main per-window corrected grid**, primary counts range from **16 to 32** (1.61–3.22%). Across the grid, 33 distinct units are selected at least once:

- **13 always selected**, all in the original 31.
- **20 change classification**: 18 original selections plus two additional units.
- **962 never selected**.

The 13 always-selected Allen unit IDs are:
951083206, 951083231, 951083397, 951083876, 951087429, 951087843, 951087886, 951088523, 951094157, 951094682, 951101812, 951105085, 951105613.

The 18 changing reference selections are:
951083173, 951083326, 951083343, 951083401, 951083408, 951083573, 951083778, 951084733, 951086967, 951087437, 951087454, 951087529, 951090313, 951091711, 951091837, 951093833, 951104389, 951104539.

Additional units 951083349 and 951083371 each appear in only 9/324 corrected settings. The CSV records all unit identifiers, original labels, selection frequencies, and correction-scope-specific classifications. Frequencies measure stability over this chosen grid; they are not probabilities of direct activation or independent replications.

Lower-level reference counts are 3 at level 1.3 and 24 at level 1.7. Their main corrected-grid ranges are 0–4 and 0–27. Uncorrected primary counts range from 30–49.

## Numerical resolution and searching across windows

The original implementation enumerates the sign-flip distribution when at most 20 trials have nonzero paired differences. Otherwise it uses 99,999 randomizations with a +1 correction, giving a minimum p-value of 0.00001. We retained this implementation and its seed for the reference and main sensitivity tables.

That resolution becomes consequential under stringent corrections. Across-window Holm at cutoff 0.01 begins near 0.01/3980 = 0.00000251, below the Monte Carlo floor. Thus loss of some strongly responding units can be numerical, rather than evidence of weaker biological responses.

An exact dynamic-programming audit of primary SST sign-flip distributions leaves the main 31-neuron reference result and the 13-unit stable core unchanged. It changes the per-window corrected range from 16–32 to 17–32, with at most one identity changed per setting. Across-window correction is much more affected: the original implementation yields 2–32 units and no unit selected in every setting, whereas exact p-values yield 15–32 and an 11-unit intersection. Up to 19 identities differ for a single setting. Exact results and p-value comparisons are provided separately in the bundle and exact-status columns of the neuron table.

At the reference window, cutoff, and effects, across-window correction selects 31 with BH, 28 with BY, and 25 with Holm using the original implementation. Do not interpret the union over explored settings as a new confirmed discovery set. Window-wide correction addresses searching among those windows, not every possible analysis decision.

## Interpretation and limits

BH targets false-discovery rate under suitable dependence assumptions; BY allows arbitrary dependence and is more conservative; Holm controls family-wise error with valid marginal p-values. None establishes that a selected unit is directly activated. The sign-flip test also relies on exchangeability/symmetry under its null; this analysis preserves that reference assumption rather than validating it experimentally.

Coverage is assumed from probe acquisition bounds and extraction-completeness evidence. Bounds enclose all analyzed events, but internal gaps and uninterrupted unit-level observation are not independently verified. Silence is retained as zero counts, never used to infer missing acquisition. No full remote source re-download was performed.

Wild-type remains descriptive because it has only 15 trials per level, below the original minimum of 20. At its highest level in the original window, mean rates across 702 units are 9.31 Hz baseline and 9.73 Hz response, with median paired rate difference 0 Hz. The bundle includes all descriptive windows and levels; no wild-type responder labels or percentages are assigned. Light-level scales and sessions are not pooled or treated as matched experimental replicates.

There is one recording per group. These neuron counts are not animal-level replication, and the sensitivity grid is exploratory because prior results were available. Short-latency responses can include indirect network activity, and sorted spike times alone cannot decisively exclude light artifacts. The original 2 ms onset mask prevents assessing earlier responses. Stable statistical selection is evidence of analytical robustness only.

## Reproducibility and files

The bundle includes runnable reproduce.py, environment versions, input hashes and structural checks, QC audits, retained trials, reproduced reference results, window statistics, setting-level counts and overlaps, sparse selected-unit membership, wild-type descriptive summaries, and the exact-p numerical audit. Run:
`python reproduce.py path/to/optotagging-reference-study.zip`

The script was executed in a fresh namespace and reproduced both main CSV outputs exactly. Original reproduction, grid monotonicity, and figure legibility checks passed. Python 3.12.14, NumPy 2.5.3, pandas 2.3.3, SciPy 1.18.0; seed 20261009.

The count figure holds the cutoff and effects at their reference values. The stability figure displays all 33 primary units selected at least once in the 324 per-window corrected settings; it omits the 962 never selected. Neither figure displays exact-audit results.
