# Blinded Kepler transit injection challenge

## Result

One hundred distinct public Kepler Quarter 1 long-cadence light curves were downloaded and parsed. Exactly five received synthetic, periodic transit-like dips. The AI reviewed anonymized numerical summaries in ten groups of ten, with no required number of flags. All 100 raw AI answers were saved before the injection key was opened. A periodic box-search rule was also fixed and saved before the key comparison.

| Method | Detected injections | Missed injections | Flags on 95 unchanged curves |
|---|---:|---:|---:|
| AI reader | 4 | 1 | 52 |
| Periodic box search | 5 | 0 | 37 |

The AI missed C013, despite its injected 9.7‰ dip at a 5.33-day period. Its raw reason says the reported search depth and score appeared inconsistent with the smaller 32-bin folded dip. At the true injected period, the close-up shows a clear dip; the compressed candidate summary obscured that evidence.

## Reproducibility and blinding

- Source: first 100 distinct products in the official [MAST Kepler Q01 long-cadence bulk script](https://archive.stsci.edu/missions/kepler/download_scripts/lightcurves/kepler_lightcurves_Q01_long.sh). `kepler_data_manifest.csv` gives each archive URL, KIC identifier, SHA-256 digest, byte count, and usable-sample count. MAST documents the [public Kepler light-curve products](https://archive.stsci.edu/kepler/download_options.html).
- Preprocessing: use finite TIME and PDCSAP_FLUX with SAP_QUALITY = 0, normalize each curve by its median, then divide by a 1.5-day median-filter trend and subtract one.
- Injection selection: NumPy `default_rng(20261008).choice(100, size=5, replace=False)` on manifest order. A multiplicative periodic trapezoid was applied to the normalized flux, with 0.025-day ingress and egress. The exact five periods, depths, durations, and epochs are in `kepler_injection_key.json`.
- Reader input: each prompt contains ten anonymized curve IDs, quality and search statistics, and 32 phase-folded median residuals. It includes no injection label and states that there is no required number of flags. `kepler_ai_inputs.json` has the exact prompts. `kepler_ai_answers.json` has all ten raw successful responses and parsed decisions. One initial Group 8 delegation failed internally; its identical-prompt retry succeeded and is identified in that log.
- AI capability: delegated reviewers inherited the Main Agent profile. The session model ID observed afterward was `gpt-6-sol`; runtime and model settings were left unchanged.
- Baseline: the saved `kepler_baseline_rule.json` fixes periods 1.5–9.5 days (321 grid values), durations 0.12/0.20/0.28 days, a robust phase-bin periodic box score, and a flag threshold of score ≥ 8 with at least three coherent events. `kepler_baseline_scores.csv` was saved before opening the key.
- Evaluation: `kepler_results.csv` contains each curve's injection status, both flags, AI reason, and search score. `kepler_metrics.json` contains the counts. The notebook execution history and these ordinary artifacts are compatible with normal session-package export.

## Figures

- `kepler_overview_100.png`: all 100 processed challenge curves. Orange identifies the five injected curves after key reveal; each small panel uses its own vertical scale.
- `kepler_injection_closeups.png`: original and injected flux folded at each known synthetic period, with injected duration shaded.
- `kepler_results_chart.png`: detected injections and unchanged-curve flags, with their separate denominators.

## Interpretation

This is an injection recovery exercise on a convenience sample from one Kepler quarter, not a planet discovery study. The five synthetic dips are not new planets. An unchanged curve may already contain an astronomical signal, an instrumental artifact, or stellar variability; a flag on one of these 95 curves is therefore **not** proof of a false astronomical discovery. Both methods were exposed to the same processed data, and the AI summaries were built from the periodic search's preferred candidate, so their errors are coupled. The score threshold is heuristic rather than a calibrated false-alarm probability. The small number of injections and their relatively deep 4.6–9.7‰ dips limit generalization to shallower transits.