EEG Seizure Detection
Seizure detection from raw 23-channel scalp EEG in the CHB-MIT database, for ECE 539 (Neural Networks) at UW–Madison. An end-to-end pipeline runs from EDF ingestion through annotation-based labels and feeds five models identical 4-second windows: three classical baselines, a 1D CNN on the raw signal, and a 2D CNN on STFT spectrograms.
Every model scored over 99.6% accuracy while catching zero seizures, because seizures make up under 5% of a recording. Sensitivity exposed the problem, and AUC suggested the CNNs had learned to rank seizure windows above background. At the time, the conclusion was that the failure was the decision threshold, not the model.
Revisiting the project later, I found two deeper problems. A labeling bug was under-counting seizure windows by about 10×. And the random window split let overlapping windows share signal across train and test, which inflated AUC to 0.98–0.999. Together they mean the headline numbers, the 99.6% included, measured leakage as much as they measured the models.
The corrected protocol uses recording-level splits, so no recording contributes to both sides, and thresholds chosen on validation data against a false-alarm budget rather than the default 0.5. That's the operating point a clinician would actually care about: how many seizures are caught at how many false alarms per hour. The lasting takeaway is the evaluation, not the models: under this kind of imbalance, accuracy says nothing, and the split has to respect where the data came from before any other number can be trusted.
My part: the 2D CNN and its spectrogram pipeline, the evaluation protocol, and the two bug fixes.
The source for this one isn't public yet. Happy to walk through it or share access — email me.
