Genlight × Strong AI Lab — Proof of Concept · Phase 1

Can background knowledge cut how much of a new user's own data NeuroTune needs?

A first, minimal walk-through of the two anchor points in the collaboration proposal: a data-efficient signal-analysis model, and an answer a user can actually check instead of a bare probability.

Stand-in data — NeuroTune's own device data is ~3 months out (per the 27 Jul meeting), so this runs on a public sleep-EEG dataset (PhysioNet Sleep-EDF), 2 EEG channels resampled to mimic NeuroTune's setup.

1 — Data efficiency

Same tiny amount of a new person's own signal, two ways to train on it — tried across eight model families, including four public EEG/sleep foundation models (LaBraM, BIOT, CBraMod, SleepFM).

Pretrained, then fine-tuned on this person's data Trained from scratch on this person's data only Pretrained, zero personal data (reference)

Each point uses a random sample of that many minutes from the person's own recording (not necessarily their first minutes of wear), evaluated on a held-out portion of the same night never seen in training. Averaged over 2 held-out people.

Model familyParametersZero-shotFine-tuned (least data)Fine-tuned (most data)Scratch (most data)

2 — An auditable answer, not just a score

Real epochs from the held-out test set — including a borderline case and an honest miss, not just wins.

User
Example