How Cognogram works,
and why it's built this way.

Seven quick tests, backed by decades of research, that take a snapshot of how your mind is working — and then let you watch it over time. The point isn't "how do you compare to everyone else." It's "how are you doing, compared to you?"

7 tests~8 minutes
Just you vs. youyour own baseline
Plain-spokenhonest about limits
Freebecause access matters
START HERE

Your brain, in plain terms

You don't need a neuroscience degree to use Cognogram or to read your results. Here's the whole idea in a few plain sentences.

Knowing is not the same as doing

Picture your brain in two rough halves. The back is where you learn and know things. The front — right behind your forehead — is where you do things: plan, start, focus, resist the distraction, stick with a goal even when it's boring. Psychologists call that front-of-brain "doing" system your executive function.

Here's the part people find surprising: knowing and doing can come apart. You can know exactly what you should do and still find it genuinely hard to make yourself do it. That's not a character flaw — it's the "doing" machinery working differently. It's why ADHD is better understood as a performance problem, not a knowledge problem: often it's not that someone doesn't know what to do, it's that the bridge from knowing to doing is under strain.

Executive function, in one line

It's the use of self-directed action — steering your own behavior — to choose goals and stick with them across time. Putting the brakes on an impulse, holding a plan in mind, seeing the future clearly enough to act for it now.

What these eight tests are actually looking at

Each test gently probes a different part of that "doing" machinery. In plain words:

  • Reaction Time — how fast your eyes-to-thumb wiring fires. The floor everything else is built on.
  • Symbol Sprint — how quickly you can scan, match, and keep going. Everyday mental speed.
  • Color Clash — your mental brakes: holding back the automatic answer to give the right one.
  • Go / No-Go — staying on task through something boring, without jumping the gun. The classic attention/ADHD measure.
  • Digit Vault — your mental sticky-note: holding information in mind, then working with it.
  • Connect — switching gears on the fly: jumping between two rules without losing your place.
  • Arrows — filtering out distraction to act on what matters, under interference.

Inside the app you'll literally see a little brain light up in the areas each test leans on as you build your session — a reminder that these aren't abstract scores, they're windows onto real machinery.

Why bother tracking it?

Because your brain changes — with a rough night's sleep, a new medication, a concussion, a stressful season, or just the slow drift of the years. Almost no one has a record of how their mind worked when they were well. Set a baseline once, and every check-in afterward tells you something a single test never could: which way, and how fast, am I moving? Think of it like a lap timer for your own head.

One honest caveat up front

Cognogram is a tool to inform, not to diagnose. It can flag "something's changed, worth a look" — it can't tell you why, and it never replaces a real clinician. Everything below explains exactly how far the numbers can and can't be trusted.

01

Purpose & positioning

In plain terms

Most brain tests ask "how do you stack up against everyone else, right now?" Cognogram asks a friendlier and often more useful question: "how are you doing compared to your own normal?"

Most computerized neurocognitive tools answer a cross-sectional question: how does this person compare to a normative sample right now? That's valuable — but the instruments that do it well tend to be expensive, clinic-bound, and visually dated. Cognogram is built around a complementary question: how is this person changing, relative to themselves, over time and across conditions?

This within-subject framing has three practical consequences that shape every design decision in the product:

  • It sidesteps the hardest normative problem. Population norms are demographically fragile and slow to build. A within-subject design uses each person as their own control — statistically powerful for detecting change, and honest about what an un-normed tool can claim.
  • It matches how the target questions are actually posed. "Did this ADHD medication help?" "Am I duller on four hours of sleep?" "Has my baseline slipped since last year?" are all change questions, answered by repeated measurement against a personal reference point — not by a single percentile.
  • It lowers the access barrier. Because the tool compares you to you, it can be free and self-administered on a phone, reaching people who'd never sit for a paid, proctored battery — while still generating data structured enough to support formal validation later.
A note on honesty

Cognogram's 0–100 domain scores are internal indices for personal trend tracking. The population percentile you'll see in your report is built from other Cognogram users and grows more trustworthy as more people test — always shown with the current sample size. How we handle that fairly is covered in measurement & fairness.

02

The eight tests

In plain terms

Seven short, well-worn tests — each one gently probes a different part of the "doing" brain: raw speed, everyday mental pace, your brakes, staying-on-task, and your mental workspace. Tap any card to go deeper.

Each is a computerized adaptation of a paradigm with decades of published use. Tap any card for the paradigm, what it indexes, the evidence, and how Cognogram implements it.

03

Why these eight — and not others

The selection is deliberate. Every inclusion earns its place against four constraints: run in ~7 minutes on a phone, cover the domains that matter, favor paradigms with strong pedigrees, and — critically — behave well under repeated testing.

Domain coverage with minimal redundancy

TestPrimary domain covered
Reaction TimeRaw psychomotor speed — the anchor
Symbol SprintProcessing speed under scanning + attention load
Go / No-GoSustained attention + impulse control + RT variability
Color ClashExecutive inhibition / cognitive flexibility
Digit VaultWorking memory — storage and manipulation

Reaction Time and Symbol Sprint are both "speed" measures, but deliberately so: the first isolates the motor/perceptual floor, the second layers scanning and attention on top. Reading them together lets a change be localized — a slowdown on both suggests a basic-speed shift; a slowdown on only the SDMT-type task points to attention or scanning.

Selected for repeatability, not just validity

A test can be perfectly valid on first administration yet useless for tracking if it improves every time from practice or drifts from boredom. Because Cognogram is fundamentally a repeated-measures tool, repeatability was a gating criterion. Reaction-time-based measures of inhibition, switching, and selective attention generally show good test–retest reliability with only gradual practice gains once a brief familiarization is included — whereas some cost scores and visuospatial working-memory metrics show poor retest reliability and were avoided as primary outputs.[13][14] The CPT specifically reaches a practice plateau after the second administration, with good-to-excellent reliability on its summary indices.[15]

What we deliberately left out (for now)

  • Trail Making B & complex set-shifting — high executive value, but touch-drag interactions are fiddly and error-prone on small phones; a future tablet build is a better home.
  • N-back working memory — powerful and sleep-sensitive, but its scoring metrics show inconsistent test–retest reliability in some computerized batteries,[14] risky as a headline tracking measure without careful local validation.
  • Long-form vigilance (10-min PVT) — the gold standard for sleep-loss research, but ten monotonous minutes is hostile to voluntary, repeated self-testing. Go/No-Go captures much of the same attention-lapse signal in a fraction of the time; a full PVT module is a candidate for research-only deployments.
  • Verbal memory / list learning — clinically rich for early Alzheimer's detection, but hard to deliver fairly without audio standardization and alternate forms; slated for a later, controlled release.

The guiding principle: ship a short, repeatable, phone-native core that's defensible today, and add higher-friction tests only where a specific research deployment justifies the cost.

04

How scores are computed

Every test yields raw, interpretable metrics — milliseconds, error counts, spans — that are always shown.

In addition, Cognogram maps performance onto five 0–100 domain indices and an overall composite, purely to make trends legible at a glance. These indices are linear rescalings of raw metrics onto a fixed range; they are not norm-referenced. The design intent is that you watch your own line move, with your baseline drawn as a reference tick on every bar. Where a test informs more than one domain — the CPT feeds both attention and impulse control — its metrics are split accordingly. Because the composite is an internal construct, its absolute value matters far less than its change, which is why the interface always shows the delta versus baseline.

05

Real-world applicability

A cognitive baseline is not interesting in itself. It becomes interesting the moment there is a question you cannot otherwise answer — and most of those questions share a shape: has something changed, and by how much? Six situations where that question comes up often enough to be worth instrumenting.

Deciding whether an ADHD evaluation is worth pursuing

This is the most common reason people arrive. Waitlists run months, private assessments are expensive, and the decision to join the queue is usually made on the strength of an internet quiz and a hunch. Cognogram cannot diagnose ADHD and does not try. What it can do is supply something a quiz structurally cannot: objective, repeatable, dated measurements to weigh alongside the self-report — and, just as usefully, tell you when your performance looks unremarkable across several sessions in varied states. That is real information in the other direction, and it is information almost nobody currently has when they make the decision.

The full argument sits in the ADHD lens and the science of ADHD testing; Convergence, in the experiment suite, is the flow that operationalizes it. What matters here is the framing. The output is a probability with an honest error bar and a record you can hand to a clinician — not an answer, and not a substitute for the developmental history, cross-setting impairment and differential diagnosis that a real evaluation requires.

Is my medication actually working?

Subjective reports of “it’s helping” are notoriously hard to calibrate, and stimulants are a particularly hard case: they reliably make people feel sharper, and that feeling is only loosely coupled to whether performance moved. CPT indices — commission errors and reaction-time variability in particular — have been used as outcome measures sensitive to stimulant medication versus placebo in ADHD.[11] Cognogram’s intervention workflow captures exactly what you’re taking and when, timestamps it, and compares the session against your own clean baseline, ideally matched for time of day.

Two honest limits travel with every such comparison. Medication frequently improves cognition without fully normalizing it,[24] so a still-elevated score does not mean the treatment is failing. And a null result does not mean it isn’t working — a browser test is coarse and easily swamped by sleep, timing and effort. What the drugs themselves actually do to the brain is covered in the ADHD pharmacopoeia.

Caffeine — the drug you are already taking

Nearly everyone who will ever use this tool is on a psychoactive drug at the moment they use it, and most of them will not think to mention it. Caffeine acts on the precise system that sustained-attention tasks measure, its effect size on those tasks is not small, and it is ordinarily invisible in cognitive data. Rather than treat that as noise to be apologised for, Cognogram treats it as the most convenient natural experiment available: a cheap, safe, well-characterized compound whose rise and fall in your bloodstream can be modelled and checked against your own performance. The caffeine section covers the pharmacology; Ignition is the paradigm built on it.

Quantifying the cost of lost sleep

Sleep loss produces one of the most reliable signatures in all of cognitive science: total sleep deprivation increases reaction time and the frequency of attention lapses,[16] and even chronic partial restriction to six hours a night accumulates deficits comparable to a night of total deprivation.[17] Sustained-attention measures are most sensitive to this, which is why the psychomotor vigilance task is the field’s reference tool.[18] A sobering, useful fact for a self-tracking app: people’s subjective sense of impairment dissociates from their measured performance — they underestimate how impaired they are.[19] An objective baseline is precisely what makes that invisible deficit visible.

This is also the most important confound in every other use on this page. Elevated reaction-time variability is produced by sleep debt just as reliably as by ADHD, which is why the honest move is to measure sleep alongside everything else and treat it as a competing explanation rather than background noise.

Coming back from something

Concussion, a bad viral illness, general anaesthesia, chemotherapy, a long stretch of untreated pain or depression — the question afterwards is always the same, and it is always unanswerable: am I back to normal? Without a “normal” recorded from when you were well, the only available comparison is memory, which is exactly the faculty least suited to the job. A dated baseline turns an unanswerable question into a tractable one.

To be explicit about the boundary: Cognogram is not a concussion assessment tool and must not be used for return-to-play or return-to-work decisions. Validated instruments exist for that purpose and are administered for good reasons. What a personal baseline offers is a slower, less consequential form of the same information — a trend line you can bring to the person actually making those calls.

Testing any intervention against yourself

The machinery is general. Anything you can start and stop — a supplement, an exercise routine, a meditation practice, a light box, a change of schedule, a diet — can be run as an A/B against your own baseline, with the same timestamped, within-subject logic. This matters most in the domain where the evidence is thinnest. The nootropics market is enormous and almost entirely unencumbered by controlled trials; the majority of what is sold has never been tested in humans for the thing it is sold for. In that environment, the single most valuable function of an objective measure is not to find what works. It is to make it harder to fool yourself — and to notice that the reference compound most of the shelf is quietly competing against costs about ten cents and is already in your kitchen.

A midlife baseline, tracked over time

The strongest case for Cognogram is preventive. Processing speed declines gradually with age, and a substantial drop can be an early flag for injury or disease rather than normal aging.[2] Yet almost no one has a personal cognitive baseline from when they were well. Establishing one in midlife — then repeating it yearly — converts an unanswerable question (“is this normal for me?”) into a tractable one (“has this moved, and how fast?”). The value isn’t any single score; it’s the slope.

06

Caveats, limitations & honest disclaimers

Cognogram is not a diagnostic instrument.

No score it produces diagnoses ADHD, mild cognitive impairment, concussion, or any other condition. It generates objective, repeatable data that can inform clinical judgment and support research. Diagnosis remains a clinical act.

Practice effects are real and must be managed

Scores on computerized batteries commonly rise on early retests from learning, not true cognitive gain — documented directly for the class of tools Cognogram resembles, with the largest gains on flexibility, processing speed, and reaction time at first retest.[20] The mitigations are well established: a brief familiarization block before each test so the steepest learning happens first,[14] and treating the first full session as a flagged baseline. Cognogram does both — every test opens with a skippable practice block, and your baseline is tagged as a first session so it can be handled carefully in analysis.

Measurement error sets a floor on meaningful change

Even reliable tests can have wide minimal-detectable-change bands — the CPT shows good-to-excellent reliability yet non-trivial random measurement error across serial assessments.[15] In plain terms: small wiggles are noise. Treat only sizeable, sustained shifts as signal; future scoring should express change against a real reliable-change threshold rather than raw point differences.

Uncontrolled conditions

Self-administration on a personal phone means variability in screen size, touch latency, lighting, distraction, effort, caffeine, and time of day. The within-subject design absorbs some of this — but only if you hold conditions roughly constant, which the app prompts you to do. Cross-person comparison of absolute scores is not supported and not advised.

Does your device's speed change your score?

In plain terms

A slow phone or laptop adds a tiny, hidden delay to every reaction-time reading. For tracking yourself on the same device, that delay is baked into every session and quietly cancels out. For comparing you to other people, it matters more — so we lean on measures that don't care much about it, and we quietly note what device you used.

This is a real and well-studied issue in web-based testing, and it's worth being precise about. Browser-based reaction-time measurement adds a lag beyond native lab software of roughly 25–45 ms depending on the system, plus small trial-to-trial jitter on the order of 5–10 ms.[21] Older comparisons put JavaScript's added latency near 88 ms with a standard deviation around 6 ms, versus roughly 50 ms for native software.[22] Timing precision also varies by operating system and browser — with macOS among the less precise for visual stimuli.[21]

The key insight is what kind of error this is. Device latency is dominated by a near-constant offset, not random noise. In a within-subject design on a stable device, a constant offset present in every session simply subtracts out when you compare session to session — so Cognogram's core use case is remarkably robust to it. This is, quietly, one of the strongest arguments for the whole design. The residual trial-to-trial jitter is small next to genuine human reaction-time variability (often 50–150 ms), and it averages out across the many trials each test runs. For cross-person comparison — the population percentile — the offset is a genuine confound, which is why Cognogram (a) leans on latency-robust measures like reaction-time variability and error rates for anything comparative, (b) logs device and browser context with each session, and (c) nudges you to test on the same device each time.

How big is this next to the things we are trying to detect? Uncomfortably big, and it is worth stating in the same units. Across the range of hardware people actually own, the fixed offset spans roughly 40 ms. A night without sleep moves mean reaction time by about 30–40 ms. A dose of caffeine moves it by perhaps 15–20 ms. The gap between a gaming PC and an older phone is larger than either effect. Anyone comparing raw reaction times between two people is, to a first approximation, comparing their equipment.

But that headline conceals something more useful, because the offset does not hit every metric equally. Since it is close to constant, it lands almost entirely in the mean and barely touches the spread: frame-quantisation jitter adds to the variance in quadrature, and next to a genuine reaction-time SD of 50–150 ms it all but disappears. The consequences are sharply different per metric, and they run in a direction most people would guess wrong.

Three results worth reading off that table. Mean reaction time is fully exposed — about ±17% across common hardware, which is why we never rank people on it. Standard deviation is effectively immune, distorted by well under one percent, which is a large part of why Cognogram leads with variability rather than speed. And the counterintuitive one: the coefficient of variation is distorted by roughly as much as the mean is — 9–25% depending on how fast and how variable you are — and in the direction that flatters worse hardware. Because CV divides by the mean and slow hardware inflates the mean, a cheaper phone makes you look more consistent than you are. Any leaderboard built naively on CV would quietly reward owning a bad device.

That is a fixable problem rather than a fatal one, and the fix is a feature we already ship. Cognogram’s timing check estimates your display and input lag before a session; subtracting that estimate from the denominator before computing CV collapses the cross-device spread from roughly 17% to a fraction of a percent in simulation, for anyone whose responses are not extremely fast and extremely tight. This is the reason that check exists and is not decoration — and it is why the app records which device you used with every session.

The design rule this produces

Rank metrics by how much of the device they carry, and prefer the robust end wherever a choice exists. Difference scores are perfect — two conditions measured on the same device in the same session cancel the offset exactly, which is why the state-regulation contrast and every on-versus-off comparison are the strongest things here. Second-scale tasks are immune: 40 ms is half a percent of an eight-second interval, so timing tasks are the only ones where a phone and a desktop are genuinely equal. Dispersion is nearly immune. Corrected CV is usable. Raw mean reaction time, between people, is not. Anything competitive or shareable we build has to sit at the top of that list, not the bottom.

What are we comparing you to? (the normative question)

In plain terms

First and foremost, we compare you to you — your own baseline. For the "where you fall on the curve" number, we start from published research norms (so it means something on day one) and gradually blend in real Cognogram data as more people test. The report always names which of the two it's leaning on, and asks you to salt it accordingly.

A percentile is only as good as the reference group behind it. With just a handful of early users, a pool-only percentile is meaningless — everyone looks average. So Cognogram uses a blended, staged approach. Each domain is anchored to published means and standard deviations from the literature — simple reaction time (~250 ms in the lab, ~284 ms web-measured; slows ~1–2 ms/yr), symbol–digit processing speed (SDMT-type), and forward+backward digit span (WAIS-IV) — age-adjusted where the data support it. Your composite percentile begins as that published estimate and is weighted toward Cognogram's own accumulating data as the sample grows; the report labels the current regime as published, blended, or Cognogram. The within-subject comparison, meanwhile, is trustworthy from day one because it never depends on anyone but you.

Take the population number with a grain of salt

Published norms were collected under conditions a phone in your hand does not reproduce. The report surfaces these confounds directly, and so do we: input device (a touchscreen adds ~25–45 ms of latency — nearly all of it a fixed offset — versus the keyboard or pencil the norms used); screen & browser timing; unproctored self-administration versus a supervised lab; adaptation (several Cognogram tasks are fast reinterpretations, not the identical instrument); and who the reference sample is (specific ages, education, and cultures — unless you set your birth year, we assume ~30). Read the percentile as a rough bearing, not a clinical result.

Two more pieces of the report follow the same philosophy. The trajectory view plots your composite across every session so you can see which way — and how fast — you're moving, which is the comparison that actually cancels all of the confounds above. And the whole report is exportable as a PDF or a shareable link, so a clinician gets a clean, dated record rather than a screenshot.

Curious how steady your device is? The app runs a quiet timing check before every lap — and you can run the full thing yourself for fun: open the tech check → It measures frame steadiness, dropped frames, and your reaction floor, then tells you whether your rig is good to go.

ADHD LENS

The ADHD-informed lens

In plain terms

Cognogram is not an ADHD test and never says "you have ADHD." It can surface the cognitive signatures that research associates with ADHD — mainly inconsistent reaction time — and let you and a clinician make sense of them.

Attention research has converged on a clear hierarchy of what computerized tasks can and can't see in ADHD. Cognogram encodes that hierarchy honestly rather than overselling it.

What actually carries the signal

Across a meta-analysis of 319 studies, the measure that most reliably separates ADHD groups from typical groups is not a bad average — it's reaction-time variability (how much your speed wobbles trial to trial). Mean reaction time comes next, then commission (impulsive) errors. Cognogram weights its "attention signatures" in exactly that order, and it also looks at a vigilance decrement — whether you slow down or slip in the second half of the Go/No-Go, the hallmark "attention fades on a boring task" pattern that the Conners CPT captures as HRT Block Change.[23]

Signatures, not a diagnosis

Each signature is reported as a plain-language observation ("your reaction time varied a lot this session") with its evidence, and classified only relative to a typical range. A radar plot puts your five-axis attention profile front and centre, drawn over the typical band; an opt-in overlay adds the average shape for people with an ADHD diagnosis, so you can see at a glance where you sit between the two. It is wrapped in an explicit warning that landing near the ADHD shape does not mean you have ADHD. Base rates matter: many people without ADHD score there, and many with ADHD score in the typical range. Only a clinician can diagnose, using history, interviews, and validated scales such as the ASRS.

The effort check

Like the Conners CPT and other clinical CPT systems, Cognogram runs a validity check: a run with implausibly fast responses, near-random error rates, or many false starts is flagged as "may not reflect your best effort," and Cognogram proactively offers to exclude it from your tracking. Effort and conditions move these numbers as much as attention does.

Self-report alongside performance

Cognitive tests measure how attention performs; they can't see how it feels across a life. So Cognogram pairs the lens with the ASRS v1.1 — the World Health Organization's public-domain Adult ADHD Self-Report Scale — the same pattern every clinical battery uses when it combines a rating scale with testing. The six-item Part A is the validated screener; a positive screen is a prompt to talk to a clinician, never a diagnosis, and Cognogram shows whether your self-report and your performance profile point the same way.

Where it's genuinely useful: tracking a change

The strongest and safest ADHD use is within-subject: run a lap off-intervention, another on it, and compare. Because reaction-time variability responds to stimulant medication, a drop from an off-meds lap to an on-meds lap is a meaningful, personal signal — the same logic clinical retest systems use when they flag a treatment response. Cognogram's compare view lays the two side by side. It's a conversation-starter for a clinician tuning a medication, never a treatment decision on its own — and worth remembering that medication often improves cognition without fully normalizing it.[24]

ADHD SCIENCE

The science of ADHD testing

The short version

ADHD is diagnosed clinically — by history, criteria, and impairment — not by any test. Cognitive tasks add a modest, objective signal that is most useful for tracking change within one person (especially on vs. off medication), and weakest as a stand-alone diagnostic. Cognogram's Convergence stack is built to reflect exactly that evidence, and to state its own limits out loud.

Because Cognogram now includes an ADHD screening-support flow, the science underneath it deserves a full accounting — what the diagnosis actually rests on, what computerized testing can and cannot contribute, and why we combine several weak signals rather than trusting any single one.

What ADHD is — and how it's actually diagnosed

Attention-Deficit/Hyperactivity Disorder is a neurodevelopmental condition defined in the DSM-5-TR by a persistent pattern of inattention and/or hyperactivity-impulsivity that interferes with functioning or development. The formal criteria require several specific things at once: a threshold count of symptoms (in adults, ≥5 of 9 in a domain rather than the ≥6 required in children), several symptoms present before age 12, symptoms evident in two or more settings (e.g. work and home), clear evidence of functional impairment, and — critically — that the picture is not better explained by another condition.[27] That last criterion is where careful evaluation lives: sleep disorders, anxiety and mood disorders, trauma, substance use, thyroid and iron abnormalities, and learning disorders can all mimic or amplify attention problems. No computer task adjudicates any of this. A diagnosis is a clinical synthesis — typically a structured or semi-structured interview, a developmental and psychiatric history, validated rating scales, and, ideally, collateral report from someone who knows the person well.[28]

Where rating scales fit — and their ceiling

Self-report scales are the workhorse of ADHD screening. The ASRS v1.1, developed with the World Health Organization, is the most widely used; its six-item Part A is a validated screener with good sensitivity for a first pass.[29] Longer instruments — the Conners scales, the Barkley scales, the DIVA-5 structured interview, and the Wender Utah Rating Scale for childhood symptoms — add depth and a developmental view. But every rating scale shares two ceilings: it measures perceived symptoms, which are subject to recall bias, current mood, and motivation (including, in some settings, incentive to over- or under-report); and a positive screen has limited specificity, because the items overlap heavily with anxiety, depression, and ordinary stress. This is precisely why a scale is a screen, not a diagnosis — and why pairing it with an objective performance measure is attractive, even if that measure is itself imperfect.

What cognitive tests actually detect — the numbers

Continuous Performance Tests (CPTs) and Go/No-Go paradigms are the most-studied objective measures in ADHD. The honest summary of decades of work is that they carry a real but modest signal. Meta-analysis puts the pooled discrimination of a CPT between ADHD and non-ADHD groups at roughly AUC 0.78, with sensitivity near 0.75 and specificity near 0.71.[11] In plain terms: better than a coin flip and clinically informative, but nowhere near sufficient to rule the condition in or out on its own. A meaningful fraction of people with ADHD perform normally on a brief, novel, engaging computer test — the very novelty that recruits attention masks the real-world deficit — and a meaningful fraction of people without ADHD score in the "impaired" range because they are tired, anxious, or simply having an off day.

The single most robust marker: reaction-time variability

Within that literature, one signal stands out. Across a meta-analysis of 319 studies, the measure that most reliably separates ADHD groups from typical groups is not slow average speed or even error count — it is intra-individual reaction-time variability, how much a person's response speed wobbles from trial to trial.[22] The pattern fits the leading cognitive-neuroscience account of ADHD: lapses of attention intrude on otherwise normal performance, producing an occasional very slow response — a "long tail" in the reaction-time distribution — rather than uniform slowing. Commission errors (responding when you should withhold) index the impulsivity/inhibition side and add signal; mean reaction time and the vigilance decrement (drifting slower or sloppier as a boring task drags on) round out the profile.[9] Cognogram deliberately foregrounds reaction-time variability — in the Go/No-Go, in the Stimulant Response view, and as the first plotted signal in Convergence — because that is where the evidence is strongest.

Why one test is never enough — the case for convergence

Here is the mechanistic heart of it. ADHD is not a single deficit; it is a profile — variable attention, weaker response inhibition, sensitivity to cognitive load, and often difficulties with working memory and interference control. Any single task probes one facet and misses the others, which is exactly why single-test AUCs plateau in the high-0.7s. The productive response is not to hunt for one magic test but to combine several partially independent signals that each capture a different facet, and read them together. This is standard clinical logic — a rating scale plus collateral report plus a cognitive profile plus, where relevant, an on/off-medication comparison — and it is the design principle behind Convergence: an ASRS-6 screener, a Go/No-Go for variability and inhibition, a Flanker for interference control, and — added in v2 — a repeated fixed probe plus a reward manipulation that measure how variability moves, lined up side by side. Several weak signals pointing the same way is far more informative than any one of them, while remaining explicitly short of a diagnosis.

Why we show the overlap

Every Convergence result plots your score against both a typical distribution and an illustrative ADHD-sample distribution — and prints how much the two overlap. That overlap is usually large. Showing it is the point: it makes visible, in one glance, why a single number can't diagnose, and inoculates against the false precision that objective-looking tests invite.

Where objective testing genuinely shines: tracking treatment

The evidence for cognitive testing is stronger for monitoring than for diagnosis. The on-medication versus off-medication change within one person is a larger, more reliable effect than the cross-sectional difference between ADHD and control groups, because it cancels the enormous between-person variation that muddies group comparisons. Stimulants measurably reduce reaction-time variability and commission errors, and a within-subject drop from an unmedicated to a medicated session is a meaningful, personal signal a prescriber can factor into titration.[22] This is the logic behind Cognogram's Stimulant Response view. Two honest caveats travel with it: medication often improves cognition without fully normalizing it,[24] so a still-elevated score on medication does not mean the medication is failing; and a browser test is a coarse instrument, easily swamped by sleep, time of day, caffeine, and effort, so runs should be matched on those conditions before they are compared.

A note on FDA-cleared tests, and what Cognogram is not

Some computerized ADHD aids are FDA-cleared as adjuncts to clinical evaluation — the QbTest, which pairs a CPT with motion tracking, is the best known. "Cleared as an adjunct" is a precise and limited claim: it means the device may support a clinician's assessment, not replace it, and its own validation emphasizes medication monitoring as much as diagnosis. Cognogram makes no such claim. It is not a medical device, is not FDA-cleared, and its distributions are illustrative rather than a calibrated clinical norm. It is a self-tracking and education tool that can help a motivated person and their clinician organize evidence and start a conversation — nothing more, and it says so at every step.

How accurate could this realistically get?

A fair estimate, stated before anyone asks. Pooling a validated self-report screener with brief unproctored cognitive measures should land somewhere around AUC 0.80–0.85 against self-reported diagnosis, and 0.75–0.82 against a structured clinical interview — informative, and a long way from diagnostic.

One structural feature helps more than it first appears: self-report and cognitive performance correlate only weakly in ADHD, which means those two blocks carry genuinely different information and combine better than the several cognitive channels do among themselves. That, rather than sheer number of signals, is the real argument for convergence. Several features cut the other way — brief novel tasks recruit attention and mask the very deficit being measured, the sample selects itself, and self-reported diagnosis is a noisy reference standard that caps any AUC computed against it. Improving that last one with a clinically characterised cohort would raise the honest ceiling further than any additional task could. So would screening the conditions that mimic ADHD, which lifts specificity where more cognitive measures cannot. The open roadmap sets out both, along with the sample sizes each would need.

What it would take to make this rigorous — our own validation path

The intellectually honest position is that Cognogram's ADHD signals are hypothesis-generating until validated on real data. Two studies would move the needle, in increasing order of difficulty. First, a within-subject treatment-response study — the achievable win — testing whether reaction-time variability reliably detects an individual's stimulant response across matched on/off sessions; this needs only modest numbers and is genuinely useful even before any diagnostic claim. Second, a diagnostic study testing whether the multivariate Convergence signature, read against clinician-established diagnoses in a properly sampled group of cases and controls, adds incremental value beyond rating scales alone. Both are exactly the kind of citizen-science questions the platform is built to help answer, with consent and appropriate safeguards. Until then, Convergence is framed for what it is: a structured, honest, disclaimer-wrapped starting point — never a verdict.

The ADHD pharmacopoeia

Cognogram does not prescribe, titrate, or advise. But it asks you what you are taking and when, and an entire feature is built on the premise that the answer changes the measurement — so it is worth being precise about what these molecules actually do. Each diagram below is a bloom: every petal is one molecular target, its length is relative binding strength, its colour is the neurotransmitter system, and a filled petal means the drug switches that target on while an outlined one means it switches it off. The shape is the drug’s personality — main effect and side effects in a single glance. Hover or tap a petal for detail.

Methylphenidate

Ritalin · Concerta
A reuptake blocker. It dams the dopamine transporter so that dopamine which was going to be released anyway lingers in the synapse.

Amphetamine

Adderall · Dexedrine
A releaser. It runs the transporter backwards, pushing dopamine out whether or not the neuron fired — a pump rather than a dam.

Lisdexamfetamine

Vyvanse · prodrug
Amphetamine wearing a lysine molecule. Inactive until the body cleaves it — which flattens the peak and makes it far harder to misuse.

Atomoxetine

Strattera · non-stimulant
A single petal. Blocks norepinephrine reuptake only — which still raises dopamine in the prefrontal cortex, but with no euphoric peak and no abuse liability.

Guanfacine

Intuniv · α2A agonist
Never touches dopamine at all. Strengthens prefrontal signalling directly by acting on postsynaptic α2A receptors.

Viloxazine

Qelbree · non-stimulant
Norepinephrine reuptake block with an added serotonergic tilt — the newest arrival, and the reason the extra petals are worth drawing.

Blocker versus releaser — the distinction that matters most

Methylphenidate and amphetamine are routinely described together as “stimulants,” which obscures the fact that they work by different mechanisms. Methylphenidate blocks the dopamine and norepinephrine transporters: dopamine that a neuron releases on its own schedule stays in the synapse longer. Amphetamine reverses those transporters and additionally unloads storage vesicles, pushing dopamine out into the synapse independent of whether the neuron fired at all.[35]

Three practical consequences follow, and all three are visible in the blooms. Amphetamine produces a larger, less physiologically-gated rise, which is part of why it carries the steeper misuse liability. The two drugs are not interchangeable: a meaningful fraction of people respond well to one and poorly to the other, which is a genuinely odd fact for two compounds aimed at the same system, and a reason trials of both are standard practice. And the mechanistic difference is why lisdexamfetamine exists at all — if the problem with a releaser is the sharpness of its peak, the fix is to make the body assemble the drug slowly.

Several different mechanisms, one clinical picture

Look at the non-stimulant blooms and notice how little they have in common with the stimulant ones, or with each other. Atomoxetine blocks a single transporter. Guanfacine is an agonist that never engages dopamine anywhere. Viloxazine adds serotonergic activity to a noradrenergic backbone. Yet all of them produce measurable improvement in the same clinical presentation.[36]

That should be read as evidence about the condition rather than about the drugs. If four distinct molecular routes all improve the picture, the picture is probably not one thing with one cause — a conclusion that arrives independently from the cognitive literature, where no single task or measure separates ADHD cleanly, and which is the whole reason Convergence combines several partially-independent signals rather than trusting any one of them.

Why more is not better: the inverted U

Prefrontal function depends on catecholamine signalling that has an optimum rather than a maximum. Too little norepinephrine and dopamine and the prefrontal cortex is underpowered; too much — the state produced by stress, or by an excessive dose — and it is impaired again, this time by a different route. The relationship is an inverted U, and it is one of the better-replicated findings in the pharmacology of attention.[37]

This is why “it stopped working, so increase the dose” is sometimes exactly the wrong move, and why subjective report is a poor guide: the felt intensity of a stimulant keeps rising past the point where performance has started to fall. An objective within-person measure will not tell you your dose — that is a prescriber’s job and requires information a browser cannot see — but it can tell you whether the thing you are optimising for has actually moved, which is a better question than whether you feel like it has.

What Cognogram can and cannot contribute here

Can: a within-subject comparison of on-medication against clean off-medication runs, matched for time of day, leading with reaction-time variability because that is the metric most reliably linked to stimulant response in the literature. A dated record of that comparison. A visible effect size rather than an impression.

Cannot: recommend, select, or titrate a medication. Detect any of the things that actually govern prescribing decisions — cardiovascular history, appetite and growth, sleep architecture, mood, tics, substance-use risk, pregnancy, drug interactions, diversion. Distinguish a genuine non-response from a bad testing day. Or substitute for the prescriber who is monitoring all of the above. Treat any output here as one line of evidence to bring to that conversation, and nothing more.

A note on why we drew these at all

You do not need receptor pharmacology to use Cognogram. It is here because the alternative — a tool that asks what you are taking, silently treats it as a label, and never explains why the label matters — teaches you nothing. The blooms are the shortest honest route from “I take Vyvanse” to understanding why its curve looks different from Adderall’s on your own timeline.

Caffeine: the drug nobody counts

Caffeine is the most widely consumed psychoactive substance on earth, and in the United States roughly eight in ten adults take it daily.[38] It is also the only psychoactive drug most people take without ever thinking of it as one. For a tool that measures attention, that is not a footnote. It is the single largest uncontrolled variable in the entire dataset — present in most sessions, rarely reported, and acting directly on the system being measured.

Caffeine

1,3,7-trimethylxanthine
One petal that matters. Caffeine is an adenosine antagonist — it blocks a receptor rather than stimulating one. The second petal, phosphodiesterase inhibition, only becomes relevant at doses far above anything you would drink.

It lifts a brake; it does not press an accelerator

Adenosine accumulates in the brain across every waking hour as a by-product of cellular energy use. As it builds, it binds A1 and A2A receptors and dampens arousal — this is a large part of what sleep pressure physically is. Caffeine is shaped closely enough like adenosine to occupy those receptors without activating them.[39] It does not add alertness. It conceals accumulated sleep debt, while the debt itself continues to accrue underneath.

That one mechanistic fact explains nearly everything that is otherwise confusing about caffeine. Why the effect fades as the day goes on and adenosine keeps rising. Why the crash arrives when occupancy falls and a backlog of adenosine finds its receptors free. Why it wrecks sleep so effectively — it is interfering with the exact signal that initiates it. And why withdrawal is a real, characterised syndrome rather than a figure of speech: receptors up-regulate in response to chronic blockade, so removing the blockade leaves an over-sensitive system.

Why it is the right probe for validating a cognitive test

If you want to check whether a browser-based attention test measures anything physiological, you need a manipulation that is legal, safe, cheap, ethically unproblematic, precisely dosable, well characterized pharmacokinetically, and known to act on the measured system. There is essentially one candidate. Caffeine acts on the same adenosine machinery that sustained-attention performance depends on; its kinetics are among the best described in all of pharmacology; and almost every user already has a supply and a habit. Ignition is the paradigm built on this: it models how much caffeine is in your body at the moment you test, then asks whether your performance tracks that curve. The falsifiability is the point — if the model says you are near-peak and nothing moves, the tool has failed a test it set for itself.

The variation between people is the argument for n-of-1

Roughly 95% of caffeine’s primary metabolism runs through a single liver enzyme, CYP1A2, and its activity varies enormously between individuals for both genetic and environmental reasons.[40] The population half-life of about five hours conceals an individual range that spans roughly two hours to more than ten. Oral contraceptives roughly double it. Smoking induces the enzyme and substantially shortens it. Pregnancy extends it dramatically, particularly in the third trimester. Common variants in the adenosine receptor gene ADORA2A track how much anxiety a given dose produces, and habitual intake shifts the whole picture again.

The consequence is that population advice about caffeine is unusually weak advice. A number derived from a group mean may be wrong for you by a factor of four in either direction. This is precisely the situation in which measuring yourself beats reading about everyone — and it is the clearest illustration in the whole whitepaper of why Cognogram is built around within-subject comparison rather than percentile ranks.

The tolerance trap — and the one question you can actually settle

Here is the most interesting unresolved question in caffeine research. Regular consumers are, by the time they reach their morning coffee, mildly withdrawn. So when performance improves after that first cup, is caffeine lifting them above their true baseline — or merely restoring them to it? The withdrawal-reversal hypothesis argues the latter: that for habitual users much of the apparent benefit is the correction of a deficit the habit itself created.[41] The evidence is genuinely mixed, with some studies finding no net benefit in people who have never been regular consumers and others finding real effects.

Population studies may never settle this cleanly, because the answer plausibly differs between people. But it is one of the very few open questions in cognitive pharmacology that an individual can meaningfully address about themselves, and it needs exactly the apparatus Cognogram already has: a stable personal baseline, a repeatable sensitive task, and a modelled drug curve. The shape of the experiment is straightforward — establish a baseline in your normal caffeinated state, taper deliberately, wait out the washout, re-baseline clean, then reintroduce a known dose and see where it lands relative to both baselines.

Take this seriously before you try it

Caffeine withdrawal is a recognised syndrome, not a metaphor. Headache is the most common feature, alongside fatigue, low mood, irritability and difficulty concentrating; onset is typically 12–24 hours after the last dose and severity peaks somewhere around a day or two in, with symptoms persisting up to about nine days.[42] It appears in DSM-5 as a diagnosis for good reason. Do not run this experiment during a week that matters, do not do it at all if you have a history that makes it unwise, and taper rather than stopping abruptly. This is a real intervention, not a app feature.

Caffeine and sleep: the arithmetic is unforgiving

Take the five-hour half-life at face value and follow a 200 mg afternoon dose — roughly a large coffee — taken at 3 pm. At 8 pm about half of it is still circulating. At 1 am, a quarter. Controlled work has found that a substantial dose taken even six hours before bed measurably reduces total sleep time, and — the finding that matters most for self-tracking — participants did not reliably notice.[43]

Then the loop closes. Worse sleep raises next-day adenosine pressure, which raises caffeine intake, which further degrades sleep. This is one of the most common self-sustaining patterns in modern cognition, it is entirely invisible from the inside, and it happens to be visible from the outside using two things Cognogram already records: your sleep and your reaction-time variability. If your variability is elevated and your sleep is short, the honest first hypothesis is not ADHD.

What we could actually pin down

Caffeine is the most tractable of the interventions here, but only for the right contrast. Morning coffee measured against a genuinely washed-out baseline is a reasonably large effect and lands within reach of roughly a dozen matched sessions. A second cup on an already-caffeinated rested morning is close to undetectable, and no amount of testing will change that — there is very little there to find. Ignition improves on a simple paired design by sampling repeatedly across the modelled curve inside a single session, which is worth perhaps three to five times the statistical efficiency. The open roadmap works through that arithmetic in full, including what it would take to settle the withdrawal-reversal question at population scale.

Caffeine and ADHD specifically

Undiagnosed adults frequently self-medicate with caffeine, and often describe it as the only thing that helps. That is worth taking seriously as a clinical observation, and it is worth being clear about mechanistically: caffeine is not a weaker version of a prescribed stimulant. It works on an entirely different system. Prescribed stimulants act on dopamine and norepinephrine transporters; caffeine blocks adenosine receptors and reaches catecholamine signalling only indirectly and diffusely. It is a blunt instrument aimed at general arousal, which is why it can take the edge off inattention without doing much for the executive picture.

The practical implication for anyone using Convergence is concrete. Caffeine state moves reaction-time variability substantially — comfortably enough to change how a session reads. Test in a consistent state, log what is on board, and treat a single caffeinated session as the weak evidence it is. Being able to say which state you were in is exactly the difference between a measurement and a number.

Honest limits

Nothing here is dosing advice. Caffeine is not benign for everyone: it aggravates anxiety and panic in susceptible people, interacts with cardiac and psychiatric conditions and with several medications, warrants restriction in pregnancy, and is genuinely dangerous in concentrated powdered form, where the gap between a normal dose and a lethal one is smaller than people assume. The pharmacokinetic model in Ignition uses population averages and will be wrong for you by an unknown margin — which, as above, is less a flaw than the reason to run the experiment on yourself rather than trusting the curve.

Creatine: the one that only works when you're broken

Creatine is the most interesting supplement in this whitepaper, for a reason that has nothing to do with how well it works. It is the clearest known case of a compound that does almost nothing to a rested brain and something measurable to a depleted one. That makes it a poor product and an excellent experiment — and it happens to be testable with the exact metric Cognogram already computes.

Creatine

methylguanidoacetic acid
Not a receptor drug. Creatine is a battery: the creatine kinase system holds phosphate in reserve and hands it to ADP to regenerate ATP faster than metabolism can. The pale outer petals are receptor interactions reported in preclinical work only — drawn small on purpose.

What it actually is

Creatine is not a stimulant and not a neurotransmitter. It is a buffer. Bound to phosphate as phosphocreatine, it sits in tissue as a rapidly accessible reserve that regenerates ATP through the creatine kinase reaction — faster than glycolysis, far faster than oxidative phosphorylation.[44] That is why it works in sprinting, and it is the whole basis of the hypothesis that it might work in thinking: the brain is metabolically expensive, and any state that drains its energy reserves is a state a bigger buffer might protect against.

The problem is getting it there. Unlike muscle, the brain is largely walled off: creatine crosses the blood–brain barrier poorly, transport depends on a single carrier (SLC6A8) that is sparse at the barrier, and the brain synthesises much of its own supply locally. This is worth being blunt about, because it is the opposite of how the popular account usually runs: an oral dose does not “bypass” the blood–brain barrier. The barrier is the entire obstacle, and it is why weeks of daily supplementation move brain creatine only slightly while muscle saturates comparatively fast.[45]

The evidence, graded honestly

Meta-analysis gives creatine a real but narrow cognitive effect, and the subgroup structure matters far more than the headline. Pooling ten randomised trials, memory improved with a standardised mean difference of 0.29 — but split by age, the effect was 0.88 in adults aged 66–76 and 0.03 in adults aged 11–31, which is another way of writing zero.[46] A larger 2024 meta-analysis of sixteen trials found memory improved (SMD 0.31, graded moderate certainty) and processing speed improved (low certainty), but reported no significant effect on global cognition or executive function — and its attention finding was subsequently corrected to null.[47] A 2024 systematic review is blunter still, arguing the literature fails to support the theoretical basis for a cognitive effect at all.[48]

So the honest summary is not “creatine makes you smarter.” It is: in rested young healthy adults, creatine does approximately nothing to cognition. Where signal appears, it appears in people whose brain energetics are compromised — by age, by depleted baseline stores, or by acute stress. And that pattern is exactly what the transport biology predicts. In healthy people at rest, supplementation barely shifts brain creatine at all.[45]

The sleep-deprivation exception

Then there is the one condition where an acute dose does something, and it is the most interesting result in the field. A Jülich group reasoned that if uptake is the bottleneck, two things might open it together: a very high concentration of creatine outside the cell, and a cell under enough energetic stress to pull harder. Sleep deprivation supplies the second. In 2024 they gave fifteen sleep-deprived participants a single 0.35 g/kg dose — around 25 g — and measured both brain high-energy phosphates and cognition through the night. Both moved.[49]

In 2026 the same group replicated it at a lower, more realistic dose: 0.2 g/kg (mean 14 g) in 29 people, double-blind, randomised, crossover, with each participant serving as their own control on two nights a week apart.[50] Creatine did not abolish the decline — everyone still got worse as the night wore on, and subjective fatigue rose just the same. What it did was flatten the slope, by something in the range of 6–12% relative to placebo.

The specific findings matter enormously for this project. Against placebo, pooled across the night, creatine improved logic (+6.1%), numeric ability (+6.2%), language processing speed (+12.3%) — and reaction-time dispersion on the psychomotor vigilance task (+9.2%). Dispersion. Not mean speed: the consistency of responding. Women benefited more than men across nearly every measure, and vegetarians showed their clearest gain in the slowest 10% of reaction times — the lapse tail.

Why we are dwelling on this

Reaction-time dispersion and the slowest-10% tail are not incidental outcomes. They are the two headline metrics Ignition and the Pulse already compute, on a psychomotor vigilance task of the same family. Of every intervention discussed in this whitepaper, creatine under sleep loss is the one whose published effect lands most precisely on the instrument we happen to have built.

What that study does not establish

Four limits, stated plainly. The effect comes from one laboratory; the 2026 paper replicates the 2024 paper, but the group has not yet been independently replicated elsewhere. The samples are small and young — 15 and 29 participants, aged 20–40 — and the authors say directly that the findings may not generalise to older adults. Many individual time-point results did not survive Bonferroni correction; the robust findings are the pooled ones. And the doses were single large boluses, 14–25 g at once, in people who were not already supplementing.

That last point deserves its own sentence, because it is where the popular summary quietly overreaches. The claim that a chronically supplemented person is “already saturated” and therefore gains nothing acutely is plausible but unestablished. Brain creatine is much harder to saturate than muscle, nobody has shown that daily dosing fills the cerebral pool, and to our knowledge no published acute study has enrolled habitual users. Whether an acute dose still helps someone already taking 10 g a day is, as of this writing, an open question with no data behind it in either direction.

Who has the least, and therefore the most to gain

Creatine comes from the diet almost entirely via red meat and seafood, so plant-based eaters carry lower peripheral stores and are the classic responder group. Be careful about how far that extends, though: lower muscle and plasma creatine in vegetarians is well established, whereas differences in brain creatine are smaller and less consistent than the popular account implies. The 2026 data fit the softer version — vegetarians improved, but the gap between vegetarians and omnivores was smaller than the gap between women and men.[50]

The sex difference is the more striking finding, and it has a mechanism worth knowing. Brain creatine in women is estrogen-sensitive and varies across the menstrual cycle, falling in the luteal phase and after menopause; and SLC6A8, the transporter gene, sits on the X chromosome, where X-inactivation and allelic variation plausibly produce different transporter expression.[50] Lower baseline, more headroom. The same logic covers older adults, who show the largest effects in the meta-analyses.[46]

Can Cognogram actually test this?

Yes — more directly than for any other supplement, because the published outcome and our primary metric are the same quantity. But not in the way Ignition works, and the difference is instructive.

Caffeine is easy to instrument because it is fast, strong and reversible: peak within an hour, a large effect, gone by tomorrow. Creatine is the opposite on all three counts. Chronic supplementation takes weeks to move brain stores and roughly a month to wash out, the effect is small, and in a rested person it is absent. A daily-dosing A/B is therefore a three-month protocol with a low prior of detecting anything — and we would rather say that than sell it.

The acute sleep-loss protocol is a different matter. It is short, it is paired, it has a published effect size, and it targets the exact metric we measure. The design mirrors Gordji-Nejad directly: an evening baseline Pulse, a dose, then Pulses at roughly three, five and a half, and seven and a half hours into a short night — run twice, once with and once without, on nights matched for bedtime and caffeine. The endpoint is fixed in advance: change in reaction-time dispersion from evening baseline, on the creatine night minus the control night. That difference-in-differences is precisely the analysis the published trials used, which means an individual result is directly comparable to the literature rather than merely suggestive.

The honest obstacle: can you afford the experiment?

A 9% shift in dispersion is not large relative to how much an individual's dispersion moves between nights for reasons having nothing to do with any supplement. Which raises the question every self-experiment should answer before it starts and almost none do: how many paired nights would it take to detect an effect this size, given how noisy you personally are?

The figure below does that arithmetic. Set the effect you are hoping to find and your own night-to-night noise, and it returns the number of paired nights required — and, more usefully, what a run of the length you are actually willing to do could and could not detect. For most people, the published creatine effect sits in the region where an honest answer is “you would need more nights than you are going to run.” That is not a reason to skip the experiment. It is a reason to know in advance that a null result will mean underpowered rather than ineffective, which is the single most common error in personal supplement testing.

The question worth collecting

There is one creatine question that a platform like this is better positioned to answer than any individual laboratory, and it is the one the literature has left completely open: does an acute dose still do anything for someone who is already chronically supplemented? Every published acute trial enrolled people who were not taking creatine. Meanwhile a large and growing number of people take it daily and take extra after a bad night, on a theory that has never been tested. Comparing acute-dose responses between habitual users and creatine-naive users, at scale, under the same paired sleep-loss protocol, would settle it. That is a genuine open question, it needs exactly the apparatus described above, and it is the kind of thing self-testing at population scale is actually good for.

Where this sits on the ladder

To be concrete about the conclusion the figure above reaches: at the published effect size, an individual would need on the order of sixty paired sleep-deprived nights to establish a creatine effect on themselves. That study does not get run. But the same sixty pairs spread across thirty people doing two nights each is entirely ordinary — which is the whole argument of the open roadmap, and the reason this module exists despite being individually underpowered. Its value is what it contributes to the aggregate, and we would rather say that than imply otherwise.

Honest limits

Creatine monohydrate has an unusually good safety record at customary doses, and the widely repeated kidney concern is largely an artefact: creatine raises serum creatinine, which is what kidney panels measure, without impairing kidney function — so tell your doctor you take it before anyone interprets a lab result. Single large doses draw water into the gut and cause cramping, bloating and diarrhoea in a substantial minority; the acute protocol above involves exactly such a dose and is a real intervention, not a app feature. Nothing here is dosing advice, none of it is a reason to deliberately lose sleep, and no supplement outperforms the night's sleep it is being used to paper over. Creatine's best-evidenced cognitive use is mitigating a deficit you would be better off not having.

Ignition: a caffeine-sensitive challenge paradigm

Most of Cognogram asks "how are you doing versus your baseline?" Ignition asks a sharper, falsifiable question: can this tool detect a known pharmacological change as it happens? The previous section makes the case for why caffeine is the right probe; this one is the apparatus. If a lightweight browser test can track a caffeine dose rising and falling in your bloodstream, that is real evidence it measures something physiologic, not noise.

The task: a psychomotor vigilance test (PVT)

Ignition is a Psychomotor Vigilance Test — the reference paradigm sleep and stimulant researchers reach for precisely because it is so sensitive to changes in alertness. A target appears at random intervals (2–10 seconds); you respond as fast as you can. Over dozens of trials it yields not just a mean, but the shape of your attention: the fast tail, the slow tail, and crucially the lapses — responses slower than 500 ms, which are the single most caffeine- and sleep-sensitive feature of the task.[25] We offer a 60-second Quick tier for frequent sampling and a 3-minute Full tier for maximum sensitivity. Anticipations (responses faster than 100 ms) are flagged and excluded from the reaction-time statistics, per standard PVT scoring.

We report the metrics the literature favors for this purpose: median reaction time, lapse count, the slowest-10% mean, reaction-time variability, and mean reciprocal reaction time (1/RT) — the last of which is preferred because it is robust to the long right-tail of lapses and behaves well statistically.[25]

The pharmacokinetics: modeling caffeine on board

Rather than treat caffeine as a simple yes/no, Ignition models how much is actually in your body at the moment you test. Caffeine follows well-described first-order kinetics: near-complete oral bioavailability, absorption peaking around 45 minutes, and an elimination half-life of roughly 5 hours in a typical healthy adult.[26] We implement a standard one-compartment model with first-order absorption — the fraction of a dose in the body over time is

A(t) = (ka / (ka − ke)) · (e−ket − e−kat)

with an absorption rate that places the peak near 45 minutes and an elimination rate set by the 5-hour half-life (ke = ln2 / 300 min). Multiplying by your reported dose gives an estimated milligrams on board for every run. The curve below is what that looks like for a single 100 mg dose — roughly one strong cup of coffee.

Modeled caffeine on board after a 100 mg dose. Peak near 45 min; about half remains at 5 hours (one half-life); effectively cleared by ~24 hours (≈5 half-lives). Individual half-life varies with genetics, smoking, pregnancy, and oral contraceptives, so we treat this as a defensible population estimate, not a personalized assay.[26]

The "negligible concentration" gate

A clean baseline run needs you to be genuinely caffeine-free — but "did you have coffee?" is too blunt, because dose and timing matter enormously. The app uses the kinetics to guide the answer honestly: within the first ~45 minutes you are still absorbing; from ~45 minutes to 5 hours you are at or near peak; by 10 hours (about two half-lives) the acute effect is largely gone; and after ~24 hours (five half-lives, ~97% cleared) you are effectively washed out. The gate tells you which window you are in rather than forcing a false binary.

Validation by design: the PK label as ground truth

Here is the methodological heart of it. Instead of validating against subjective "I feel caffeinated," Ignition uses the pharmacokinetic estimate itself as the ground-truth label: a run is designated ON if modeled caffeine on board exceeds a threshold you set, OFF otherwise. Performance is then the predictor, and the question becomes a clean classification problem — how well does the test's reaction-time signal separate PK-labeled ON runs from OFF runs? The in-app Caffeine Lab computes this as a sensitivity/specificity trade-off with a sweepable threshold, reporting the threshold-independent area under the ROC curve (AUC) as the headline. This mirrors exactly how the CPT literature quantifies clinical utility — pooled AUC around 0.78 for ADHD identification, for instance.[11]

The second, richer analysis is the time-course: plotting performance against minutes-since-dose and overlaying the modeled caffeine curve. If your reaction time improves and decays in step with the pharmacokinetics — better near the 45-minute peak, drifting back over the following hours — that temporal correspondence is far stronger evidence of a genuine pharmacodynamic effect than any single before/after pair.

Honest limits

This is deliberately framed as n-of-1, hypothesis-generating work — not validation. Time of day and sleep pressure move PVT performance as much as a modest caffeine dose, so runs must be paired at matched clock times to mean anything, and the primary metric and threshold should be fixed in advance rather than chosen after seeing the data. Because the PK label derives from your own self-reported dose and timing, the strongest version of this design adds a blinded arm — caffeinated versus visually identical decaf, where you don't know which — which is why the data model already carries a blinded field. Until then, Ignition validates whether the test can detect the PK-predicted effect, which is a real and useful claim, distinct from "detects caffeine better than simply asking."

A citizen-science invitation

Every Ignition run — caffeinated or not — is a data point in a question that genuinely hasn't been answered at scale: how sensitively can a two-minute browser task detect everyday pharmacology? Runs are exportable as CSV for real statistical analysis. Contributing a handful of your own paired runs, honestly logged, is a small act of open science.

FEATURES

The experiment suite, in practice

In plain terms

Cognogram is organized into two halves. Tests are the cognitive battery you run against your own baseline. Experiments are where the citizen science lives — structured tools that ask sharper questions: can a quick test detect a drug? Is your medication working? Do several attention signals line up? This section is the manual for each.

Everything here is built on the same eight tests and the same honesty principles — it just packages them to answer specific, falsifiable questions. Each tool states what it can and cannot conclude, right in the interface.

The Ignition Lab — your analysis bench

Once you've logged a handful of Ignition runs, the Lab turns them into a proper analysis. It treats the pharmacokinetic model as ground truth — the modeled milligrams of caffeine (or stimulant) on board define whether a run was "on" or "off" — and then asks how well your performance predicts that label. A sweepable threshold lets you watch sensitivity and specificity trade off in real time, and the tool computes an ROC curve and its AUC (via the Mann–Whitney relationship) so you can see, in one number, how separable your caffeinated and clean runs actually are. A labeled time-course plots each run's score against hours-since-dose with the modeled drug curve overlaid. This is the same apparatus a pharmacology study would use, scaled to an n-of-1. It is explicitly hypothesis-generating: a single person's runs can suggest sensitivity, never establish it.

Stimulant Response — treatment tracking, done soberly

This is the clinical sibling of Ignition, and it deliberately looks different — quieter, instrument-like — because its question is serious: does your ADHD medication measurably steady your attention? It compares your on-medication runs against your clean, off-medication runs, within yourself, and leads with reaction-time variability because that is the metric most reliably linked to stimulant response in the literature.[22] The readout is a within-subject effect size (Cohen's d) with a plain-language magnitude, an on-versus-off comparison, and a per-run distribution strip. It needs at least two on-med and two clean off-med runs, ideally matched for time of day to control circadian effects. Two honest limits travel with every result: medication frequently improves cognition without fully normalizing it,[24] so a still-elevated score doesn't mean the medication is failing; and a "no clear difference" result doesn't mean it isn't working — a browser test is coarse and easily swamped by sleep, timing, and effort. It is a conversation-starter for a prescriber tuning a dose, never a treatment decision.

Convergence — the ADHD screening-support stack

Convergence is the feature that operationalizes this whitepaper's core argument about ADHD. Version 1 asked: do several weak signals point the same way? Version 2 asks a sharper question — is your variability a trait or a state? — and changes three things to answer it.

It measures variability at two timescales. A fixed simple-RT probe (the Pulse) is repeated two or three times across the session. It runs 40 trials, about three minutes — raised from 14 in August 2026 because the sampling error on a dispersion estimate falls only as 1/√(2(n−1)), so 14 trials carried roughly 20% error before any biology entered, larger than most of the effects being hunted. Forty brings that to about 11%. Runs recorded under the shorter protocol are retained, labelled, and not pooled with the new ones for dispersion. It is deliberately never varied: its entire value is being the same instrument every time it appears. One probe gives the familiar trial-to-trial number; repeating it gives minute-to-minute drift, and — because you now have several independent readings of the same quantity — a genuine confidence interval on the measurement itself. A wide interval is reported as such, because a session that measured you imprecisely should say so rather than quietly returning a number.

It provokes variability rather than only observing it. A four-choice reaction-time task (Four Corners) runs twice: once slow and unrewarded, once fast and scored, after the design of the Fast Task.[32] The research finding is not simply that ADHD groups are more variable — it is that they are markedly more variable under low arousal and improve more when event rate and incentive are raised. So the quantity of interest is the difference between conditions, not the level in either. That matters for an unproctored browser tool in particular: a within-session difference cancels device latency, motor speed, age, and motivation-to-look-good, which are precisely the confounds we cannot otherwise control. We report it as a state-regulation index and label it exploratory — it has no published operating characteristics, so it carries deliberately little weight.

It reports a probability, not a count. Version 1 counted how many of four signals leaned ADHD-ward. That framing weights every channel equally and hides the base rate. Version 2 starts from adult population prevalence and updates it with an approximate likelihood ratio per channel, showing the full table of what each contributed. Two honesty mechanisms are built into the arithmetic. First, the cognitive channels are down-weighted, because they are correlated with one another — they share method variance and a common latent attention factor, and treating them as independent evidence would produce a confident, wrong posterior. Second, the result carries a structural-uncertainty band: how far the answer moves if those channels overlap more, or less, than assumed. It is wide, and it is meant to be.

The consequence is an asymmetry we print on the result screen rather than hide: with a starting rate near 5% and instruments this coarse, the stack is structurally better at lowering the estimate than raising it. A clean sweep is meaningful reassurance. A full house of elevated signals still lands well short of a diagnosis — it lands at “worth a proper evaluation.”

Two further decompositions are reported for mechanism rather than classification. The reaction-time distribution is fitted with an ex-Gaussian model, separating the ordinary bulk (μ, σ) from the long right tail (τ) that carries the ADHD signal in the literature — because the finding is not uniform slowness but ordinary performance punctuated by lapses.[33] And the two-alternative Arrows task is passed through a closed-form EZ-diffusion model, splitting performance into quality of evidence accumulation (drift rate), response caution (boundary), and everything non-decisional — perception, thumb, screen latency.[34] That last term is useful in a way specific to browser-based testing: device lag lands in non-decision time, leaving drift rate as the least device-dependent number Cognogram produces. We are explicit that neither decomposition improves the call — τ does not out-discriminate plain reaction-time variability[33] — and both are fitted on far fewer trials than the methods want. They explain the mechanism; they do not sharpen the verdict.

Every screen states plainly that this is screening support, not a diagnosis and not a medical device: a real evaluation needs a clinician, DSM-5-TR criteria, cross-setting impairment, childhood onset, and differential diagnosis.[27] The result closes by naming the three things that would actually settle the question — repeat testing across varied days, ruling out the conditions that mimic ADHD, and bringing a dated record rather than a verdict to a clinician.

The Arcade — the same paradigms, scored as games

Collecting repeated measurements at population scale requires people to come back, and a ten-minute battery is not something anyone runs daily. The Arcade is three paradigms with gamified scoring and no account required: Steady (simple RT), Dead Reckoning (time reproduction) and Hold Fire (go/no-go). Only the scoring is invented; the paradigms are the same ones described above.

The scoring is constrained by the hardware problem set out in section 06. Since device lag inflates the mean and leaves the spread alone, Steady is scored on standard deviation rather than speed or CV — CV would quietly reward a slower phone. Hold Fire reports impulsivity, inattention and variability as three separate numbers rather than one blended score. And Dead Reckoning, measured in seconds where 40 ms is a rounding error, is the only one of the three that is genuinely device-fair, which is why it is the one suited to comparison between people.

One methodological point is stated in the app as plainly as here: these games reward the player, so they are a different condition from the quiet probes and their data is never pooled with an unincentivised Pulse. Used deliberately, the gap between them is itself the state-regulation contrast that Convergence rests on.

Result cards — seeing yourself on the curve

Throughout the experiments, results can be rendered as shareable cards that plot your score against a typical distribution and, for the ADHD-relevant metrics, an illustrative ADHD-sample distribution — with your percentile and, crucially, the overlap between the two curves printed on the card. That overlap is the teaching device: it is usually large, and seeing it makes the central lesson unmissable — that landing near the ADHD shape is common in people without ADHD, and vice versa, which is exactly why a single number cannot diagnose. The distributions are illustrative, assembled from typical published ranges rather than a calibrated norm, and every card says so. Cards export as images to save or share, turning a private result into a small, honest piece of science communication.

How the pieces fit — a map

The mental model: Tests establish and track your baseline across five cognitive domains. Ignition and its Lab ask whether the instrument can detect a known change (caffeine) — validating the whole premise. Stimulant Response applies that same within-subject logic to the question people most want answered — is my medication working? Convergence assembles multiple attention signals into a structured ADHD screen. And result cards wrap it all in honest, shareable visualizations. Each is a different question asked of the same underlying measurements, and each carries its own explicit statement of limits.

07

The open roadmap

This section exists because most self-tracking tools skip it. They ship a measurement, imply it means something, and never state what it could not have detected. What follows is the arithmetic in the other direction: what this instrument can pin down, what it cannot, and what would change that. Some of the answers are unflattering. They are here anyway, because a tool that will not say what it is blind to cannot be trusted about what it sees.

Start with the denominator

Every question of the form “did this do anything?” is a ratio: the size of the effect over the size of your own noise. Almost all the attention in this field goes to the numerator — how big is the effect of caffeine, of medication, of a supplement. The denominator decides whether any of it is knowable, and it is almost never reported.

Our denominator has an awkward property. Cognogram's headline metric is variability, and variability is expensive to measure. The sampling error on an average shrinks with the square root of your trial count; the sampling error on a measure of spread shrinks about four times more slowly. That is not a flaw in the design, it is arithmetic — and it is why the reference psychomotor vigilance task in sleep research runs for ten minutes rather than one.

Read off the current build and the number is uncomfortable: a single 14-trial Pulse carries roughly ±20% sampling error on its dispersion estimate before any biology enters the picture. Pooling three Pulses brings that to about ±11%. Layering on realistic day-to-day biological swing and unproctored device variability puts a plausible total at ±18% per session for variability, against about ±8% for median reaction time.

Those two numbers — and the honest fact that we have not yet measured them on real users — determine everything below. Establishing them empirically is the first item on this roadmap for a reason.

What one person can actually find out

Combine those noise estimates with published effect sizes and you get a ladder. It is worth sitting with, because it reorders the intuitive ranking almost completely.

Four things fall out of that picture.

Sleep loss is the clear winner, and it is not close. The effect is two to four times larger than anything else here, which is why it is the one condition a handful of sessions can call with confidence. It is also, not coincidentally, the intervention with the strongest literature behind it.[18] If Cognogram is good at one thing, it is telling you what a bad night costs you.

Stimulant response is genuinely within reach for a clear responder — on the order of a dozen matched pairs, which is a month of alternating days. Impulsive errors move alongside variability, giving a partly separate second channel and making the answer more robust than the bar alone suggests.

Caffeine depends entirely on the contrast you choose. Morning coffee against a genuinely washed-out baseline is a decent-sized effect. A second cup on top of an already-caffeinated rested morning is close to nothing. Ignition is more efficient than the ladder implies, because it samples repeatedly across the modelled curve within one session rather than comparing two sessions — worth perhaps three to five times a simple paired design.

Creatine, individually, is out of reach. The published acute effect of about 9% on reaction-time dispersion[50] would need something like sixty paired sleep-deprived nights to establish in one person. Nobody runs that study. We would rather say so plainly than let a two-night null read as evidence of nothing happening, when it is really evidence of nothing measurable.

The most common error in personal supplement testing

Treating an underpowered null as a negative result. If you run four sessions on something with a 9% effect, you had roughly a one-in-six chance of detecting it even if it works perfectly. “I tried it and felt nothing” and “it does nothing” are different claims, and the gap between them is exactly the arithmetic above. Every experiment in Cognogram states its power budget before you begin, for this reason.

Then the split that changes everything

Here is the part that makes this worth building. An individual is stuck with their own noise. A population averages it away — and the required evidence does not grow, it merely gets redistributed. Sixty paired sessions is an impossible ask for one person and a trivial one for thirty people contributing two nights each.

Conflating those two situations is the standard error in this entire category. Consumer tools routinely claim population-scale findings for an individual sample size, and researchers routinely dismiss self-tracking because the individual case is weak. Both are looking at the wrong axis.

Notice where the curve stops being interesting. Somewhere around ten thousand participants, statistical power ceases to be the binding constraint and bias takes over — self-selection into an attention-testing app, unverified self-reported diagnoses, unmatched testing conditions, and the ordinary chaos of unproctored measurement. Past that point more people do not buy more truth. That is why the roadmap below front-loads reliability and reference-standard work rather than recruitment: the ceiling on this project is not sample size, it is measurement quality.

The questions we could actually answer

Not aspirations — specific, falsifiable, pre-registerable claims, each with the sample it needs and the reason nobody has settled it yet.

~30 people
Does creatine do anything for a sleep-deprived brain outside one laboratory?The entire acute literature comes from a single group in Jülich.[49][50] Independent replication is the most valuable cheap thing available in this whole document.
~400 people
Does an acute dose still work if you are already taking it daily?Every published acute trial enrolled non-users. Millions take creatine daily and take extra after a bad night on a theory with no data behind it in either direction. This is a genuinely open question and a self-testing platform is the natural instrument for it.
~1,000 people
Does day-to-day variability predict ADHD better than within-session variability?Trial-to-trial inconsistency is the most replicated cognitive marker in ADHD.[20] Whether between-day inconsistency adds anything on top is unknown, because a clinic sees one day. This is the finding that would belong to whoever collects the data first.
~1,000 people
Is the state-regulation contrast better than raw variability?The slow-versus-rewarded difference cancels device, age and motivation by construction.[32] If it out-predicts raw variability, that is a better screening measure than the field currently has.
~2,000 people
Can a retest after sleep remediation separate trait from state?ADHD is persistent; sleep-driven variability should resolve. One session cannot tell them apart. A structured retest can — and no clinic is positioned to run it.
~10,000 people
Who responds, and who does not?Sex, diet, age and baseline stores all appear to moderate creatine response;[50] metabolic genotype does the same for caffeine.[40] Moderator effects need roughly an order of magnitude more people than main effects, which is why almost none are established.

The staged programme

  1. Reliability first, and publicly. Establish test–retest reliability and practice-effect curves in a stable cohort under controlled conditions, and derive a reliable-change index and minimal-detectable-change band for every metric. This replaces the assumed denominator above with a measured one. Until it exists, every number on this page is an estimate, and we will keep saying so.
  2. Convergent validity. Administer alongside an established reference battery in a subset, to estimate how far Cognogram's indices track their validated counterparts.
  3. Known-groups sensitivity. Test whether the battery separates groups it should — sleep-restricted versus rested sessions within person is the cleanest available check, and by the ladder above it is also the one most likely to succeed.
  4. A clinically characterised cohort. Even 150 participants with structured diagnostic interviews would change the class of claim Convergence can make, and would replace self-reported diagnosis as the reference standard. This is the highest-value external partnership available to the project.
  5. Pre-registration, then publication either way. Primary endpoints fixed before collection. Negative results published — particularly if the exploratory measures fail to beat plain variability, which is the outcome the existing literature predicts.[33]
  6. Local norms only where earned. Reference distributions built solely for populations with adequate sampling, and always shown with their limits.

What has been built since this roadmap was written

Three of the items above have moved from plan to product, and the roadmap should say so rather than continue to promise them.

Reliability first — partly delivered. Every task now records what a learning curve needs: which exposure this is for that task, the gap since the last one, device class, and time of day. With five or more runs the app fits an exponential approach to asymptote, y(n) = A + Be−(n−1)/τ, by grid search on τ with A and B solved by least squares at each step. It then reports the split that matters: raw change minus the change practice predicts equals the change actually worth attributing. Against simulated curves the fit recovers τ to within about 5% at r² 0.95–0.99, and a flat series correctly yields no learning.

This carries a limitation we surface in the app rather than bury. A real change occurring during the learning phase is partly absorbed into the fitted curve. In simulation, a genuine 15-unit shift injected early in a 20-run series is recovered at roughly 3%; the same shift late in the series recovers at about 77%. The correction therefore systematically under-detects early change, and a small residual while someone is still climbing should not be read as nothing having happened. The practical instruction that follows — get past the plateau before measuring anything you care about — is the single most useful thing the feature outputs.

Screening the mimics — delivered. The differential panel takes about three minutes: sleep hours and quality, the full eight-item STOP-BANG apnoea screen, PHQ-2 and GAD-2 with automatic expansion to PHQ-9 and GAD-7 when positive, substances and acute stress, common physical contributors, a cognitive-disengagement block, a five-domain impairment scale, and a childhood-onset probe. Impairment and onset are included because they are formal DSM-5-TR requirements that no reaction-time task can reach, and because distractibility without impairment is not a disorder.

The output is deliberately not a flag. It is a ranked list of competing explanations, each with an experiment that could falsify it — sleep debt paired with “three nights of seven-plus hours, then retest at the same time of day”, and so on. Where a positive mood screen includes any endorsement of self-harm, a support block renders before any result, ahead of every cognitive number on the page.

Session quality — delivered. Between roughly a third and a half of adults presenting for ADHD evaluation fail at least one performance-validity check, so this is a data-quality problem rather than an edge case. Cognogram derives its indices from data already collected: responses under 120 ms, response timing too regular to be human (the script or rhythm signature), non-response rate, window-focus loss during a task, and false starts. Sessions are labelled good, limited or questionable with the specific reason given, and questionable sessions are kept in the person’s own record but excluded from trend lines and from any aggregate. The framing is deliberately “was this a good measurement?” rather than an accusation.

The motion channel — the cluster nothing measured

The battery sees inattention (missed targets) and impulsivity (commission errors). Until now it has been blind to the third symptom cluster in the diagnostic criteria: restlessness. That gap is not incidental — it is shared by essentially every browser-based attention test.

It is also the one axis where a phone is genuinely better than a clinic desktop rather than merely cheaper. The FDA-cleared QbTest pairs a continuous performance task with infrared tracking of a head-mounted marker; its smartphone version reports that adding motion and camera-derived features to task data exceeded traditional computer-based accuracy. A phone accelerometer is a cruder instrument than an infrared marker, but it is already in the person’s hand.

Cognogram samples DeviceMotionEvent at up to 50 Hz during a Pulse, high-passes the signal (gravity is a slow constant near 9.81 m/s²; movement is the deviation from it) and retains four summary features: root-mean-square motion, burst rate per minute against a personal threshold, fraction of time immobile, and integrated movement path. The raw trace is discarded when the task ends.

The question we find most interesting is not how much someone moves but when. Movement in the 1.5 s window before each response is correlated with reaction time at lags of −2 to +2 trials, which asks whether restlessness precedes an attentional lapse, arrives with it, or follows it. Movement building ahead of a lapse would fit the arousal accounts of ADHD; fidgeting after one looks more like a consequence of disengaging. To our knowledge nobody holds that pairing at scale, and it needs exactly this combination of movement and response data.

Two constraints, both deliberate

Nothing leaves the device. No raw accelerometer trace is retained or transmitted under any circumstance; only the summary features above are eligible for contribution, and only if cognitive data is opted in. Capture requires an explicit gesture, which iOS enforces anyway.

A phone on a table reads as perfectly still. That is an absent measurement rather than a calm person, and scoring it as low restlessness would be straightforwardly wrong. Sessions whose trace is near-flat throughout are flagged as stationary-device and excluded from both the person’s summary and any aggregate. The channel is exploratory, carries no weight in any screening estimate, and has no published operating characteristics on commodity hardware.

How contribution would actually work

Publishing aggregate findings requires pooled data, and pooled data requires consent that is worth the name. The mechanism is now built and can be inspected in the app under Settings. It is off by default, it is not required to use anything, and it is split into five categories a person accepts or declines individually — cognitive task data, session context, coarse demographics, screening questionnaires, and self-reported diagnosis. The last two are the most sensitive and the most scientifically load-bearing, which is precisely why they are separable rather than bundled.

De-identification is applied before anything leaves the device: no name, email or account identifier; a random pseudonym instead; timestamps reduced to the date with clock time dropped; age reduced to a five-year band; device fingerprint reduced to touch or keyboard with no screen size or model; questionnaire results reduced to totals rather than item-level answers; and free text never included at all. The app shows the person the actual record that would be sent — generated from their own data, not a description of it — and lets them download it.

Where this honestly stands

The pipeline exists; the collection does not. No aggregate dataset is being gathered, because that requires ethical review and a published data-governance policy, and neither should be taken on trust from a website. Until both exist, the feature prepares, displays and exports a contribution and transmits nothing. When collection begins it will be announced rather than switched on quietly, and consent will be sought again rather than inherited.

The device registry — controlling the largest noise source we can

Section 06 establishes that display and input latency spans roughly 40 ms across ordinary hardware, larger than a night of lost sleep. That analysis was previously only an argument; it is now enforced in software.

Every session records a device signature derived from operating system, pointer type, screen dimensions and — measured directly from frame timing rather than reported — refresh rate. From that we estimate the fixed lag each device contributes, on the order of 13 ms for a fast laptop and 34 ms for an older touchscreen. The consequences are applied rather than merely described: reaction speed is never compared across devices, differences in consistency are, because a near-constant offset shifts the mean and leaves the spread intact, and any two-session comparison spanning devices carries that caveat inline. When the device in front of the person differs from their recent norm, the dashboard says so before they read anything into a change.

We detect rather than ask. Self-report about which device was used is worse on every axis — people misremember, nobody can report their own refresh rate, and a question at the start of a session is friction that gets clicked past. The only thing detection cannot supply is a name, so that is the sole thing the app asks for.

The anonymous first run

An honest look at the funnel showed that requiring an account before a person has any evidence the tool is worth using loses roughly four out of five of them, at the step that gates everything downstream. Since five of the six questions in this roadmap require paired within-person data, and a single orphan run answers almost none of them, the design problem is not capture but return.

So the first run takes sixty seconds, asks a single context question — hours slept, because sleep is simultaneously the largest real effect and the largest confound — and requires no account at all. A device-local token pairs run one with run two without any sign-up. The account is offered only after the second run, once the person has seen the comparison, which is the actual product.

The first result says, in as many words, that the number means nothing on its own. That is not modesty; it is the literal measurement situation, since a meaningful fraction of any single figure is hardware, posture and the day. It also happens to be the strongest available reason to come back.

The Open Trial — any compound, properly tested

The supplement literature runs on a recognisable asymmetry: a press release with a large percentage, a small sample, and no practical way for a reader to check. The Open Trial closes that loop for an individual, with four constraints that distinguish an experiment from an impression.

Pre-registration. The expected effect size, the number of sessions, and the deciding metric are fixed before the first dose and cannot be edited afterwards. Self-experimentation typically fails not through bad measurement but because the definition of success is settled after the data arrives.

A power budget computed from the person’s own measured noise, not an assumption, and shown before they commit. Most published supplement claims turn out to be undetectable by any individual: at a typical 18% session-to-session noise, a 19% effect needs about 17 paired sessions, a 5% effect needs over 200, and a 2% effect more than a thousand. Where a claim falls in the untestable range the app says so plainly, which is more useful before three months of dosing than after.

A randomised schedule generated in advance, in balanced on/off pairs, because otherwise people take the compound on days they already expect to go well. And a null result reported as underpowered rather than negative, alongside the smallest effect that run could actually have detected.

The launch example, appraised

The feature ships with a real and current claim: astaxanthin at 12 mg/day for four weeks, reported to improve task-switching reaction time by 19% and performance on a distraction-heavy task by 23% in recreationally active women after a fatigue challenge. The design is sound in principle — double-blind, placebo-controlled, plausible mechanism, since astaxanthin is among the few antioxidants crossing the blood–brain barrier. But the trial enrolled 25 participants in total, roughly a dozen per arm; it was announced by the ingredient manufacturer rather than through independent reporting, with no journal citation or trial registration given; and a 19–23% shift in reaction time would be larger than the effect of caffeine or of a night without sleep. We present it as worth testing rather than as established, which is precisely the distinction the feature exists to let people draw for themselves.

Open by default

Three commitments, because a roadmap without them is a marketing document. Methods are public — this whitepaper is the specification, and every paradigm, threshold and likelihood ratio in the app appears somewhere in it. Analyses are pre-specified, and where a figure lets you move a threshold interactively, that is a teaching tool and is labelled as distinct from a confirmatory analysis. Aggregate data will be publishable, so that findings can be checked rather than believed. If a claim here turns out to be wrong, the useful thing is for someone else to be able to demonstrate it.

What would help

The bottleneck is sessions, not ideas.

Every question in the grid above is limited by one thing: how many people run a consistent protocol and keep running it. A single session contributes almost nothing. Ten matched sessions from one person contributes a usable noise estimate — the denominator this entire page is currently guessing at. A few hundred people doing that turns a list of open questions into a set of answered ones.

If you take creatine, drink coffee, take a stimulant, or simply sleep badly and want to know what it costs you: the most useful thing you can do is establish a baseline and be boring about it. Same time of day, same device, same routine, repeatedly. That is unglamorous, and it is the entire input this needs.

Cognogram is a research preview, not a medical device and not a diagnostic instrument. Data stays on your device unless you choose to sync it. No aggregate analysis will be published without appropriate ethical review, informed consent and data governance — and nothing on this page should be read as a promise that any of these questions will resolve in the direction we find interesting.

08

References

Full bibliographic details should be verified against source records before external publication; journal and year are given as located.