Research agenda

Open questions in observation architecture

If Layer 1 validates — if passive longitudinal observation produces calibrated, target-specific lift beyond the R1 population-prior and R2 own-routine baselines — the next question is how to observe a person continuously over months to years. That is a different optimization problem from any current product. Every question below is answerable with the apparatus already specified: fix the task set, vary the sensing arm, read the evidence-efficiency curve.

Manuscript and benchmark release V1.0 — pre-pilot protocol release. Synthetic harness only. No human-subject results.

The hypothesis ladder

Eight claims, in order of strength — each separately falsifiable

Each rung presupposes the ones below it, and the program can stall or fail at any rung — a failure that is itself an informative result. Nothing above mechanics is claimed today.

H0
Generic predictabilityFuture outcomes can be predicted above population priors (R1). Phase 0 mechanics; Phase 1 examines.
H1
Routine predictabilityA person's own history improves prediction over R1 (an admitted R2). Phase 0 mechanics; Phase 1 examines.
H2
Target-specific informationThe system beats R2 and loses skill under wrong-target permutation. Phase 1 examines; Phase 2 confirms.
H3
Longitudinal dynamicsCorrect temporal order beats shuffled history, especially during transitions. Phase 1 examines; Phase 2 confirms.
H4
Observation valueAdded passive or multimodal evidence beats digital exhaust and self-report. Phase 1 examines; Phase 3 confirms.
H5
Evidence efficiencyComparable gated skill at less energy, data, burden, or privacy exposure — minimum sufficient observation. Phase 3.
H6
Cross-task personal transferA frozen personal representation improves sealed predictions on held-out task families without labeled outcomes from them — the test that separates a reusable model of the person from a per-task rule. Phase 4.
H7
Downstream utilityA validated personal model improves assistance, under a separate intervention and safety protocol. Phase 5.

The staged roadmap. Phase 0 — synthetic harness (complete): proves the pipeline, nothing about people. Phase 1 — five-person, thirty-day feasibility pilot (next): proves the protocol runs on real, consented lives and yields variance estimates; not powered for effects. Phase 2 — powered target-specificity study, sized by simulation from Phase-1 estimates. Phase 3 — evidence-tier and model–evidence frontier study. Phase 4 — cross-task transfer. Phase 5 — downstream assistance utility. No phase borrows the license of a later one, and sample sizes for Phases 2–5 are outputs of Phase 1, not promises.

Four clusters

The agenda

Sensing sufficiency

  • What is the minimum sensing needed to model a person well?
  • Which modalities add measurable predictive lift — and which mostly re-encode routine?
  • Is continuous video necessary, or can sparse images match it?
  • How much audio is sufficient before lift saturates?
  • How much context is enough for each transition type?

Sampling & scheduling

  • What capture schedule maximizes predictive information per joule?
  • Can event-triggered sensing outperform continuous sensing?
  • How should missing data be represented so missingness neither leaks the outcome nor biases the prediction?
  • What battery life is necessary before charging gaps censor the evidence that matters?

Compute placement

  • How much local computation is actually required?
  • Can inference happen offline — deferred to a daily upload — without losing time-sensitive evidence?
  • What evidence value is lost to on-device redaction and privacy filtering?
  • What architecture maximizes calibrated skill per watt?

Uncertainty & representation

  • How should sensor error and coverage gaps propagate into the prediction distribution rather than being imputed away?
  • How should a belief state represent what the device did not see?
  • Which derived representations (transcripts, entities, commitments) preserve predictive information, and which discard it?

Status. These are open research questions, stated as formulation. None has been answered; no experiments have run. The proposed pre-pilot feasibility study (five consenting participants, thirty days, audio-first, daily sealed predictions) is designed to validate only that the harness can run — it is not powered to answer any question on this page.

Contribute

How to work on these questions

Run the harness

The synthetic reference path exercises the full loop — sealing, resolution, scoring, permutation. Run it.

Run a sensing ablation

Fix the task set, vary the sensing arm (tier, schedule, device), and report Skill vs R1/R2 per arm. The playbook.

Study the neighbors

Reality Mining, digital phenotyping, GLOBEM, PULSE, Ego4D, and the memory-benchmark family each solved part of this problem. The reviewed landscape maps what each measures and how it could participate.

Propose an instrument

Evaluate a capture device against the hardware framework: coverage, evidence, research fitness — and pre-register an evidence-efficiency estimate.