Meeting recorders, AI wearables, smart glasses, memory assistants, digital companions, enterprise copilots, passive audio devices, and lifelogging systems ship as separate categories with separate metrics. Their trajectories are not separate. Each accumulates longitudinal context and advertises the same destination: a persistent model of a particular person. TargetSpace is the shared measurement layer that tests whether any of them is actually building one.
Manuscript and benchmark release V1.0 — pre-pilot protocol release. Synthetic harness only. No human-subject results.
Different categories bring different evidence and architectures — but all can face the same target-specific evaluation contract, and the surrounding research landscape is reviewed at Related work. The internal representation is unconstrained: a personal world model is one architectural hypothesis, and TargetSpace requires no world model, memory graph, or simulator — only sealed probabilistic predictions under the protocol. The placement of categories is illustrative, not a ranking. The convergence claim is observational: evidence that the capability class is arriving, not validation of any method and not a claim that any vendor misuses data.
Interaction comes after understanding, not before. The sequence is observe → infer → predict → validate → interact: observation produces evidence, inference maintains a belief over the latent target-state, prediction commits to a distribution before the outcome exists, validation scores it against the sealed outcome, and interaction then acts on a model that has earned its claims. Today's products largely run observe-and-interact, with understanding asserted in between rather than measured. Interaction quality is a legitimate product metric; it is not evidence of a target-specific model — which is why the benchmark scores the sealed prediction and nothing else. Here, as everywhere in TargetSpace, “understanding” means calibrated predictive skill about future observable states and nothing more.
Every architecture and product argument above reduces to one comparison: the same sealed tasks, one factor changed, the difference in gated target-specific skill read off at a disclosed cost. Memory on/off answers whether persistence improves the model of the user or only retrieval. Evidence-tier ablation answers whether audio, video, location, or physiology earns its cost. Fixing evidence and varying the architecture answers whether a larger model — or an on-device one — extracts more from the same observations (the model–evidence frontier). The minimum-sufficient-tier report answers which raw data can be deleted without losing validated skill. And because evaluation runs federated, these answers come from your own users' data without exporting it. The full decision-to-control map is the paper's builder table; the starter experiment is the two-arm version any team can run.
Recorders and notetakers. Does meeting audio predict commitment follow-through beyond the calendar?
Wearable memory layers. Retrieval or modeling? An enable/disable A–B over sealed tasks answers it.
Always-on capture. How much audio is sufficient, and does sparse sampling match continuous?
Egocentric audiovisual context. Does vision add calibrated lift over audio alone?
Physiology, mobility, location. Marginal lift beyond the routine the R2 baseline already captures?
Copilots over organizational exhaust, for consenting individuals. TS-Enterprise itself is research-status.
Accessibility, memory support, care coordination — under the strictest consent architecture.