Benchmark / Overview

What TargetSpace-Bench measures

TargetSpace-Bench evaluates target-specific prediction under partial observation: whether a system can maintain and use a target-specific predictive representation for a specific target from passive, multimodal, non-IID observation — measured through calibrated, sealed prediction.

1capability: maintain a target-specific predictive representation
R2headline baseline: the target's own routine
5+6domain tracks & evaluation splits
bitsskill, by strictly proper scoring

Definition & scope

A target

A persistent, evolving entity tracked over time: a person, agent, system, organization, environment, or process. The flagship track targets a consenting individual; other tracks target patients, energy assets, embodied agents, and projects.

A target state

The latent configuration a target acts to reach or maintain — a commitment, priority, constraint, or regime — evaluated only through its externally observable consequences. The object to prediction is the transition between target states.

Passive longitudinal observation

Evidence accrued over time without acting on the target: metadata, text, audio, passive multimodal streams, location, physiology. The benchmark measures whether richer, correctly-ordered observation improves target-specific predictions.

The capability under test

Can a model maintain and use a target-specific predictive representation, rather than merely a model of generic scenes or generic events? A high score means calibrated, prospective, target-specific skill — nothing about a target's inner life.

A worked example, end to end synthetic · illustrative

One pass through the TargetSpace loop on a made-up target with made-up numbers — no participant data, no empirical result. The loop:

observation representation prediction calibration evidence minimization validation

1 · Target & passive evidence stream

The target is a synthetic person. Two instrumented streams record timestamped evidence — calendar metadata and app-focus samples — without the target reporting, recalling, or performing anything for the measurement itself:

  • days 1–7 — a recurring “draft spec review” block sits each weekday at 09:00; on days 3, 5, and 7 it is rescheduled to later the same day.
  • day 7, 16:10 — app focus shifts to a competing deliverable and stays there for the rest of the session.
  • day 7, 18:02 — a meeting request from another team lands on the day-8 09:00 slot.
  • day 8, 07:35 — the morning’s first app-focus samples are on the competing deliverable again.

2 · Cutoff, query & possibility space

Evidence seals at day 8, 08:00. The sealed query: does the “draft spec review” commitment complete, defer, cancel, or get replaced? Before the resolution window closes, several future target states remain operationally possible:

complete defer cancel replace

The benchmark asks one question: does the system assign better-calibrated probabilities over these possibilities than the baselines do?

3 · Sealed distributions vs the baselines

Possible stateModel (sealed)R1 population priorR2 own-routine
complete0.140.550.70
defer  resolved outcome0.780.250.18
cancel0.040.100.06
replace0.040.100.06

4 · Resolution & per-instance score

Resolved outcomedefercalendar item moved before its start time — pre-registered, deterministic rule
Skill vs R1, this instance+1.64bits — log2 0.78 − log2 0.25
Skill vs R2, this instance+2.12bits — log2 0.78 − log2 0.18

Illustrative single-instance numbers only. A reportable run aggregates many sealed instances per target, and skill counts only after the calibration and wrong-target permutation gates — genuine target-specific skill collapses under wrong-target permutation.

Evidence cost. Of the two streams, the app-focus samples are the more invasive — and that stream must earn its lift. The objective is validated lift from minimum sufficient observation, not maximal capture: at equal gated skill, the less invasive configuration wins.

A full end-to-end prototype (registry, sealing, scoring, evidence efficiency) runs at targetspace.ai/platform.

Different from generic dynamics

TargetSpace is not a video-realism, robotics-manipulation, or intuitive-physics benchmark. Those evaluate generic dynamics, realism, and control. TargetSpace evaluates target-specific dynamics: whether prediction improves when a model is given the correct target's history in the correct temporal order. A system can render plausible scenes, generate fluent continuations, or execute dexterous control and still be unable to track this target and prediction where it turns next.

How it complements other benchmark families

The families are complementary, not competing. Each is the right instrument for a different question; TargetSpace is compared with them only on shared axes.

Benchmark familyPrimary object of evaluationTypical inputTypical outputWhat it missesHow TargetSpace complements it
Physical reasoning / intuitive physicsphysical plausibility, object permanence, causality, spatial continuityshort scenes / clipsplausibility or violation judgmenta persistent target; longitudinal adaptation; calibration over timeadds target-specific dynamics over generic physical law
Video generation / world simulationvisual realism, temporal consistency, plausible scene evolutioncontext framesgenerated continuationtarget identity; sealed prospective scoring; proper calibrationscores latent target-state transitions, not surface reconstruction
Embodied roboticsaction utility, policy evaluation, manipulation / control successproprioception, sensors, actionsactions; achieved configurationpassive longitudinal inference; calibration; an own-routine baselinescores passive consequence predictions, not control
Symbolic / event predictioningprobabilistic prediction of public eventsquestion + contextcalibrated probabilitya tracked individual target; own-routine R2; permutation specificityadds the target as the unit, with R2 + permutation controls
Agent memory / personalizationrecall, preference modeling, retrieval QAhistory / profile + queryheld-out response / preferenceprospective sealing; calibration; transition predictioningscores why an episode matters and predictions the next transition, sealed
TargetSpace-Bench (this work)target-specific prediction under partial observationpassive multimodal observation up to sealed Tcalibrated prediction over target-state transitionsby design: physical realism, control, generation fidelityis the complementary layer the other families omit

Built like the benchmarks researchers trust

We borrow the structure and seriousness of established efforts — not their branding.

Challenge & leaderboard ARC-style

A clear mission, a public leaderboard, and explicit benchmark versions — with contamination-resistant, prospective evaluation rather than a static answer key.

Submissions & splits SWE-bench-style

Defined task splits, a submission pipeline, and a verification path so leaderboard credibility rests on reproducibility, not self-report.

Transparent evaluation HELM-style

An explicit evaluation philosophy: scenarios (tracks/splits), multiple metrics reported side by side, and calibration treated as first-class.

Governance & verification MLCommons-style

Versioned rules, official vs unofficial (public/verified/private-eval) submissions, and organizer-run private evaluation for high-stakes claims.