Inputs
Egocentric or external video, robot episodes, sensor streams, state, actions, force or tactile data when present, task protocol, and target schema.
Turn demonstrations and robot logs into synchronized episodes with task phases, actions, object state, outcomes, failures, and uncertainty represented as separate signals.
Human video, robot state, controller actions, force, and outcomes are different signals. The program only claims what the source data can support.
Egocentric or external video, robot episodes, sensor streams, state, actions, force or tactile data when present, task protocol, and target schema.
Capture planning, episode filtering, temporal annotation, object and pose work, language grounding, synchronization review, and schema mapping.
Accepted episodes, stream references, phases, interactions, outcomes, failures, tracks, provenance, rejection reasons, and a dataset card.
Protocol tests, file and stream checks, timestamp audits, boundary review, failure review, duplicate detection, and representative loader validation.
The buyer's native episode format remains authoritative, with an agreed compatibility export when useful and technically supported.
The program distinguishes capture, curation, annotation, and conversion so a buyer can see where each signal came from.
Design buyer-defined protocols, qualify participants and devices, validate views and task resets, and record consent, rights, and allowed metadata.
Find broken runs, missing streams, invalid resets, duplicates, interventions, recovery attempts, rare failures, and underrepresented outcomes.
Mark task phases, actions, contacts, object states, outcomes, failures, recovery windows, poses, tracks, and grounded instructions as scoped.
Review camera and sensor timing, map native fields, preserve explicit nulls, and test representative episodes in the buyer's loader.
Physical data gets expensive when an unusable view, missing stream, or ambiguous outcome is discovered after the run.
Name the embodiment, environment, inputs, expected signals, outcome rule, rights, and target model job.
Output: task and data specificationReview representative successes, failures, resets, missing streams, timing issues, and native loader behavior.
Output: capture or curation mapRun a small batch, test collection or labeling rules, and confirm what can be observed versus inferred.
Output: accepted sample and guideValidate incoming data, reject unusable runs, annotate the supported signals, and route ambiguous cases to review.
Output: reviewed episode setCheck synchronization, schema, manifests, explicit nulls, and a representative delivery in the buyer's loader.
Output: episodes, QA report, and rework logForcing unusable recordings or broken logs through annotation creates false confidence. The rejection record is part of the deliverable.
Representative sample
This example shows the delivery shape. It does not claim that human video contains controller state or that visible motion proves contact.
Bring the task, embodiment, environment, available signals, and the way the model team loads an episode.
No. Human video can show observations, motion, visible interaction, and outcomes. Robot controller actions, state, force, and tactile signals must come from their own sources.
Collection can be scoped when the task, environment, participant, rights, device, safety, and quality requirements are defined and supported.
The buyer's native representation is the first delivery target. Flinket maps and tests a representative sample before production scale.
Failures, interventions, recovery attempts, incomplete runs, and unusable episodes receive explicit reasons rather than being silently excluded or forced into a success schema.
Receive a capture or curation plan, one sample annotated episode, a rejection taxonomy, and a scale estimate.