EGO DATA/solutions Request samples

Egocentric video for embodied AI

Recorded from
inside the task.

First-person video of people doing ordinary physical work — dishes, repairs, bedding, yard — on head-mounted cameras, in real homes and worksites. Every clip is commissioned: we meet with you, build the shot list around what your model is missing, and record to that spec.

Now showingDriving screws
Taxonomyrepair / assembly
Resolution1920 × 1440
Frame rate29.97 fps
Frames381
Length12.71 s

How we work

We don't sell a catalog. You tell us where your model falls over, we sit down with you, and we go record exactly that.

It starts with a conversation

Before any shot list exists, we want to see the failure cases — the tasks where the policy stalls, the objects it can't place, the grip it never recovers from. That conversation is what the collection gets built around.

Pilot first, then scale

Commission an hour. Look at it properly, on your own loader, against your own eval. If the framing, the pacing, or the level of mess isn't right, we adjust and reshoot before you commit to volume.

Your schema, not ours

Labels, task naming, segment boundaries, and manifest structure follow whatever your pipeline already expects. We'd rather match your format than hand you a conversion job.

What we collect

Everyday physical tasks, in the places they actually happen.

No studio, no scripts, no set dressing. A wearer puts on the camera and does the thing they were going to do anyway, at their own pace, with their own tools. That means both hands stay in frame, objects are where people really leave them, and the clip keeps the fumbles — the dropped sponge, the screw that strips, the second attempt. Recovery behavior is usually the part that's missing from staged datasets.

To specEvery clip commissioned
1 hourMinimum pilot batch
WeeklyRolling delivery once running
NE OhioWearer network, expanding

Sample clips

Five clips from recent collection.

Untouched apart from a resolution pass and stripped audio for the web. Hover or tap any plate to play. Full-resolution files and the JSON manifests that ship with them are in the sample set.

Play

repair / assembly / power driver

Driving screws

Length12.71 s
Frames381
Res1920×1440
Rate29.97
Play

kitchen / dishes / hand-wash

Washing dishes

Length12.28 s
Frames368
Res1920×1440
Rate29.97
Play

cleaning / surfaces / glass

Cleaning a window

Length11.24 s
Frames337
Res1920×1440
Rate29.97
Play

laundry / bedding / duvet

Making a bed

Length6.51 s
Frames195
Res1920×1440
Rate29.97
Play

exterior / yard / clearing

Clearing the yard

Length11.21 s
Frames336
Res1920×1440
Rate29.97

How a collection runs

From task list to delivered clips.

A one-hour pilot batch can be shot and delivered inside a week, which is usually the fastest way to find out whether the footage is what you pictured. Larger collections run as rolling weekly deliveries, so you can start training before the set is finished.

01

Meet and scope the shot list

A call or a visit, and a walk through the cases your model handles badly. That becomes a written shot list with a per-clip spec: what has to be visible, from what angle, how long, how many repeats, how much variation between wearers. You approve it before anyone puts a camera on.

02

Recruit and consent

Wearers are recruited for the environments and demographics your spec calls for. Everyone signs a release that explicitly covers machine-learning training and redistribution of the footage, and everyone is paid.

03

Capture in situ

Head-mounted capture in the wearer's own kitchen, garage, or yard. We don't tidy up first and we don't re-shoot for neatness. The environment is the point.

04

Review, redact, label

Every clip is watched end to end. Bystander faces, screens, and documents are blurred on request. Labels go to your schema — task boundaries, subtask segments, hand–object contact, tool in hand, or spoken narration.

05

Deliver and iterate

Clips plus a JSON manifest, pushed to your bucket. You flag what's short and the next batch corrects for it.

Delivery specs

What lands in your bucket.

Defaults are listed below. Most of it is negotiable — tell us what your loader expects and we'll match it rather than making you write a conversion step.

Resolution
1920 × 1440, 4:3 native sensor crop. Downscaled derivatives on request.
Frame rate
29.97 fps constant, matching the capture hardware. Higher rates for fast-manipulation tasks on request.
Video codec
H.264 high profile, MP4 container. ProRes or camera-original on request.
Audio
Stripped by default. Included or transcribed on request.
x
Clip length
From 5-second quick actions to uncut multi-hour sessions, per your spec.
Annotation
Task and subtask boundaries, hand–object contact frames, tool-in-hand flags, free-text narration. Your schema, not ours.
Manifest
One JSON per clip: task path, duration, frame count, environment type, wearer ID, consent record ID, redaction log.
Licensing
Perpetual, non-exclusive, worldwide for model training and evaluation. Exclusive collection available per project.
Delivery
S3, GCS, or Azure blob. Signed-URL download for smaller sets.

Start a collection

Tell us what you need recorded.

Send a task list, or just a description of the behavior your model keeps getting wrong — the second one is fine, and it's usually where the useful conversation starts. We'll set up a call, come back with a shot list, a per-hour price, and a date. The sample set goes out the same day you ask.