Dataset Report

Project Bumblebee

14,407 curated hours across four synchronized streams, measured on the axes that decide whether egocentric data is trainable.

  • Curated hours14,407 across all streams
  • Capture streamsFour synchronized modalities
  • FormatFHD–4K · 30–120 fps
  • DuplicatesZero across all modalities
Representative egocentric capture from the household split.

Household egocentric

Curated hours
11,778
Clips
29,247
Participants
1,802
Tasks / domains
50 / 9

Factory egocentric

Curated hours
2,544
Clips
1,283
Participants
482
Tasks / families
52 / 10

Tri-camera

Curated hours
75
Clips (100×3)
300
Tasks
100
Viewpoints
Head + 2 wrists

Motion capture

Curated hours
9.7
Clips (72×4)
288
Tasks
18
Output
Metric 3D joints
Foreword

Editor’s Note

Navaneeth Bodla, Head of AI, Deccan AI

Robotics has a data problem, but it isn't a data-volume problem.

Large language models scaled because the internet already contained dense semantic information: ideas, instructions, conversations, code, and reasoning. Video foundation models have benefited from a similar advantage. Internet video contains rich priors about motion, geometry, causality, object permanence, human behavior, and everyday physics.

Robotics builds on those priors rather than relearning the physical world from scratch.

The missing layer is not more generic pretraining data. It is high-quality mid-training data that adapts video foundation models from broad world understanding to egocentric behavior and, eventually, robot action. This data needs dense interactions, clear hand visibility, long-horizon tasks, diverse environments, failures, recoveries, and signals beyond video such as body motion, contact, pressure, and force.

The path may seem indirect: first learn from real-world video, then from embodied human behavior, and finally from robot-specific data. But the alternatives are difficult to scale.

Autonomous driving benefits from using nearly the same vehicles, sensors, and action spaces for both collection and deployment. General robotics does not. Robots differ in cameras, kinematics, grippers, control rates, and physical limits, making large-scale teleoperation expensive and platform-specific. Simulation is essential, but it also depends on grounded data about contact, materials, friction, and failure.

There is also a perception bottleneck. Many action models running on robots today operate on relatively low-resolution visual inputs, making precise manipulation difficult when objects, fingers, or contact points occupy only a small part of the frame. Multi-camera data helps close this gap by combining global context with close-range wrist views, much like autonomous vehicles use multiple cameras to cover different parts of the road and reduce blind spots.

Humans remain the most scalable source of rich manipulation behavior, but human demonstrations must be captured in a form that preserves what can transfer across embodiments: intent, task structure, trajectories, contact events, coordination, and recovery strategies.

Project Bumblebee is built around this premise.

It combines egocentric video, synchronized head-and-wrist views, motion capture, and embodied interaction signals to create an adaptation layer between video foundation models and robot-specific training. The goal is not the largest archive, but deliberate coverage of useful tasks, environments, viewpoints, interactions, failures, and recoveries.

I am grateful to the researchers, engineers, operators, reviewers, and program teams across Deccan AI who contributed to Bumblebee. It reflects the care required to build robotics data with measurable quality, clear provenance, and meaningful coverage.

Navaneeth Bodla Former Tech Lead, Adobe Video Generation Models
Reader’s Guide

Guide to Reading the Report

This report is written to be read end to end. If you are short on time, start where your role sits.

// 01 — Research & ML Engineering (VLA researchers, world-model teams, perception engineers)

Section 2 for the quantitative portrait. Sections 3 through 5 for lighting, composition, and hand visibility in depth. Appendices A through D carry the full per-task distributions you will filter and sample against.

// 02 — Data & Procurement (Heads of data, legal, dataset licensing)

Section 6 for the collection protocol, Section 9 for consent, ethics, and licensing terms, Section 7 for the curation funnel your auditors will want to walk through. Independent-auditor access under NDA is described in Section 9.

// 03 — Strategy & Leadership (Founders, CTOs, Heads of AI)

The abstract and Section 2 for what the corpus is. Sections 10 and 11 for where it is going: the hand-object interaction layer, new capture modalities, and the data pyramid this corpus anchors.

Overview

What an hour of footage actually contains

Egocentric datasets are usually sold by total hours. For robotics that is the wrong headline. A manipulation policy needs hand-visible frames, a world model needs temporal structure, and an eval team needs certainty that no two clips are the same video twice. Bumblebee is built and reported against those properties.

14,407Curated hours across all four streams
31,118Released clips, each with versioned metadata
2,284Household and factory participants, plus 4 MoCap performers
100%Informed consent, commercial use included
The four streams

Each stream answers a different question

Two egocentric domains give breadth and precision. A tri-camera rig adds per-hand close-up observation. Marker-based motion capture supplies the metric skeleton everything else is validated against.

Household egocentric

11,778 hours across nine activity domains

Weighted toward the cleaning, organization, and cooking work that dominates domestic settings. The distribution is intentionally non-uniform: it follows where hands-on manipulation is densest rather than sampling domains evenly.

Household recorded hours by activity domain: cleaning 4,248; organization 2,387; cooking and food prep 1,820; laundry management 1,355; daily routines 612; assistance and navigation 538; preparation and occasions 381; hosting 258; home maintenance 179.
Activity domainClipsHours%
Cleaning12,0214,24836%
Organization6,6972,38720%
Cooking & food prep2,4191,82015%
Laundry management4,4881,35512%
Daily routines7096125%
Assistance & navigation1,9335385%
Preparation & occasions4453813%
Hosting3052582%
Home maintenance2301792%
Total29,24711,778100%

50 goal-directed tasks, grouped into these nine domains. Household footage also carries a tri-granular labeling of verb, object, and goal.

Factory egocentric

2,544 hours across ten operation families

Grouped by the manipulation skill each station demands rather than by location — all of this footage comes from one environment, the assembly floor. Three families carry 65% of factory hours: the tight-tolerance, high-repetition operations where policy performance matters most.

Factory recorded hours by operation family: component insertion and assembly 561; inspection and quality verification 549; precision soldering and joining 546; electrical test and debug 289; packaging and kitting 160; material processing and forming 155; machine tending and material handling 119; surface preparation and cleaning 64; marking and labeling 64; line supervision and unclassified 37.
Operation familyTasksHours%
Component insertion & assembly556022%
Inspection & quality verification954922%
Precision soldering & joining454621%
Electrical test & debug828811%
Packaging & kitting51606%
Material processing & forming61556%
Machine tending & handling51205%
Surface preparation & cleaning4653%
Marking & labeling3643%
Line supervision & unclassified3371%
Total522,544100%

52 distinct tasks recorded on active production lines. Per-task clip counts were not recorded for this split, so hours are reported without a clip column; the split comprises 1,283 clips in total.

Factory · deeper cut

The fifteen heaviest of 52 tasks

Top fifteen factory tasks by recorded hours, led by manual soldering and touch-up at 478 hours and manual insertion at 319 hours, down to marking on LCD at 33 hours.

Manual soldering and touch-up alone carries 478 hours — nearly a fifth of the factory split in a single operation. The concentration is what makes this data useful: deep coverage of the few operations that dominate a real line, rather than thin coverage spread evenly across all 52.

The remaining 37 tasks account for 579 hours. The full per-task listing, grouped by family, is in the report appendix.

Tri-camera

Head plus one camera per wrist

All three views are first-person. The wrist views resolve grasp and contact detail at a scale the head camera cannot, and they keep the hand in frame when the head view is blocked by the body or the object. This is the observation structure bimanual policies are usually conditioned on — hand-level resolution, not an external vantage point.

100Standardized tasks
3Synchronized views
300Clips (100 × 3)
75Curated hours

Each task is captured from all three views, so 100 tasks × 3 views = 300 clips. Hours are footage summed across the three views, at 15 minutes per view.

Motion capture

Metric ground truth, 18 task categories

Marker-based optical capture gives metric 3D joint trajectories — the reference that pixel-space pose estimation is validated against, and the basis for retargeting to humanoid and bimanual kinematic chains.

Motion-capture recorded minutes by task, led by packing at 106.3 minutes and arranging household items at 59.0 minutes, down to box transfer at 7.7 minutes.

18 tasks × 4 performers = 72 takes; each take is captured from 4 synchronized views, so 72 × 4 = 288 clips carrying 580 minutes. Team work and box transfer are two-performer interaction tasks.

What the data looks like

Frames from each capture configuration

One sample per configuration, at full width. Every frame below is from the released corpus.

Grid of representative egocentric frames spanning household scenes, actors, and lighting conditions.
Household egocentric Representative frames across household scenes, actors, and lighting conditions — the spread the lighting distribution describes numerically.
Grid of tri-camera frames. Each row is a bimanual manipulation task; the three frames in a row are the same time step from the head-mounted camera and one camera on each wrist.
Tri-camera · head + dual wrist Each row is a different bimanual task. The three frames within a row are the same time step, seen simultaneously from the head-mounted camera and one camera on each wrist.
Grid of motion-capture output. Rows show the three released streams: 3D pose render, solved skeleton, and RGB reference video, five samples each.
Motion capture · three released streams One row per released stream: 3D pose render (top), solved skeleton (middle), and RGB reference video (bottom), with five samples drawn from each.
Quality signals

Representative Quality Analysis

Hand visibility, lighting, and temporal integrity are measured on a random subset of 7,800 household clips drawn across all nine activity domains.

Hands

Hand visibility, quantified

Hand-visible frames (≥ 1 hand) 68%
Videos · both hands ≥ 20% of frames 45%
Videos · both hands ≥ 50% of frames 26%

Hand visibility is the most common filter teams apply to egocentric footage. Both-hands coverage, the signal that matters for bimanual work, is released per video at every threshold so a team can pick its own operating point and trade coverage against volume.

Lighting

Covering the deployment envelope

Scene-brightness distribution: share of videos by average luma normalized 0 to 1, a broad bell-shaped curve centred near mid-range with mean and median marked.

Scene brightness is measured directly as mean frame luminance rather than binned into coarse categories. The distribution is bell-shaped and centred near the middle of the normalized 0–1 range, corpus mean 0.44, with meaningful mass in both the darker and brighter tails. Deployed robots do not control their lighting, so spread is the property that matters — not a corpus shot under convenient studio light.

Temporal structure

Long-horizon material at scale

24Mean clip minutes
3,593Recordings ≥ 40 min
4,225Long-horizon hours
Histogram of clip durations across the corpus, with the mean and median marked. Duration histogram across all long-horizon recordings of at least 40 minutes, extending into multi-hour material.

The distribution is bimodal: short manipulation vignettes (88% of clips, mean 18 min) and long-horizon recordings (12% of clips, mean 71 min). The latter are an eighth of the clips but carry over a third of all hours.

Capture format

Native fidelity preserved

Share of parsed clips by native resolution: 1080p 70.2%, 720p 11.5%, 1080p portrait 7.4%, 4K UHD 5.6%, 720p portrait 1.5%, other 3.8%. Share of parsed clips by frame rate: 30 fps dominant, with substantial 60 fps and 50 fps and a 120 fps tail.

Clips ship at native capture resolution and frame rate rather than transcoded to a common target, so teams select the fidelity their pipeline expects and can study robustness to capture conditions instead of having it flattened away. The 120 fps tail resolves fast hand motion without blur.

Temporal integrity

Cuts are documented, not hidden

43%Clips free of jump cuts
0Undocumented cuts

Jump cuts and unannounced scene breaks corrupt sequence models. Clips that contain them ship with exact cut timestamps so downstream users can split or mask cleanly. Every cut identified in QC is documented with a timestamp rather than left implicit.

Uniqueness

Every clip is distinct

0Duplicates in the release
218Candidates removed pre-release

Duplicates inflate apparent scale, leak across eval splits, and let models overfit specific footage. Duplicate detection was run across the corpus and 218 candidates were removed before release.

  • Frame-exact synchronization across all four streams
  • Chain-of-custody documented from capture to release
  • No identifiable non-consenting individual in any frame
Task taxonomy

Two domains, two taxonomies

Household footage carries a tri-granular labeling — verb, object, goal. Factory tasks carry station-level identity instead, because a factory task name already denotes a fixed, repeatable operation on a known part.

510Verbs
2,140Objects
50Household goal tasks
9Household domains
52Factory tasks
10Factory operation families
Most-mentioned verb clusters
pick upwipeplacepressopen adjustfoldmovepourarrange closesmoothhandleinsertrotate removeinspectsprayturn onrub cutrepositionhangdipattach gathersortstirscoopspread
Most-mentioned object clusters
clothdrawerbowlclothesshirt bottlehandfloorcontainerstove pillowmopsinkbaglid cableboxbucketframeitem

Failure and recovery events — dropped objects, mis-grasps, corrections — are retained rather than filtered, across both splits. Curated imitation-learning data often removes them.

Curation

The QC behind the numbers

Every shipped clip cleared six stages in order, with attrition documented per stage so the material screened out is accounted for rather than hidden.

Funnel

Six stages, attrition documented

Six-stage curation funnel: ingestion validation 100%, automated quality screen 84%, VLM semantic verification 68%, human review 54%, duplicate elimination 44%, statistical audit and release 34%.

Per-stage attrition is documented and released to licensees. The ordering is deliberate: reported numbers change when the pipeline changes, not when a metric falls outside its target range.

Stages

What each stage rejects

  • 1Ingestion validation. Container, codec, timestamp monotonicity, per-frame decode.
  • 2Automated quality screen. Motion blur, sustained under- or over-exposure, incoherent ego-motion, audio corruption, encoding artifacts.
  • 3VLM semantic verification. A vision–language model checks scene-label agreement, hand-presence consistency, and safety-relevant properties. It is calibrated against human-in-the-loop review, and low-confidence cases escalate.
  • 4Human review. Uniform random-sample audit for calibration, plus targeted review of anything flagged upstream.
  • 5Duplicate elimination. Re-encodings, format conversions, and clips sharing substantial content despite differing framing or compression.
  • 6Statistical audit & release. Confidence intervals on every published metric.

Model choices, prompt structures, thresholds, and reviewer rubrics are not disclosed. Independent auditors are granted full access under NDA — the balance we think is right between verifiability and not handing over the methodology wholesale.

Annotation

Every clip is tagged for action, hand, and object

Our curation funnel decides which clips ship. This pipeline, built and run in house, decides what each one says: the action, its sub-steps, which hand, and which object. A vision–language model supplies breadth across every clip and frame; trained annotators set the standard of correctness, and every correction they make is versioned on top of the machine output.

The layer, playing

Hand boxes, chunk captions, and a subtask timeline, per episode

Laundry preparationBath mat folded and placed into a hamper; human and model chunks side by side.
Cooking preparationVegetables carried to the sink and rinsed; subtasks segmented at contact boundaries.
Room navigationBucket repositioned between rooms; scene class changes with the room.

Screen captures of the review tool on three released household episodes. Left and right hand boxes are drawn per frame; the panel carries the model's chunk caption, its scene class, and the timestamped subtask breakdown, with the human-written chunk for the same span above it. The two tracks at the foot of each clip are the human and model chunk timelines for the whole episode, so the granularity difference between the two passes is visible directly. Pause or scrub any clip to read a panel in full.

Machine pass

Two scales, one frame store

  • 1Frame store. Each approved episode is decoded once at ~1 fps with a manifest of exact frame timestamps. Every annotation is keyed to that manifest, never to container frame counts, because fragmented sources report those wrong and an annotation layer that trusts them drifts silently.
  • 2Chunking. The timeline is split into contiguous 15–45 s action units (30 s target): long enough to hold an approach, a contact and a release, short enough that one description of it stays specific.
  • 3Chunk-scale reasoning. Per chunk: a caption, a scene class from a closed vocabulary with confidence, and a segmentation into 1–6 contiguous, non-overlapping subtask intervals (reach, grasp, move, place, release, adjust), each with its own bounds and confidence. Idle stretches are labelled idle, not forced into an action.
  • 4Frame-scale reasoning. One short hand–object sentence per frame: what the hands are doing and what they touch or hold. Handedness is resolved from position in the frame with no dominant-hand default and an explicit unlabelled fallback. Side inversion is invisible in aggregate statistics and poisonous to bimanual policies.
  • 5Spatial track. Hand and manipulated-object boxes plus hand → object interaction links come from a specialist egocentric detector on GPU, not from the language model. That is a measured call: VLM box grounding was not at shippable quality, and object names are paid for once per chunk rather than once per box.
  • 6Reduction. Tracks are joined on frame timestamp into a frame table, chunk records, and a manifest naming which tracks ran, with model versions and timings. The manifest is written last, so its presence marks completion and reruns stay idempotent.

Model output is requested as strict JSON and parsed defensively: spans clamped to their chunk, confidences clamped, malformed items dropped, complete entries salvaged from truncated replies. Tracks fail soft: a spatial-track failure is recorded and the episode still ships its language layer.

Human in the loop

Refined, not overwritten

  • Machine output is version 0 and is never overwritten; refinements are new versions on top of it
  • Annotators work in a player that draws boxes, subtask timeline and captions over the video; editing is keyboard-first, so frame-level corrections stay quick and comfortable to make
  • Box move, resize, add, delete and propagate-forward; label rename with autocomplete over labels already seen in the episode
  • Subtask retiming, relabelling, split and merge; caption, scene and hand-description rewriting; mark-reviewed
  • One short-lived writer lease per episode; every save is a single transaction conditioned on a live lease and the version the annotator actually loaded
  • Client-generated edit ids mean a retried save is applied once, not twice; a save on a stale version is surfaced as a conflict so the annotator can reapply it on the current version, never silently merged
  • Append-only, attributed history: any released annotation traces back through every human edit to the machine version it started from, with credit to the people who made it
  • Uniform audit sample plus targeted review of low-confidence and flagged episodes

Disagreements found in review are what drive prompt and threshold revisions, which is why annotation review is part of the pipeline rather than a check appended to its end. Models, prompt text, vocabularies and thresholds are not published; auditors get access under NDA.

Schema

The annotation layer, per episode

Annotations {
  chunks[]  : { chunk_id, t_start_s, t_end_s, caption,
               scene_class, scene_confidence,
               subtasks[]: { t_start_s, t_end_s, label, confidence } }
  frames[]  : { frame_idx, ts_s, hand_object_description,
               detections[]: { track_id, label, confidence,
                              box2d{ x1, y1, x2, y2 },
                              interacts_with, interaction_prob } }
  manifest  : { schema_version, tracks_completed[], model_versions,
               timings_s, frame_count, chunk_count }
  provenance: { version, base_version, author, edited_at }
}

One row per decoded frame is emitted even when both tracks are empty, so consumers never have to special-case gaps. track_id carries the role (left hand, right hand, first object), while label carries the open-vocabulary object name. Boxes are normalized to frame coordinates. The layer ships against the same versioned schema as the per-clip metadata, additively.

Quality signals

Measured, not asserted

Hand visibility, lighting, and temporal integrity are measured on a random subset of 7,800 household clips.

Hands

Hand visibility, quantified

Frames with at least one visible hand 68%
Videos · both hands > 20% of frames 45%

Hand visibility is one of the most common filters teams apply to egocentric footage, so we measure it directly. Every frame in a random subset of household clips is scored for hand presence and aggregated per video, giving both the at-least-one-hand figure and the stricter both-hands figure above. These are visibility measurements, and we state their sampling basis alongside them.

  • Wrist-mounted cameras on the tri-camera stream keep the hand in frame when the head view is blocked, so visibility does not depend on head orientation

Next release. Visibility is the outer layer. Our immediate next deliverable reports the interactions themselves, contact and grasp structure across the released corpus: see Hand–object interaction analysis.

Lighting

Covering the deployment envelope

Scene-brightness distribution: share of videos by average luma normalized 0 to 1, a broad bell-shaped curve centred near mid-range with mean and median marked.

Scene brightness is measured directly as mean frame luminance rather than binned into coarse categories. The distribution is bell-shaped and centred near the middle of the normalized 0–1 range, corpus mean 0.44, with meaningful mass in both the darker and brighter tails. Deployed robots do not control their lighting, so spread is the property that matters, not a corpus shot under convenient studio light.

Temporal structure

Long-horizon material at scale

24Mean clip minutes
3,593Recordings ≥ 40 min
4,225Long-horizon hours
Histogram of clip durations across the corpus, with the mean and median marked. Duration histogram across all long-horizon recordings of at least 40 minutes, extending into multi-hour material.

The distribution is bimodal: short manipulation vignettes (88% of clips, mean 18 min) and long-horizon recordings (12% of clips, mean 71 min). The latter are an eighth of the clips but carry over a third of all hours.

Capture format

Native fidelity preserved

Share of parsed clips by native resolution: 1080p 70.2%, 720p 11.5%, 1080p portrait 7.4%, 4K UHD 5.6%, 720p portrait 1.5%, other 3.8%. Share of parsed clips by frame rate: 30 fps dominant, with substantial 60 fps and 50 fps and a 120 fps tail.

Clips ship at native capture resolution and frame rate rather than transcoded to a common target, so teams select the fidelity their pipeline expects and can study robustness to capture conditions instead of having it flattened away. The 120 fps tail resolves fast hand motion without blur.

Temporal integrity

Cuts are documented, not hidden

43%Clips free of jump cuts
0Undocumented cuts

Jump cuts and unannounced scene breaks corrupt sequence models. Clips that contain them ship with exact cut timestamps so downstream users can split or mask cleanly. Every cut identified in QC is documented with a timestamp rather than left implicit.

Uniqueness

Every clip is distinct

0Duplicates in the release
218Candidates removed pre-release

Duplicates inflate apparent scale, leak across eval splits, and let models overfit specific footage. Duplicate detection was run across the corpus and 218 candidates were removed before release.

  • Frame-exact synchronization across all four streams
  • Chain-of-custody documented from capture to release
  • No identifiable non-consenting individual in any frame
Task taxonomy

Domain-specific taxonomies

Household footage carries a tri-granular labeling: verb, object, goal. Factory tasks carry station-level identity instead, because a factory task name already denotes a fixed, repeatable operation on a known part.

510Verbs
2,140Objects
50Household goal tasks
9Household domains
52Factory tasks
10Factory operation families
Most-mentioned verb clusters
pick upwipeplacepressopen adjustfoldmovepourarrange closesmoothhandleinsertrotate removeinspectsprayturn onrub cutrepositionhangdipattach gathersortstirscoopspread
Most-mentioned object clusters
clothdrawerbowlclothesshirt bottlehandfloorcontainerstove pillowmopsinkbaglid cableboxbucketframeitem

Failure and recovery events, including dropped objects, mis-grasps and corrections, are retained rather than filtered, across both splits. Curated imitation-learning data often removes them.

Curation

The QC behind the numbers

Every shipped clip cleared six stages in order, with attrition documented per stage so the material screened out is accounted for rather than hidden.

Funnel

Six stages, attrition documented

The six curation stages in order: ingestion validation, automated quality screen, VLM semantic verification, human review, duplicate elimination, statistical audit and release.

Each stage is applied in sequence to every candidate clip. Per-stage attrition is documented and released to licensees rather than published here.

Stages

What each stage rejects

  • 1Ingestion validation. Container, codec, timestamp monotonicity, per-frame decode.
  • 2Automated quality screen. Motion blur, sustained under- or over-exposure, incoherent ego-motion, audio corruption, encoding artifacts.
  • 3VLM semantic verification. A vision–language model checks scene-label agreement, hand-presence consistency, and safety-relevant properties. It is calibrated against human-in-the-loop review, and low-confidence cases escalate.
  • 4Human review. Uniform random-sample audit for calibration, plus targeted review of anything flagged upstream.
  • 5Duplicate elimination. Re-encodings, format conversions, and clips sharing substantial content despite differing framing or compression.
  • 6Statistical audit & release. Final audit before packaging, with structured metadata attached to every clip.

Model choices, prompt structures, thresholds, and reviewer rubrics are not disclosed. Independent auditors are granted full access under NDA. That is the balance we think is right between verifiability and not handing over the methodology wholesale.

Delivery

What you actually receive

Every clip in all four streams ships with structured metadata, versioned alongside the data release. Schema changes stay backward-compatible within a major version.

Schema

Per-clip metadata

Clip {
  id          : uuid
  stream      : household | factory | tricam | mocap
  timestamps  : { capture_utc, duration_ms }
  scene       : { activity_domain | operation_family, task }
  lighting    : { avg_luma_norm }
  hands       : { pct_visible, both_hands_pct }
  task        : { verb, object, goal }
  quality     : { blur, exposure, ego_motion }
  integrity   : { jumpcut_count, cut_timestamps[], dedup_hashes }
  views       : { mount, sync_offset_ms }
  consent     : { participant_id, commercial_ok, jurisdiction }
}

Tri-camera metadata adds the mount position of each view and inter-view sync offsets. Motion-capture metadata adds per-take timecode, performer identity, approval status, and the synchronized reference-video pointer.

Coverage & licensing

Commercial-training ready

Core metadata 100%
Narration 100%
Hand-visibility stats 100%
Informed consent 100%
  • License supports commercial policy training and downstream model release
  • Informed consent documented for every participant, across all four streams
  • Bystander and PII handling per jurisdiction-appropriate policy
  • Filled datasheet and chain-of-custody with every release
  • Independent audit of consent, redaction, and provenance under NDA
What comes next

A living corpus

Collection continues beyond this release. The next deliverable goes one level deeper than clips and tasks, to the interactions themselves.

Immediate next release

Hand–object interaction analysis

A dedicated report over the released corpus: which objects are contacted, when contact begins and ends, how each grasp is formed and released, and how bimanual coordination is distributed across tasks. Contact structure is what separates footage a policy can learn manipulation from, and footage that merely shows a task being completed — and it is the axis teams most often have to annotate themselves before egocentric data becomes trainable.

Because the analysis runs over the corpus already released, it requires no re-collection and no change to the data licensees hold. It ships as an additive layer against the same versioned metadata schema.

01 — Richer modalities

More of what teams ask for

Instrumented tactile gloves for contact, grip force, and finger pose. UMI-style handheld grippers that record from the gripper's point of view rather than the wearer's, which transfers more directly to real arms. Eye-gaze tracking as a signal of attention and intent. Denser depth. Every new modality plugs into the same synchronization and metadata schema, so the corpus grows as one dataset rather than splitting into incompatible pieces.

02 — Long-tail environments

Coverage, not raw hours

Hours are no longer the bottleneck; rare situations are. We use the corpus's own metadata and clip embeddings to find environments that appear less often in our data than in the real world — cramped or dimly lit kitchens, shared and multi-generational households, festival preparation, aging appliances, unusual factory cells. Each gap becomes a collection brief, so rare settings are captured deliberately instead of by luck.

03 — The data pyramid

Human base, synthetic middle, robot top

The base is large-scale human data; Bumblebee sits at its high-quality end, first-person and synchronized and task-directed rather than scraped. The middle is simulation and generative augmentation, cheap to scale but only approximating real physics. The top is teleoperated and autonomous robot data, exact but expensive. Because every layer shares one task taxonomy and schema, data from one can train and validate models built on another.

04 — Adaptive distribution

Balance while collecting, not after

Today balance comes from collection briefs and after-the-fact curation. We are closing the loop: dashboards already show in near real time how much each task and scene contributes, and we are wiring those numbers back into task assignment. Once a category hits its target share, contributors are redirected to categories still short — so drift surfaces as an alert during collection instead of a surprise at release.

Access

We would be glad to hear from you

Bumblebee is available under a license that supports commercial policy training. A filled datasheet and chain-of-custody documentation are included with every release, and independent auditors are welcome to review consent, redaction, and pipeline detail under NDA.

License
Commercial-training ready
Cadence
Versioned · additive
Datasheet
Ships with release
Audit
Available under NDA